# Make ↔ Numu Orchestration Contract (repository-based)

> Read-only specification. No Make scenarios or backend wiring are implemented.
> Every endpoint, payload, and behavior below is verified against the actual
> repository (routes, controllers, services). The repository is the source of truth.

**Verified against:** `feat/agent-knowledge-cache` @ `8e892f9`.
**Transport for all endpoints:** prefix `/api/v1/ai`, middleware `McpAuth:full` + `ai.cb` (circuit breaker) + `throttle:600,1`.
**Auth:** OAuth 2.1 bearer (`naat_…`, scope `numu:write`, audience-bound) **or** Sanctum PAT (ability `numu:full`). Missing/invalid token → **401**; insufficient scope → **403**; token in query string → **400**. (`app/Http/Middleware/McpAuth.php`.)
**Envelope:** success `{ "data": …, "meta": { request_id, correlation_id, actor } }`; error `{ "error": { code, message, details } }`. (`AgentApiController`.)

---

## 1. Full execution sequence

```
[Make trigger: a source AgentRun has completed / an entity is ready]
      │
      │ (1) create the run record
      ▼
POST /api/v1/ai/agent-runs                         → 201 (new) | 200 (existing by idempotency_key)
      │   owns: Numu (run store).  idempotent: idempotency_key.
      ▼
GET /api/v1/ai/agents/{agent_key}/knowledge/meta?environment=production
      │   owns: Numu (Knowledge Cache).  read-only.  404 if no active package.
      │   Make copies knowledge_version + checksum to stamp on the run (audit).
      ▼
POST /api/v1/ai/agent-runtime/run                  → 200 proposal | 4xx/502 fail-closed
      │   { agent_key, startup_id, agent_run_id, environment=production }
      │   owns: Numu Runtime.  RETURNS A PROPOSAL ONLY (executed:false).
      │   retry-safe: no state mutation beyond re-stamping traceability.
      ▼
[Make receives the validated proposal: analysis, proposed_action, confidence,
 reasoning_summary, proposed_note, proposed_handoff]
      │
      │ (map proposal → decision payload)
      ▼
POST /api/v1/ai/agent-runs/{run}/decision          → 200 { policy, decision, executed, … }
      │   { action_slug, confidence, reason, payload, environment=prod, risk_level? }
      │   owns: Numu decision pipeline (AgentDecisionService).
      │   NOT blindly retry-safe → check-before-retry (see §3).
      ▼
[Existing Numu pipeline — Numu owns everything from here]
   AgentDecisionService::decide()
     → recordAnalysisResult (freeze proposed_action + decision_payload)
     → pause/disable gate (AgentRuntimeService)
     → ActionPolicyService::resolve  (override → designed → default blocked)
     → branch:
         always_allow   → ExecutionRequestService::claim → ExpectedStateValidator
                          → AgentActionExecutor::executeForRun → recordExecutionResult
         needs_approval → ApprovalService::createFromRun (PENDING; human decides later)
         blocked        → recordExecutionResult(blocked)
      │
      ▼
[Optional, stage-dependent] POST /api/v1/ai/handoffs  → 201 (new) | 200 (existing)
      │   for intake_triage: only when the decision is move_to_review (AUTO-025 §3.1).
      │   owns: Numu (Handoff store).  idempotent: logical tuple (see §6).
      ▼
GET /api/v1/ai/handoffs/latest?startup_id=…          → verify the written record
```

**Ownership boundaries per step:** run store, knowledge, runtime, decision/policy/approval/execution, and handoff persistence are **all Numu**. Make's role is to *sequence the calls*, copy values between responses, and stop on errors.

---

## 2. Backend authority boundaries

**Make MAY:**
- Create AgentRuns (`POST /agent-runs`) with a caller-supplied `idempotency_key`.
- Read knowledge metadata (`GET …/knowledge/meta` / `…/knowledge`).
- Call the runtime (`POST /agent-runtime/run`) and receive a proposal.
- Submit the proposal to `POST /agent-runs/{run}/decision`.
- Read state (`GET /agent-runs/{run}`, `GET /handoffs/latest`, `GET /approvals`).
- Create/read handoffs (`POST /handoffs`, `GET /handoffs*`) — idempotent-safe.

**Make MUST NOT:**
- Call the Anthropic model directly (the Runtime owns model access).
- Assemble prompts, validate model output, or repair output (Runtime owns this).
- Decide policy, approve, execute native actions, change startup group/status, or send notifications (Numu owns this).
- Persist knowledge packages or write to the Knowledge Cache (Loader/admin only).
- Treat the proposal as final truth or claim execution occurred.
- Re-submit a decision blindly (may duplicate an approval — see §3).

**Numu EXCLUSIVELY owns:**
- Business authority: policy resolution, approval, expected-state validation, native action execution, notifications, audit.
- Idempotency + deduplication (run key, execution claim, handoff logical tuple, active-policy guard).
- Structured-output validation + fail-closed enforcement.
- Knowledge Cache (immutable versions, active pointer, rollback).

**Where authority lives:** business authority = `AgentDecisionService` + policy/approval/execution services (Numu). **Idempotency** = `AgentRunService` (run key), `ExecutionRequestService` (execution claim), `HandoffService` (logical tuple), `AgentActionPolicy` guard. **Validation** = `HandoffController::validateSourceRun` + `StructuredOutputValidator` (runtime) + each controller's `$request->validate` + policy resolution.

---

## 3. Retry semantics (per endpoint)

| Endpoint | Safe to blind-retry? | Idempotency mechanism | Duplicate handling | Timeout / compensation |
|---|---|---|---|---|
| `POST /agent-runs` | **Yes** | `AgentRunService::create` returns the existing run for a repeated `idempotency_key` (race-caught) | No duplicate run | Retry with the **same** `idempotency_key` |
| `GET …/knowledge/meta` | **Yes** | Pure read | n/a | Retry freely |
| `POST /agent-runtime/run` | **Yes** | No state mutation except re-stamping traceability on the run; produces a fresh proposal | No side effects | Retry freely; each call is a new proposal |
| `POST /agent-runs/{run}/decision` | **No — check first** | `recordAnalysisResult` freezes the decision once analysis is terminal; `always_allow` execution is claim-idempotent (`agent-run-{id}:{slug}`); **`needs_approval` is NOT idempotent — `createFromRun` always inserts a new approval** | Blocked branch harmless; always_allow safe; needs_approval may duplicate | On timeout, **`GET /agent-runs/{run}` first**: if `analysis_status` is terminal → decision applied (inspect `execution_status`/`approvalRequests`), do **not** re-POST; if still `queued`/`running` → safe to retry |
| `POST /handoffs` | **Yes** | `createIdempotent` on the logical tuple → returns existing (200); DB unique guard closes the race | Returns existing row, no new row/audit | Retry freely |
| `POST /handoffs/{id}/supersede` | **No — check first** | Guarded: re-superseding an already-superseded handoff → **409** | Second supersede rejected | On timeout, `GET /handoffs/latest` to see if the correction landed |

**Replay handling:** run-create replay → same run; runtime replay → new proposal (safe); handoff replay → existing row (200). **Compensation strategy:** the run is the coordination anchor — `GET /agent-runs/{run}` reveals `analysis_status`, `execution_status`, `approvalRequests`, `executionRequests`, so Make can determine outcome before any retry. **Never** re-POST a decision for a run with a terminal `analysis_status`.

---

## 4. Runtime contract

**Request** (`POST /api/v1/ai/agent-runtime/run`):
```json
{ "agent_key": "intake_triage", "startup_id": 123, "agent_run_id": 456, "environment": "production" }
```
Validation: `agent_key` required ≤80; `startup_id`, `agent_run_id` required integers; `environment` nullable ≤40 (knowledge namespace, default `production`).

**Response `200`** (`data`):
```json
{
  "agent_key": "intake_triage", "agent_code": "AUTO-025", "stage": "intake_triage",
  "environment": "production", "startup_id": 123, "agent_run_id": 456,
  "proposal": {
    "analysis": "…", "proposed_action": "move_to_review", "confidence": 0.82,
    "reasoning_summary": "…", "proposed_note": { … }|null, "proposed_handoff": { … }|null
  },
  "traceability": {
    "knowledge_package_id": 38, "knowledge_version": "v1", "knowledge_checksum": "…",
    "knowledge_registry_version": "1.0", "knowledge_loaded_at": "…", "source_stale": false,
    "prompt_contract_version": "v1", "runtime_provider": "anthropic", "runtime_model": "claude-sonnet-5",
    "runtime_latency_ms": 7767, "runtime_token_input": 1375, "runtime_token_output": 435,
    "runtime_request_id": "…", "runtime_calls": 1
  },
  "output_repaired": false, "executed": false
}
```

- **Proposal schema:** `analysis` (non-empty), `proposed_action` (allowed slug per `config('agent_runtime.allowed_actions.{agent}')` — for intake_triage: `move_to_review|reject|stop`), `confidence` ∈ [0,1], `reasoning_summary` (non-empty), `proposed_note` (obj|null), `proposed_handoff` (obj|null). Unknown fields stripped; execution-claim fields rejected.
- **`runtime_calls`:** 1 on first-valid output; 2 when a schema-repair occurred. Tokens/latency are **summed** across both calls.
- **Repair-call behavior:** invalid output → exactly one repair call → re-validate → fail closed if still invalid (`output_repaired=true` when the repair succeeded).
- **Fail-closed scenarios (before any state change):** run missing → **404**; run∉startup or agent_key/env mismatch or terminal run → **422 precondition_failed**; no active knowledge → **409 knowledge_unavailable**; invalid output after repair → **422 invalid_model_output**; provider timeout/error → **502 provider_error**. `executed` is always `false`.

---

## 5. Decision submission contract

**Request** (`POST /api/v1/ai/agent-runs/{run}/decision`) — validation from `AgentRunController::decision`:
```json
{ "action_slug": "move_to_review", "confidence": 0.82, "reason": "…", "payload": { … },
  "environment": "prod", "risk_level": "expected" }
```
Rules: `action_slug` required ≤120; `confidence` nullable numeric 0–1; `reason` nullable ≤2000; `payload` nullable array; `environment` nullable `in:dev,test,prod` (default `prod`); `risk_level` nullable `in:none,expected,high`.

**Mapping proposal → decision:**
| Proposal field | Decision field | Notes |
|---|---|---|
| `proposed_action` | `action_slug` | direct |
| `confidence` | `confidence` | stored in `decision_payload.confidence` |
| `reasoning_summary` | `reason` | stored in `decision_payload.reason` |
| (agent context) | `payload` | optional structured context → `decision_payload.payload` |
| — | `environment` | Make maps runtime env → decision env (§8) |
| — | `risk_level` | optional; if omitted, `decide()` defaults `reject→high`, else `expected` for the approval branch |

- **Confidence handling:** persisted on the run's `decision_payload`; also copied to the approval's `confidence_score` when `needs_approval`.
- **Reason handling:** persisted on `decision_payload.reason`; copied to the approval's `reason`.
- **Note handling:** `proposed_note` is **not** persisted by the decision endpoint. If a note is required, Make persists it via the existing note surface (`POST /api/v1/ai/notes`) or via `PATCH /agent-runs/{run}/analysis-result` (`note_payload`). The runtime/decision never auto-creates a Note.
- **Handoff handling:** `proposed_handoff` is **not** persisted by the decision endpoint. For intake_triage it is persisted via `POST /api/v1/ai/handoffs` **only when the decision is `move_to_review`** (AUTO-025 §3.1). Reject/Stop produce **no** handoff.

**Response `200`** (`data`): `{ policy, policy_source, decision, executed, run, … }` where `decision ∈ {executed, duplicate, state_mismatch, failed, pending_approval, blocked, agent_paused, agent_disabled}` and extras include `approval_id` / `execution_request_id` / `mismatches` per branch.

---

## 6. Handoff contract

- **Creation:** `POST /api/v1/ai/handoffs`. Semantic validation (`validateSourceRun`): the `source_run_id` must belong to the subject, be completed, and match `from_agent`/`stage` (= the run's `agent_key`). → **201** on create.
- **Duplicate:** the same logical tuple `(startup_id|investor_id, source_run_id, from_agent, to_agent, stage, schema_version)` → **200** with the existing row (no new row, no second audit, no supersede). Backed by app-level `createIdempotent` + the `handoffs_active_logical_unique` DB guard.
- **Supersede/correction:** `POST /api/v1/ai/handoffs/{id}/supersede` → new immutable corrected row; original gets `superseded_by`; the correction becomes the active head; `GET latest` returns it. Re-superseding an already-superseded row → **409**.
- **Replay behavior:** a replayed CREATE returns the current active head (the correction if one exists). A replayed supersede of a stale head → 409.
- **Audit behavior:** `agent.handoff_created` on create; `agent.handoff_superseded` on supersede; **no** audit on an idempotent duplicate hit.

---

## 7. Failure matrix

| Failure | Detected by | Signal | Recovery owner | Retry allowed? |
|---|---|---|---|---|
| Runtime unavailable | Make (HTTP) | `502 provider_error` from `/agent-runtime/run` | Make | **Yes** — backoff; runtime is side-effect-free |
| Invalid model output | Numu Runtime | `422 invalid_model_output` (after one repair) | Make | **Yes** (bounded) — no decision was submitted; re-run runtime |
| Policy blocked | Numu | decision `decision=blocked`, `executed=false` | Numu (terminal) | **No** — deterministic terminal outcome |
| Approval required | Numu | decision `decision=pending_approval` + `approval_id` | Human (via `/approvals/{id}/decide`) | **No auto-retry** of the decision |
| Execution failed | Numu | decision `executed=false`, `execution_status=failed` | Numu / operator | **No blind retry** — `always_allow` is claim-idempotent; a genuine failure needs investigation, not re-POST |
| Handoff duplicate | Numu | `POST /handoffs` → **200** (existing) | none needed | **Yes** — idempotent |
| Startup changed (stale expected-state) | Numu | decision `decision=state_mismatch` + `mismatches` | Make | **Yes** — re-read state, re-run with fresh expected values |
| Stale knowledge | Numu Runtime | `409 knowledge_unavailable` (no active pkg) or `source_stale` flag | Make/admin | **After refresh** — refresh/activate knowledge, then retry |
| Timeout | Make | no response | Make | **After state check** — `GET /agent-runs/{run}` then decide whether to retry |
| Retry (generic) | Make | — | Make | Governed by the per-endpoint rules in §3 |

---

## 8. Environment handling (current split)

Two vocabularies exist in the repository (see `ENVIRONMENT_RECONCILIATION_ANALYSIS.md`):

```
knowledge / runtime :  dev | staging | production   (default: production)
decision / policy   :  dev | test    | prod          (default: prod)
```

- The **runtime** call (`/agent-runtime/run`) and knowledge reads use the **knowledge** namespace (`production`).
- The **decision** call (`/agent-runs/{run}/decision`) uses the **policy** namespace (`prod`) — validated `in:dev,test,prod`.

**Make MUST map between them per call until reconciliation:**

| Logical env | Runtime / knowledge value | Decision value |
|---|---|---|
| Production | `production` | `prod` |
| Development | `dev` | `dev` |
| Staging | `staging` | `test` *(no `staging` policy namespace exists)* |

**Fail-safe:** a wrong value cannot cause a cross-environment action — the decision endpoint rejects `production` (`422`, `in:dev,test,prod`), and the runtime returns `409` for an unknown knowledge env. Never silently mis-scoped. Reconciliation to a single set is a documented **post-integration** hardening item (non-blocking).

---

## 9. Sequence diagrams

**9.1 Successful flow (always_allow → executed)**
```
Make → Numu : POST /agent-runs {idempotency_key}          → 201 {run.id}
Make → Numu : GET  …/knowledge/meta?environment=production → 200 {version,checksum}
Make → Numu : POST /agent-runtime/run {run.id,production}  → 200 {proposal, executed:false}
Make → Numu : POST /agent-runs/{run}/decision {action_slug,confidence,reason,prod}
Numu        : resolve policy = always_allow → claim → validate-state → execute
Numu → Make : 200 {decision:"executed", executed:true, execution_request_id}
Make → Numu : POST /handoffs (if move_to_review)           → 201 {handoff.id}
Make → Numu : GET  /handoffs/latest                         → 200 {handoff.id}
```

**9.2 Approval-required flow**
```
… (run, knowledge, runtime as above) …
Make → Numu : POST /agent-runs/{run}/decision {action_slug=reject,prod}
Numu        : resolve policy = needs_approval → create PENDING approval (no execution)
Numu → Make : 200 {decision:"pending_approval", executed:false, approval_id}
Human → Numu: POST /approvals/{approval_id}/decide {approve|reject}   (out of band)
Numu        : on approve → Make re-validates state + executes via the pipeline
Make        : MUST NOT re-POST the decision (would duplicate the approval)
```

**9.3 Blocked flow**
```
Make → Numu : POST /agent-runs/{run}/decision {action_slug,prod}
Numu        : resolve policy = blocked (or agent paused/disabled) → record blocked
Numu → Make : 200 {decision:"blocked", executed:false}
Make        : terminal — no retry
```

**9.4 Duplicate-handoff flow (idempotent)**
```
Make → Numu : POST /handoffs {logical tuple}   → 201 {handoff.id=N}
… (retry / at-least-once redelivery) …
Make → Numu : POST /handoffs {same tuple}      → 200 {handoff.id=N}   (same row)
Numu        : no new row, no second audit, no supersede
```

**9.5 Retry flow (decision timeout → check-before-retry)**
```
Make → Numu : POST /agent-runs/{run}/decision …   → (timeout, no response)
Make → Numu : GET  /agent-runs/{run}
  ├─ analysis_status = completed  → decision applied; inspect execution/approval; DO NOT re-POST
  └─ analysis_status = queued     → decision not applied; safe to re-POST the decision
```

---

## 10. Final architecture statement (confirmed against the repository)

- ✅ **Make remains orchestration-only** — it sequences calls, copies values, and stops on error; it holds no business logic, no model access, no prompt/validation ownership, no execution authority.
- ✅ **Claude remains proposal-only** — invoked solely inside the Runtime; its output is validated and used only as a recommendation.
- ✅ **Runtime remains propose-only** — `RuntimeRunner` returns a proposal, stamps traceability, and performs a single AgentRun metadata write; it never decides/approves/executes (`executed:false`).
- ✅ **Numu remains the sole business authority** — policy resolution, approval, expected-state validation, native execution, notifications, audit, and all deduplication live in Numu services.
- ✅ **No parallel decision pipeline** — the only entry to business action is `POST /agent-runs/{run}/decision` → `AgentDecisionService`; the runtime feeds it, never bypasses it.
- ✅ **No execution logic outside Numu** — `AgentActionExecutor`/`ExecutionRequestService`/`ExpectedStateValidator` are Numu-internal; Make cannot execute actions.

**Status:** specification only. No Make scenario or backend wiring is implemented. Awaiting approval before any Make integration; all future work stays on `feat/agent-knowledge-cache` with the existing implementation — no new architecture, service, repository, package, module, or execution path.
