# AI Agent Runtime Contracts Draft

| Field | Value |
|---|---|
| Document Type | Runtime Contracts Draft (Reconciliation) |
| Document Version | 0.2 (Draft — reconciled) |
| Status | Draft for Review — design/contract only |
| Repository | Azure DevOps `numuinvestment / numu / numu-old` |
| Authoritative base | `main` |
| Foundation Path | `/ai-agents` |
| Unification | Phase 5 backend implementation **and** the `/ai-agents` foundation module are already unified on `numu-old/main`. |
| Last Updated | 2026-07-07 |
| Authoring Note | Reconciled against the current `numu-old/main` clone: the `/ai-agents` foundation module (README, technical plan, placeholder folders) and the Phase 5 AI backend (Laravel `app` / `routes` / `database`). Verified from repo evidence (migrations + routes), the workspace shared-knowledge design docs, and Fahad's confirmed decisions (2026-07-06/07). |

> **Reconciliation note (2026-07-07):** This revision reconciles the contract draft against the code **already present** on `numu-old/main`. The earlier feature branches `feature/ai-agent-runtime-foundation` and `feat/agent-knowledge-cache` were merged and are no longer active; `main` is the single authoritative base (HEAD: `Merged PR 1: Add AI Agent Runtime foundation module`, plus the Phase 5 AI backend). This is still a **design/contract draft for review** — it is **not** production-ready, **not** deployed, **not** wired to Make production, **not** runtime-activated, and AUTO-026 is **not** production-registered.

---

## 0. Implementation Status Legend

Labels are used **only where helpful** to show how each contract area stands against `main`:

- **Implemented Backend Contract** — present in the Phase 5 backend on `main`; must be **verified/aligned** during controlled validation (not built from scratch).
- **Draft / Needs Final Contract** — shape agreed here; final contract still to be confirmed with the developer.
- **Operational Open Item** — an operational/configuration item still to be resolved.
- **Product Decision Resolved** — decided by Fahad; not reopened here.
- **Controlled Validation Item** — to be exercised/confirmed during controlled production validation.

Repo evidence used for the *Implemented Backend Contract* labels (migrations under `database/migrations`, routes under `routes/api.php`): `create_agents_table`, `create_agent_runtime_config_table`, `create_agent_runs_table` (+ `add_knowledge_columns_to_agent_runs`, `add_context_to_agent_runs`, `add_runtime_metrics_to_agent_runs`), `create_handoffs_table` (+ `add_investor_sequence_unique_to_handoffs`, `add_handoff_logical_idempotency_guard`), `create_agent_action_policies_table` (+ `add_active_policy_guard_to_agent_action_policies`), `create_execution_requests_table`, `create_agent_approval_requests_table` (+ `add_agent_decision_fields_to_approval_requests`, `add_active_approval_guard_to_agent_approval_requests`), `create_agent_knowledge_cache_table` (+ `add_environment_to_agent_knowledge_cache`, `augment_knowledge_traceability`), `add_agent_run_id_to_logs_and_notes`; and AI routes `POST/GET /api/v1/ai/agent-runs`, `GET /api/v1/ai/agent-runs/{run}`, `PATCH …/analysis-result`, `POST …/decision`, `POST …/execution-results`, `POST …/cancel`, `POST /api/v1/ai/agent-runtime/run`, plus `ai.cb` circuit-breaker and audit/rollback via `AiActivityLogger`. Exact column names/types must still be confirmed against the models/migrations during validation.

---

## 1. Purpose

This document defines the **draft runtime contracts** for the AI Agent Runtime that lives inside the existing Numu Dashboard repository at `/ai-agents`, **reconciled against the Phase 5 AI backend already present on `numu-old/main`**.

Its purpose is to give a single, reviewable definition of the runtime boundaries, data shapes, and orchestration rules — clear enough for developer alignment, Make.com orchestration planning, agent registration, controlled production validation, product review, and approval/execution-boundary review.

Explicit status:

- **Design/contract status only** — contracts (shapes, boundaries, rules), reconciled with existing backend code.
- **Not production-ready**, **not deployed**, **not connected to Make production**, **not runtime-activated**, **AUTO-026 not production-registered**.

Approving this draft agrees the reconciled contracts as the basis for controlled validation and Make wiring. It does **not** authorize production activation or any live write.

---

## 2. Scope

### 2.1 Covered

Core runtime boundaries; agent key naming (canonical AUTO-026 key); reconciled contracts for Registry, Knowledge Loader, AgentRun, Structured Model Output, Handoff, Notes, Approval, Execution Layer; controlled production validation plan; idempotency/duplicate-prevention; rollback/pause; Make.com orchestration contract; open items; implementation sequence; non-goals; review checklist.

### 2.2 Not covered

No production deployment or migration authoring; no Make scenario/module/webhook/Data Store construction; no secret creation; no live API call, AI-model call, startup status change, or notification send; no production activation of AUTO-025 or AUTO-026.

### 2.3 Relation to `/ai-agents` and to `main`

`/ai-agents` is the runtime foundation area at the repo root (confirmed on `main`). Its folders (`runtime`, `registry`, `knowledge-loader`, `agent-runs`, `handoffs`, `approvals`, `notes`, `execution-layer`, `shared`, `tests`) are documentation placeholders (`.gitkeep`); the **executable Phase 5 AI backend** lives in the Laravel application (`app`, `routes`, `database`, `mcp-server`). This draft maps each contract concept onto the corresponding existing backend contract. The confirmed canonical foundation path is **`/ai-agents`** (any older `/backend/ai-agents` reference is superseded — corrected in the technical plan).

### 2.4 Relation to Numu Dashboard

The runtime is **inside** the existing dashboard repo — no new project, repo, or separate service. It consumes and proposes changes to dashboard/backend entities only through approved backend/API contracts and the Execution Layer, never by direct mutation.

### 2.5 Relation to Make.com

Make.com is **orchestration only** (see §3 and §15) — it triggers and sequences calls, but is not the source of truth, does not call the AI model, does not decide policy, and does not execute business mutations except through an approved backend endpoint.

---

## 3. Core Runtime Boundaries

Each capability is owned by exactly one layer; authority must not leak across boundaries.

| Layer | Owns | Must NOT do |
|---|---|---|
| **Agent Runtime** | Loads the active knowledge package; assembles the prompt; calls the AI model; validates structured output; returns a **proposal** only. | Execute business actions; mutate startup status; persist Handoffs/Notes; make approval decisions. |
| **Registry** | Source of truth for which runtime-loadable packages exist per `agent_key + environment`. | Store file contents/secrets; grant production activation by itself. |
| **Knowledge Loader** | Retrieve → validate → checksum → assemble → version → (atomic) activate the runtime package; fail closed. | Analyze startups; make business decisions; expose content beyond the requested agent+environment. |
| **AgentRun** | Immutable record of one run: inputs, knowledge provenance, proposal, policy result, approval/execution status, telemetry. | Be edited retroactively; act as an execution channel. |
| **Handoff** | Immutable, versioned, machine-readable context package between agents. | Be created by unmanaged model output; be treated as current-state proof. |
| **Approval** | Human decision record when policy says `needs_approval`. | Be bypassed by Make; be inferred by the model. |
| **Notes** | Operational, human-readable records. | Replace Handoffs; claim an action executed unless it actually was. |
| **Execution Layer** | The only layer that performs approved mutations, with precondition (expected-state) checks and full audit. | Run inside the Runtime; be bypassed by Make. |
| **Make.com** | Orchestration/trigger sequencing and alerting. | Call the model; assemble core prompts; validate structured output; decide policy; execute business actions; mutate startup status directly; bypass approval/policy. |
| **Human approval** | Authorize `needs_approval` actions. | Be simulated by any automated layer. |
| **Backend/API** | Authoritative entity state, persistence, audit history, official endpoints. | — |

**Make.com is orchestration only.** Make must not call the AI model, assemble core prompts, validate structured model output, decide policies, execute business actions, mutate startup status directly (except through an explicitly approved Execution-Layer endpoint), or bypass approval/policy checks.

---

## 4. Agent Key Naming Contract

An agent key is the stable technical identifier used by the Registry, Runtime, AgentRun, Handoff routing, and Make mapping.

Standard: **lowercase** ASCII; **stable** (never reused/renamed casually); **no spaces**; **no ambiguity** (one canonical spelling; no aliases for new work); the **technical key is separate from the human-readable name** (`agent_name`).

Current keys: `intake_triage` (AUTO-025, registered in `AGENT_KNOWLEDGE_REGISTRY.json`, enabled); `prescreen` (AUTO-026).

**AUTO-026 key — Product Decision Resolved:**

> The approved technical `agent_key` for AUTO-026 is **`prescreen`**. The historical **`pre_screen`** key is superseded and must **not** be used for new runtime, registry, routing, handoff, or Make mapping work. Any legacy records carrying `pre_screen` should be treated as **legacy compatibility data**, not as an active naming decision.

This is settled and is **not** an open decision. Legacy `pre_screen` data compatibility (e.g. historical Handoff records) is a data-migration/compatibility note for the developer, handled during controlled validation — not a naming choice to reopen.

---

## 5. Registry Contract

**Purpose:** the Registry is the **source of truth for runtime-loadable agent packages** — it lists exactly which files compose each agent's runtime package and how to assemble them (references and rules only; never file contents, tokens, or signed URLs). Aligns with `AGENT_KNOWLEDGE_REGISTRY.json` (registry_version `1.0`).

Draft per-agent structure — **Draft / Needs Final Contract**, reconciled with the backend `agents` / `agent_runtime_config` tables (**Implemented Backend Contract**):

| Field | Required | Notes |
|---|---|---|
| `agent_key` | yes | Canonical lowercase key (`intake_triage`, `prescreen`). |
| `agent_name` | yes | Human-readable display name. |
| `version` / `package_version` | yes | Runtime package version (e.g. `v01`); missing → fail closed. |
| `status` / `enabled` | yes | Whether the agent may load/run. |
| `environment` | yes | `development` \| `staging` \| `production`. |
| `runtime_package_path` | yes | Training/prompts folder (`04 Prompts`). |
| `knowledge_package_path` | yes | Shared + runtime file references. |
| `required_files` | yes | Must all be present/non-empty; else fail closed. |
| `optional_files` | no | May be absent. |
| `output_contract` | yes | References to Note authoring + Handoff authoring contracts. |
| `allowed_actions` | yes | Agent-specific `proposed_action` values (§8). |
| `approval_policy_reference` | yes | Pointer to the policy (backend `agent_action_policies`). |
| `handoff_schema_reference` | yes | Pointer to the agent's `HANDOFF_SCHEMA.md`. |
| `note_schema_reference` | yes | Pointer to the agent's `NOTE_GUIDE.md`. |
| `owner` | yes | Accountable owner. |
| `last_updated` | yes | Date of last registry change. |
| `checksum/hash placeholder` | optional | Bundle/identity hash placeholder; content hashing is owned by Loader/Cache. |

Supporting rules preserved: `load_order`, `excluded_runtime_files` (audit-only), `cache_policy`, `loader_rules` (`max_package_bytes`, `no_wildcard_loading`, `explicit_files_only`, `store_contents = false`, `store_secrets = false`).

Clarifications: only explicitly listed files may enter a prompt (no folder/wildcard loading); **Registry approval ≠ production activation**; **Training Package approval is separate from Runtime Registry approval**.

---

## 6. Knowledge Loader Contract

Aligns with `AGENT_KNOWLEDGE_LOADER_CONTRACT.md` and `AGENT_KNOWLEDGE_CACHE_SCHEMA.md`. The **Knowledge Cache is an Implemented Backend Contract** (`agent_knowledge_cache` table with `environment` and knowledge-traceability columns) and must be **verified/aligned** during controlled validation.

- **Source of content:** SharePoint/OneDrive is the source of truth for file **content**; the Registry is the allow-list. The Loader resolves each registered path under the approved `Startup Agents` root, validates, checksums, assembles, versions, and (except `validate_only`) atomically activates. The Runtime reads only the active, valid package for its `agent_key + environment` — never SharePoint during a run.
- **Ordering:** strict registry `load_order`; mismatch → `LOAD_ORDER_MISMATCH` (fail closed).
- **Required vs optional:** all `required_files` present/non-empty; `optional_files` may be absent.
- **Missing file:** `REQUIRED_FILE_MISSING` / `REQUIRED_FILE_EMPTY` → fail candidate, keep active package; never activate a partial package.
- **Checksum/version:** per-file `content_checksum` + package `bundle_checksum` (sha256) + `package_version` + `registry_version`. Identical force-refresh reproduces the same checksum → `unchanged` (no duplicate active package).
- **`source_stale`:** source unavailable **and** active valid package exists → continue on last-known-good with `source_stale = true`; source unavailable **and** no active package → hard fail (`SOURCE_UNAVAILABLE_NO_ACTIVE`).
- **Cache boundary:** Numu-owned, versioned, immutable, encrypted-at-rest packages — not Make Data Store, not Handoff storage. Lifecycle `candidate → validating → valid → active → superseded`, `rejected` terminal; one active pointer per `agent_key + environment`.
- **Fail closed on:** agent not registered/disabled; invalid registry JSON; missing `package_version`; missing/empty required file; duplicate path/load-order; path escape; manifest referencing unregistered file; unsupported encoding; size-limit exceeded; per-file/bundle checksum failure; environment mismatch; activation transaction failure; source unavailable with no active package. On failure `active_package_unchanged = true` + refresh-run audit row.
- **Warn (not fail):** serving last-known-good on transient source failure (`source_stale`) — never presented as fresh validation.

Clarifications: project-level `/knowledge` folders are **context only** unless explicitly registered; the Runtime loads **only approved runtime package content**.

---

## 7. AgentRun Schema Draft

**Implemented Backend Contract** — `agent_runs` exists on `main` with knowledge, context, and runtime-metrics columns; verify exact column names during validation. Draft fields:

| Field | Requiredness | Notes |
|---|---|---|
| `agent_run_id` | required | Backend-issued id (integer). |
| `agent_key` | required | Canonical key. |
| `startup_id` / `subject_id` | required | V1 subject `startup`; `subject_id` generalizes (investor subject columns exist). |
| `subject_type` | required | `startup` in V1. |
| `trigger_source` | required | e.g. `make` \| `manual` \| `scheduler` \| `backend`. |
| `trigger_id` | optional | Upstream trigger reference. |
| `idempotency_key` | required | Prevents duplicate runs (see §14). |
| `environment` | required | `development` \| `staging` \| `production`. |
| `status` | required | Overall run lifecycle. |
| `analysis_status` | required | Model/analysis phase (see `PATCH …/analysis-result`). |
| `proposed_action` | optional | Agent-specific allowed value once analysis completes. |
| `confidence` | optional | Model confidence. |
| `policy_result` | required once evaluated | e.g. `needs_approval` \| `always_allow` \| `blocked`. |
| `approval_status` | conditional | Required when `policy_result = needs_approval`. |
| `execution_status` | conditional | Set by the Execution Layer (see `POST …/execution-results`). |
| `error_code` / `error_message` | optional | On failure. |
| `correlation_id` | required | Cross-layer trace id. |
| `created_by` / `created_at` / `updated_at` | required | Actor + timestamps. |

Traceability / telemetry (knowledge columns + runtime metrics on `agent_runs`): `knowledge_package_id`, `knowledge_version`/`knowledge_package_version`, `knowledge_checksum`/`knowledge_bundle_checksum`, `knowledge_registry_version`, `knowledge_loaded_at`, `source_stale`, `prompt_contract_version`, `runtime_provider`, `runtime_model`, `runtime_latency_ms`, `runtime_token_input`, `runtime_token_output`, `runtime_request_id`, `runtime_calls`.

Reconciliation with the foundation technical plan §3.5 ("AgentRun Logging"): the plan's shorter suggested set maps onto the above (`model` → `runtime_model`; `prompt_package_version` → `prompt_contract_version`/`knowledge_version`; `knowledge_hash` → `knowledge_checksum`; `recommendation`/`output_json` → `proposed_action` + structured output; `completed_at` → completion timestamp). Final persisted column names are a **Controlled Validation Item** (confirm against the migrations/models).

---

## 8. Structured Model Output Contract

The Runtime requires a single structured object (**Implemented Backend Contract** — structured-output validation and the runtime proposal path exist via `POST /api/v1/ai/agent-runtime/run` and `PATCH …/analysis-result`; verify during validation).

| Field | Notes |
|---|---|
| `analysis` | Structured analysis vs the agent's stage rules. |
| `proposed_action` | One value from the agent's **allowed_actions**. |
| `confidence` | Confidence for the proposal. |
| `reasoning_summary` | Short rationale (no chain-of-thought dumping). |
| `proposed_note` | Optional proposed Note content (§10). |
| `proposed_handoff` | Optional proposed Handoff draft (§9). |
| `validation_errors` | Empty on clean output. |
| `requires_human_review` | Boolean routing signal. |

Rules:

- **Allowed actions are agent-specific.** Known startup-stage actions: `move_to_review`, `request_more_info`, `request_meeting` (approved design; live startup Action-catalog availability is an **Operational Open Item**), `reject`. Actions outside the allow-list are rejected.
- **Invalid output fails closed** — no business decision, Handoff, or guessed result.
- **One repair attempt** — the Runtime may re-prompt once for a schema-valid object; a second failure fails closed. (**Implemented Backend Contract** — verify the one-repair behavior during validation.)
- **Execution never happens inside the Runtime** — proposal only; no mutation, no persistence, no approval.
- Output must **never claim an action was executed** — agent-authored content is intent only; execution facts are added later by the Execution Layer (§9 `decision_summary.execution`).

---

## 9. Handoff Schema Draft

Aligns with `HANDOFF_JSON_SCHEMA_V1.json` (`schema_version: "v1"`) and `HANDOFF_PAYLOAD_CONTRACT_V1.md`. **Implemented Backend Contract** — `handoffs` table exists with investor-subject support, a unique sequence, and a **logical idempotency guard** (`add_handoff_logical_idempotency_guard`, 2026-07-05); verify/align during validation.

| Field | Notes |
|---|---|
| `handoff_id` | Server-managed id. |
| `source_run_id` | Parent AgentRun id (required). |
| `startup_id` / `subject_id` | Exactly one subject (`startup_id` in V1; `investor_id` supported). |
| `subject_type` | `startup` in V1. |
| `from_agent` / `to_agent` | Agent keys (use canonical `prescreen` for new work). |
| `stage` | Stage label. |
| `schema_version` | `v1`. |
| `facts` | `{key, value, source_type, source_ref}`. |
| `findings` | `{summary, evidence_refs[]}` — a finding requires evidence. |
| `risks` | `{summary, severity(low\|medium\|high), evidence_refs[]}`. |
| `open_questions` | `{question, reason, blocking}`. |
| `documents_used` | `{document_type, document_ref, version}` — references only. |
| `decision_summary` | `{agent{proposed_action, expected_resulting_group, expected_resulting_status, rationale}, execution?{policy_result, approval_outcome, actual_executed_action, actual_group_after, actual_status_after, source_run_id, execution_source_ref}}`. |
| `retention_class` | Server-managed. |
| idempotency/logical key | entity + `from_agent` + `to_agent` + `stage` + `source_run_id` (+ `sequence_number`) — backend guard present. |
| `supersedes_handoff_id` / `superseded_by` | Supersede lineage (server-managed `superseded_by`). |
| `created_at` | Server-managed. |

Server-managed fields (`sequence_number`, `superseded_by`, `created_at`) and `retention_class` are not part of the submitted payload.

Clarifications:

- **Authoring vs persistence are separate.** The agent/Runtime may **propose** a Handoff (`decision_summary.agent` = intended outcomes only). Persistence is owned by backend/Execution Layer; `decision_summary.execution` is authored by the Execution Layer from verified Numu evidence only.
- **Semantic checks beyond JSON Schema** (proposed action/group/status match; valid key/transition; `source_run_id` ownership + parent-run success; current-state verification; duplicate detection) are enforced by a Runtime/Loader application pass.
- **Duplicate prevention** is now an **Implemented Backend Contract** (logical idempotency guard) — verify it behaves as expected during controlled validation; any Make-side creation path must still send a stable logical key and check before retrying (§14).
- **Supersede/correction** — supersede lineage fields exist (**Implemented Backend Contract**); confirm supersede/correction behavior/endpoint during validation.

---

## 10. Notes Persistence Contract

**Implemented Backend Contract — verified (Saad, 2026-07-06); no longer unresolved.** Notes link to a run via `agent_run_id` (`add_agent_run_id_to_logs_and_notes`).

Endpoints:

- Create: `POST /api/v1/ai/notes`
- Edit: `PATCH /api/v1/ai/notes/{note}`
- Investor alias (create): `POST /api/v1/investors/{investor}/notes`

Create schema:

- `entity_type`: `investor` \| `startup`
- `entity_id`
- `body` (1..5000)
- `rationale?`
- `agent_run_id?`

Create response: `note_id`, `entity`, `admin_id`, `created_at`.

Edit schema: `body`, `rationale?`. Edit response: `changed`, `note_id`, `updated_at`.

Rules:

- **Create is append-only and is NOT idempotency-keyed.** `agent_run_id` links notes to a run for traceability and dedup responsibility (the caller/orchestrator must avoid creating duplicate notes — see §14/§15).
- **Edit is content-idempotent.**
- **Approval is not required for notes.**
- An **audit row is written atomically** with create/edit.
- **Delete is intentionally not offered.**
- The Runtime may **propose** note content; persistence is performed by the backend via these endpoints (the Runtime does not itself persist).
- Notes must **not** claim an action was executed unless it actually was.

---

## 11. Approval Request Contract

**Implemented Backend Contract** — `agent_approval_requests` exists with agent-decision fields and an **active approval guard** (`add_active_approval_guard_to_agent_approval_requests`, 2026-07-05); verify during validation.

| Field | Notes |
|---|---|
| `approval_request_id` | Backend-issued id. |
| `agent_run_id` | Originating run. |
| `startup_id` / `subject_id` | Subject. |
| `proposed_action` | Action awaiting approval. |
| `policy_result` | Why approval is required (`needs_approval`). |
| `risk_level` | Risk classification. |
| `approval_status` | `pending` \| `approved` \| `rejected` \| `expired`. |
| `approver_id` | Human approver. |
| `approval_reason` | Rationale. |
| `approved_at` / `rejected_at` / `expires_at` | Timestamps/expiry. |
| idempotency behavior | One-shot / idempotent (active approval guard). |
| audit trail | Full audit of the decision. |

Clarifications:

- **Approval is required whenever policy returns `needs_approval`.**
- **AUTO-026 approval-required actions** (foundation technical plan §3.8): `request_more_info`, `request_meeting`, `reject` — all require human approval before execution.
- **`move_to_review` for first controlled validation = `needs_approval` (Product Decision Resolved; §7 / §3.7 policy).** It must create an approval request and must not auto-execute.
- **Make must not bypass approval** — Make may only create/poll approval requests.
- **Decisions are one-shot / idempotent.** Human approval is separate from the model proposal.

---

## 12. Execution Layer Contract

**Implemented Backend Contract** — `execution_requests` exists; execution is recorded via `POST /api/v1/ai/agent-runs/{run}/execution-results`, and the decision path is `POST /api/v1/ai/agent-runs/{run}/decision`; audit/rollback via `AiActivityLogger`. Verify during validation.

| Field | Notes |
|---|---|
| `execution_request_id` | Backend-issued id. |
| `agent_run_id` | Originating run. |
| `approved_action` | Authorized action. |
| `expected_current_state` | Required state before mutation (Group/Status). |
| `target_state` | Intended state after. |
| precondition checks | Verified before any write. |
| `execution_status` | `pending` \| `executed` \| `failed` \| `skipped`. |
| `executed_by` / `executed_at` | Actor + timestamp. |
| `rollback_note` | How to revert / what was preserved. |
| `audit_log_id` | Link to the immutable audit record (`ai_activity_logs`). |
| idempotency key | Prevents duplicate execution (§14). |

Clarifications:

- **The Execution Layer owns mutations** — the only layer that changes business state.
- **The Runtime does not execute** — proposals only.
- **Make does not directly execute** business actions unless explicitly routed through an approved backend endpoint owned by the Execution Layer.
- **Expected-state validation before mutation** — mismatch → stop and record a contradiction rather than force the write.
- **All mutations are auditable.**

---

## 13. Controlled Production Validation Plan

**Product Decision Resolved:** Fahad selected **controlled production validation using clearly marked test startup records**. A dedicated staging environment is **not** a blocking open question; staging is mentioned only as an **optional/future** path if it becomes available later.

**First controlled production validation policy:** `move_to_review = needs_approval`. Any `move_to_review` decision during controlled production validation or first limited activation must create an approval request and must **not** auto-execute. `always_allow` is **not** approved for first controlled validation and may only be used later if Fahad explicitly approves it.

Controlled production validation sequence (based on Saad's runbook) — **Controlled Validation Items**:

1. Create or identify a marked test startup record.
2. Prefer the real public signup flow if appropriate, or a controlled DB insert if handled by developer/admin.
3. Mark the record clearly, e.g.: name starts with `ZZ_TEST_AI_`; `source = ai_controlled_test` if available.
4. Record: `TEST_STARTUP_ID`, `BASELINE_GROUP`, `BASELINE_STATUS`.
5. Force or verify `needs_approval` for `move_to_review`.
6. Verify policy resolution before running.
7. Create the AgentRun.
8. Run the runtime proposal if needed.
9. Submit the decision.
10. Expect an approval created and execution **not** triggered.
11. Verify: exactly one approval; execution not started; startup group/status unchanged; audit logs present.
12. Cleanup: reject the test approval; cancel the run; revoke the scoped override if used; soft-delete/archive the test startup only; **never hard-delete** (because `agent_runs` / `approvals` / `handoffs` may reference it).

**Scoped-override correction:** if the global policy is already `needs_approval`, a scoped override is **optional safety only**. After revoking a scoped override, verify that policy resolution falls back to the global/default policy. For the first controlled validation, the expected fallback may still be `needs_approval`.

General guardrails: **no broad production activation before a clean controlled end-to-end run**; **test records must be clearly marked**; **production writes require explicit approval**.

---

## 14. Idempotency and Duplicate Prevention Contract

| Operation | Idempotency owner | Behavior |
|---|---|---|
| AgentRun creation | Backend (via `idempotency_key`) | Same logical trigger → same run. Make must send a stable key. **Implemented Backend Contract.** |
| Runtime execution / retry | Runtime | A retried model call must not create a second AgentRun; reuse the existing run. |
| Decision submission | Backend (`…/decision`) | Re-submitting the same decision is a no-op. **Implemented Backend Contract.** |
| Approval | Backend (active approval guard) | One-shot/idempotent; re-submission does not duplicate or flip a settled decision. **Implemented Backend Contract.** |
| Execution | Execution Layer (idempotency key + expected-state) | Same approved action executes once. **Implemented Backend Contract.** |
| Handoff creation | Backend (logical idempotency guard) | Duplicate logical key suppressed. **Implemented Backend Contract — verify during validation.** |
| Note creation | Caller/orchestrator (append-only, NOT keyed) | Backend does **not** dedup notes; `agent_run_id` carries dedup responsibility. Make/caller must not blindly recreate notes. |

Rules for Make:

- **May retry blindly:** idempotent/read-only or key-protected calls where the backend guarantees dedup (AgentRun creation with a stable `idempotency_key`, GETs, decision submission).
- **Must check before retrying:** any write where dedup depends on the caller — above all **Note creation** (append-only, not keyed) and any execution/mutation call; and confirm the Handoff guard result before re-posting.
- **Never retry blindly:** business mutations without idempotency key/expected-state, approval decisions, and note creation.

---

## 15. Make.com Orchestration Contract

> Reference basis: the Make Skills (`make-scenario-building`, `make-module-configuring`, `make-mcp-reference`, `make-api-shell-connection-workflow`) are the operating reference for any future Make design. **No Make scenario, module, webhook, Data Store, or connection is built or edited by this document.**

Make's future role is orchestration only.

**Make may:** trigger the flow; create the AgentRun (stable `idempotency_key` + `correlation_id`); call the knowledge **metadata** endpoint; call the runtime endpoint for a proposal; submit the decision; poll approval status; request Handoff creation if the backend contract allows (rely on the logical idempotency guard, send the logical key); read back final state; alert on failure.

**Make must not:** execute decisions directly; bypass backend policy; bypass approval; call the AI model directly; modify statuses outside the approved Execution Layer; store secrets unnecessarily; create duplicate state without idempotency keys. Because **note creation is append-only and not idempotency-keyed**, Make must **not** blindly create duplicate notes (dedup via `agent_run_id`/pre-check). Handoff/decision/approval retry behavior must follow the backend contract.

**Recommended Make sequence:**

```
trigger
  → create run (idempotency_key, correlation_id)
  → knowledge metadata check
  → runtime proposal
  → decision submission
  → approval handling (only if policy = needs_approval)
  → execution result check
  → handoff / note handling (respect backend idempotency; no blind note duplication)
  → audit / alert
```

MCP/auth/visibility, connection, timeout, and permission concerns for any future Make↔Numu bridge should be validated against `make-mcp-reference` and `make-api-shell-connection-workflow` before that section is built.

---

## 16. Rollback and Pause Contract

**Product Decision Resolved:** rollback/pause ownership is **manual control by the platform admin** — this is not reopened. Rollback surfaces are **admin dashboard actions**, not public `/api/v1/ai/*` APIs, and are protected by permission **`ai.rollback`**.

Pause / resume (admin-protected dashboard/backend actions, not necessarily public `/api/v1/ai/*` endpoints):

- Pause all agents: `POST /agents/pause-all`
- Resume all: `POST /agents/enable-all`
- Pause one agent: `POST /agents/{agentKey}/pause`

Admin rollback surfaces:

- `admin/ai/activity/{id}/rollback`
- `admin/ai/sessions/{id}/rollback`
- `admin/ai/tasks/{id}/rollback`
- `admin/ai/rollback/range`

Minimal incident handling: **pause all → investigate → rollback → restore state → log incident.** (Backend audit/rollback is backed by `AiActivityLogger` / `ai_activity_logs`.) Ownership, alerting, and security specifics are intentionally **not** expanded here.

---

## 17. Open Items

Only direct, unresolved items that matter for running the project (resolved decisions are **not** reopened — see §18):

1. Final review and sign-off of this reconciled contract by Fahad and Saad.
2. Approval to execute the controlled production validation run.
3. Make wiring plan/build **after** contract approval (per §15; no scenario built yet).
4. Any remaining endpoint/path or field-name mismatch actually found in the repo during validation (confirm exact `agent_runs`/`handoffs`/etc. column names against migrations/models).
5. `request_meeting` live Action-catalog availability (Operational Open Item).
6. `/ai-agents` vs `/backend/ai-agents` path correction — applied to the technical plan in this patch; confirm no other stale reference remains.

Explicitly **not** reopened here: canonical key (`prescreen` resolved), testing path (controlled production validation resolved), `move_to_review = needs_approval` (resolved), note persistence endpoint (verified), rollback/pause ownership (platform admin, resolved). Broad questions about no-training, secret ownership, rotation, incident owner, alert destination, and staging availability are out of scope for this reconciliation.

---

## 18. Recommended Implementation Sequence

1. Reconcile the contract draft with the existing Phase 5 implementation and `/ai-agents` foundation (this document).
2. Review with Fahad and Saad.
3. Commit the contract draft under `/ai-agents`.
4. Prepare the Make wiring plan.
5. Configure controlled production validation with marked test startup records.
6. Set/verify `move_to_review = needs_approval`.
7. Run controlled production validation only after Fahad approval.
8. Review audit logs and cleanup.
9. Limited activation if approved.
10. Wider activation only after clean validation.

---

## 19. Non-Goals

This draft does **not**: deploy production; activate AUTO-026; build Make scenarios; create secrets; send notifications; change startup status; bypass approvals; replace the developer implementation plan.

---

## 20. Review Checklist

For Fahad, Saad, and implementation reviewers:

- [ ] Repo/base context correct: `numu-old` / `main`, foundation at `/ai-agents`, Phase 5 + foundation unified on `main`.
- [ ] Core boundaries (§3) correct; Make = orchestration only.
- [ ] `prescreen` accepted as canonical; `pre_screen` = legacy compatibility data only (§4).
- [ ] Phase 5 areas correctly framed as **existing backend contracts to verify** (§0, §5–§12), not built from scratch.
- [ ] AgentRun fields (§7) reconcile with `agent_runs`; exact column names confirmed during validation.
- [ ] Structured output (§8) — allowed actions, fail-closed, one repair — confirmed.
- [ ] Handoff (§9) — logical idempotency guard and supersede lineage verified during validation.
- [ ] Notes (§10) — endpoints/schema accepted; append-only dedup responsibility understood.
- [ ] Approval (§11) — `move_to_review = needs_approval` confirmed for first validation.
- [ ] Execution Layer (§12) — decision/execution endpoints + expected-state + audit confirmed.
- [ ] Controlled production validation plan (§13) accepted; test records marked; no hard-delete.
- [ ] Idempotency ownership (§14) agreed, especially append-only notes and the Handoff guard.
- [ ] Make orchestration boundaries/sequence (§15) approved.
- [ ] Rollback/pause (§16) — admin-only, `ai.rollback`, not public APIs — confirmed.
- [ ] Open items (§17) have owners; resolved decisions not reopened.
- [ ] Implementation sequence (§18) accepted.
- [ ] Non-goals (§19) confirmed — approving this draft does not imply production activation.

---

*End of reconciled draft. Design/contract only — not production-ready, not deployed, not runtime-activated. Reconciled against `numu-old/main` (Phase 5 backend + `/ai-agents` foundation).*
