# Phase 5 — Completion Report (Managed Agents backend)

> Final completion report for the Numu Managed Agents Phase-5 backend. The
> repository is the source of truth. Branch `feat/agent-knowledge-cache` (not
> pushed, no PR). Date: 2026-07-05.

## 1. Architecture summary

```
SharePoint / OneDrive (long-term source of truth)
        │ (Loader, registry-driven, env-switchable)
        ▼
Knowledge Cache (immutable, versioned, per-environment, Numu-owned = runtime source of truth)
        │
Make (orchestration ONLY) ─────────────────────────────────────────────────────────┐
  create AgentRun → read knowledge meta → call Runtime → submit decision → (handoff) │
        │                                   │                    │                   │
        ▼                                   ▼                    ▼                   ▼
   POST /agent-runs              POST /agent-runtime/run   POST /agent-runs/{run}/decision
   (idempotency_key)             (PROPOSAL only,           (AgentDecisionService)
                                  executed:false)                │
                                                                 ▼
                          Numu business pipeline (sole authority; unchanged):
                          ActionPolicyService → ApprovalService → ExpectedStateValidator
                          → ExecutionRequestService → AgentActionExecutor → HandoffService
```

Runtime **proposes**; Numu **decides and executes**; Make **orchestrates**. The AI
never executes. One entry to business action: `POST /agent-runs/{run}/decision`.

## 2. All commits (since the `a9fa600` baseline)

| Commit | Type | Summary |
|---|---|---|
| `29f916a` | feat | Agent Runtime endpoint (propose-only) — Phase 5 Task #3 |
| `47621d4` | docs | Repository question answers + reconciliation |
| `e6fbfc7` | docs | Live smoke test + env/handoff analyses + Phase 5 plan |
| `4e9c236` | docs | G8 intake_triage policy analysis |
| `dde6349` | feat | G8 active-policy uniqueness safeguard |
| `8e892f9` | feat | Handoff idempotency / replay protection (idempotent-200) |
| `2aa01b7` | docs | Make ↔ Numu orchestration contract |
| `83ea211` | docs | Decision idempotency analysis |
| `51d9bb8` | feat | Decision idempotency (Option B + D) |
| `eaed88d` | docs | E2E validation report + environment reconciliation plan |

Folded into the `a9fa600` baseline (earlier in the same effort): registry-driven
Knowledge Loader refactor + Handoff P1 hardening. Prior base: `4b4854f`
(Knowledge Cache), `33d8e49` (per-environment pointers + content validation).

## 3. All implemented features

**Knowledge Cache + Loader**
- Immutable, append-only versioned cache (`agent_knowledge_cache`); monotonic `vN` per `(agent_key, environment)`; pointer-only rollback; audit chain; checksum reproducibility.
- Registry-driven Loader (`AGENT_KNOWLEDGE_REGISTRY.json`): ordered multi-folder package, no hardcoded file names; rejects missing/empty required, duplicate path/order, wildcard, traversal, absolute paths, oversized/binary/script content.
- Env-switchable source adapter (Local now, Graph/SharePoint ready-but-inert).
- Read API: metadata/manifest vs runtime-content split (`/knowledge`, `/knowledge/runtime`, `/versions`, `/meta`, `/{version}`).

**Agent Runtime (Phase 5 Task #3)**
- `POST /api/v1/ai/agent-runtime/run` — loads active knowledge from the cache (never SharePoint), builds PII-minimized context, assembles a registry-ordered prompt, calls Claude via a config-selected provider (Anthropic; Fake for tests), validates structured output (fail-closed + one schema-repair), stamps 12 traceability fields, returns a PROPOSAL (`executed:false`).
- Deterministic fail-closed errors: 404 / 422 precondition / 409 knowledge_unavailable / 422 invalid_model_output / 502 provider_error.

**Handoff subsystem**
- P1 hardening: semantic `source_run_id` validation, supersede-chain 409 guard, investor sequence uniqueness, sanitized `QueryException` mapping (409/413/422).
- Idempotency: logical-tuple dedup → replay returns the existing row (200), backed by an active-scoped unique guard.

**Decision / policy / approval / execution**
- G8: active-policy uniqueness (one active policy per `agent/action/environment`).
- Decision idempotency (B + D): `decide()` short-circuits on an already-decided run (`agent_run_id` boundary); one pending approval per run; execution already claim-idempotent.

**Migrations added this phase** (all additive, reversible):
`2026_07_04_120000_add_runtime_metrics_to_agent_runs`,
`2026_07_05_120000_add_active_policy_guard_to_agent_action_policies`,
`2026_07_05_130000_add_handoff_logical_idempotency_guard`,
`2026_07_05_140000_add_active_approval_guard_to_agent_approval_requests`.
(Plus `2026_07_02_150000_augment_knowledge_traceability` + `2026_07_02_160000_add_investor_sequence_unique_to_handoffs` in the baseline.)

## 4. Production safety guarantees

- **Immutability:** knowledge versions and handoffs are append-only; content mutation/deletion throws; only pointers/links flip.
- **Fail-closed runtime:** every precondition checked before the model call; invalid output (after one repair) submits no decision; provider errors are deterministic; no prompt/response persisted; API key never logged/returned.
- **PII minimization:** runtime context uses an explicit allowlist (excludes email/phone/applicant name/social URLs).
- **Sanitized DB errors:** `QueryException` mapped to 409/413/422 without leaking SQL.
- **Single business authority:** all decision/policy/approval/execution logic is Numu-internal; Make cannot execute.
- **Auditability:** every create/decision/supersede/rollback/runtime-run writes an `activity_logs` event with actor + correlation IDs; replays add no duplicate audit.
- **Reversibility:** every guard migration has a working `down()`; rollbacks are pointer/versioned, never destructive.

## 5. Idempotency guarantees (per hop)

| Hop | Guarantee | Mechanism |
|---|---|---|
| Create run | Same `idempotency_key` → same run | `AgentRunService::create` (unique key + race-catch) |
| Runtime | Retry-safe; no side effects | propose-only; re-stamps traceability only |
| Decision | One decision per `agent_run_id`; replay (even different action) → same outcome, no re-execution, no duplicate approval | Option B short-circuit + Option D unique pending-approval guard |
| Execution | Native action + notifications run once per `(run, slug)` | `ExecutionRequestService` unique claim key |
| Approval | One pending approval per run | `approvals_active_pending_unique` + firstOrCreate |
| Handoff | Replay → existing row (200); at most one active per logical tuple | `createIdempotent` + `handoffs_active_logical_unique` |
| Policy | One active policy per `(agent, action, env)` | `policies_active_unique` |

**Net:** the full orchestration chain is safe under at-least-once delivery with no reliance on caller retry discipline.

## 6. E2E validation results

- **Live smoke test** (`LIVE_RUNTIME_SMOKE_TEST_REPORT.md`): real `claude-sonnet-5`, knowledge from cache only, single traceability write, no side effects, rolled back.
- **Controlled E2E** (`E2E_VALIDATION_REPORT.md`): full chain — create run → knowledge meta → runtime (live) → decision (test env) → handoff — validated idempotent at every hop; `executed:false`; 0 approvals / 0 executions / exactly 1 handoff; rolled back (0 rows persisted).
- **Automated tests:** 79 Agents feature tests passing (Knowledge Cache, Runtime, Handoff hardening + idempotency, Decision idempotency, Managed-Agents invariants).

## 7. Remaining work — classified

**Backend work (Numu):** none required for orchestration. Optional/deferred only:
- Environment reconciliation (`ENVIRONMENT_RECONCILIATION_IMPLEMENTATION_PLAN.md`) — additive, post-integration, fail-safe today.
- Handoff `retention_class` / `schema_version` allow-lists (hardening; non-blocking).
- Runtime context expansion (notes/documents, full handoff history) — deferred (§8.8).
- Graph/SharePoint production source enablement (once Azure consent granted).
- A live *repaired-output* runtime run (repair path is unit-tested, not yet live-exercised).

**Make.com work (external, no backend change):**
- Build the scenarios per `MAKE_NUMU_ORCHESTRATION_CONTRACT.md`: trigger → create run → knowledge meta → runtime → submit decision → (handoff on move_to_review) → verify read-back.
- Use a dedicated service token; map `production↔prod` / `staging↔test` per §8; honor per-endpoint retry rules.

**Operational work:**
- Provision a dedicated Make service account + token (rotation policy).
- Confirm the `move_to_review = always_allow` product intent (G8 §7) before any live `prod` decision.
- Provider data-retention / no-training account settings; alert destinations + incident-log location.
- Controlled production test-record marking before any live `prod` write.

**Optional future improvements:**
- Push trigger (webhook/event) for AgentRun completion (currently poll-only).
- Least-privilege handoff/knowledge scopes; read/validation-failure audit rows.
- Multi-agent rollout (AUTO-027+) after `intake_triage` is live end-to-end.

## 8. Architecture invariants — reconfirmed against the repository

- ✅ **No parallel pipeline** — the only entry to business action is `POST /agent-runs/{run}/decision` → `AgentDecisionService`; the runtime feeds it, never bypasses it. No second decision/approval/execution/handoff path exists.
- ✅ **No business logic in Make** — Make sequences calls and copies values; it holds no policy, prompt, validation, or execution logic (verified: no such code path is Make-facing).
- ✅ **Numu is the sole business authority** — policy resolution, approval, expected-state validation, native execution, notifications, audit, and all deduplication are Numu-internal services.
- ✅ **Claude is proposal-only** — invoked solely inside the Runtime; output is validated and used only as a recommendation.
- ✅ **Runtime is propose-only** — returns a proposal, stamps traceability (a single AgentRun metadata write), and performs no decide/approve/execute (`executed:false`).

## 9. Status

**Phase 5 backend: COMPLETE, production-safe, and end-to-end validated.** All work
is additive on `feat/agent-knowledge-cache`; no new architecture, service,
repository, package, module, or parallel execution path was introduced. Not
pushed; no PR opened. Next decision points: push the branch / open a PR / proceed
to external Make scenario implementation.
