# Phase 5 — Handover Package (Managed Agents backend)

> Final handover for Numu Managed Agents Phase 5. Branch `feat/agent-knowledge-cache`
> (not pushed, no PR). The repository is the source of truth. This document is the
> single entry point; deep detail lives in the referenced Phase-5 artifacts.

---

## 1. Executive summary

The Managed Agents backend is **complete, additive, and end-to-end validated** on
`feat/agent-knowledge-cache`. It delivers: a registry-driven immutable Knowledge
Cache, a propose-only Agent Runtime (live-verified against `claude-sonnet-5`),
full AgentRun traceability, and production-grade idempotency at every hop
(run-create, decision, approval, execution, handoff, policy). All backend
architectural invariants were independently re-verified in the final audit.

- **Verdict:** `PHASE_5_COMPLETE_WITH_KNOWN_LIMITATIONS` (confidence 92/100).
- **Backend-code blockers:** 0. **Tests:** 79 Agents feature tests passing.
- **Remaining work is external** (Make scenario wiring) plus operational/product
  approvals — no further backend code is required for orchestration.

Reference artifacts (all committed): `PHASE_5_COMPLETION_REPORT.md`,
`FINAL_REPOSITORY_ANSWERS.md`, `FINAL_PHASE_5_AUDIT.md`,
`MAKE_NUMU_ORCHESTRATION_CONTRACT.md`, `E2E_VALIDATION_REPORT.md`,
`LIVE_RUNTIME_SMOKE_TEST_REPORT.md`, `HANDOFF_GAP_ANALYSIS.md`,
`DECISION_IDEMPOTENCY_ANALYSIS.md`, `ENVIRONMENT_RECONCILIATION_ANALYSIS.md`,
`ENVIRONMENT_RECONCILIATION_IMPLEMENTATION_PLAN.md`,
`G8_INTAKE_TRIAGE_POLICY_ANALYSIS.md`, `REPOSITORY_QUESTION_ANSWERS.md`.

## 2. Final architecture flow

```
SharePoint / OneDrive (long-term source of truth)
        │  Loader (registry-driven, env-switchable, validated, checksummed)
        ▼
Knowledge Cache (immutable, versioned, per-environment)  ← runtime source of truth
        │
Make (orchestration ONLY) ───────────────────────────────────────────────────────┐
  1 create AgentRun ─ POST /agent-runs (idempotency_key)                            │
  2 read knowledge ─ GET /agents/{key}/knowledge/meta                               │
  3 propose ──────── POST /agent-runtime/run  → PROPOSAL (executed:false)           │
  4 submit ───────── POST /agent-runs/{run}/decision                                │
        ▼                                                                           │
  AgentDecisionService (SOLE entry to business action)                              │
    → ActionPolicyService → { always_allow → execute | needs_approval → approval    │
                              | blocked }                                           │
    → ExpectedStateValidator → ExecutionRequestService → AgentActionExecutor        │
    → (human approve) → ApprovalExecutionService → AgentActionExecutor              │
  5 handoff (if move_to_review) ─ POST /handoffs (idempotent) → GET /handoffs/latest │
```

Runtime **proposes**; Numu **decides + executes** (single `AgentActionExecutor`
engine, two triggers: auto-allow and post-approval); Make **orchestrates**. The AI
never executes. One entry to business action.

## 3. All commits produced during Phase 5 (branch `feat/agent-knowledge-cache`)

Baseline `a9fa600` (registry Loader refactor + Handoff P1) on top of `4b4854f` /
`33d8e49` (Knowledge Cache base). Since the baseline (oldest → newest):

| # | Commit | Type | Summary |
|---|---|---|---|
| 1 | `29f916a` | feat | Agent Runtime endpoint (propose-only) — Task #3 |
| 2 | `47621d4` | docs | Repository question answers + reconciliation |
| 3 | `e6fbfc7` | docs | Live smoke test + env/handoff analyses + Phase 5 plan |
| 4 | `4e9c236` | docs | G8 intake_triage policy analysis |
| 5 | `dde6349` | feat | G8 active-policy uniqueness safeguard |
| 6 | `8e892f9` | feat | Handoff idempotency / replay protection (200) |
| 7 | `2aa01b7` | docs | Make ↔ Numu orchestration contract |
| 8 | `83ea211` | docs | Decision idempotency analysis |
| 9 | `51d9bb8` | feat | Decision idempotency (Option B + D) |
| 10 | `eaed88d` | docs | E2E validation report + env reconciliation plan |
| 11 | `be59a6c` | docs | Phase 5 completion report |
| 12 | `b45f0a1` | docs | Final repository answers |
| 13 | `52fd42b` | docs | Final Phase 5 consistency audit |
| + | *(this)* | docs | Phase 5 handover package |

Migrations this phase (all additive, reversible): `2026_07_04_120000_add_runtime_metrics_to_agent_runs`,
`2026_07_05_120000_add_active_policy_guard_to_agent_action_policies`,
`2026_07_05_130000_add_handoff_logical_idempotency_guard`,
`2026_07_05_140000_add_active_approval_guard_to_agent_approval_requests`.
(Baseline also: `2026_07_02_150000_augment_knowledge_traceability`, `2026_07_02_160000_add_investor_sequence_unique_to_handoffs`.)

## 4. Implemented features

- **Knowledge Cache + Loader:** `agent_knowledge_cache` (immutable, per-env, monotonic versions, pointer-only rollback, audit, checksum no-op); registry-driven ordered multi-file loading; env-switchable Local/Graph source; metadata/runtime read-API split.
- **Agent Runtime:** `POST /agent-runtime/run` — cache-only knowledge, PII-minimized context, registry-ordered prompt, config-selected model provider (Anthropic + Fake), structured-output validation (fail-closed + one repair), 12 traceability fields, propose-only (`executed:false`), deterministic fail-closed errors.
- **Handoff:** P1 hardening (semantic `source_run_id`, supersede 409 guard, investor sequence unique, sanitized `QueryException`), idempotency (logical-tuple, 200 replay), supersede/correction.
- **Decision/policy/approval/execution:** G8 active-policy uniqueness; decision idempotency (B: `agent_run_id` short-circuit; D: one pending approval per run); execution claim-idempotent (unchanged).

## 5. Idempotency guarantees

| Hop | Guarantee | Mechanism |
|---|---|---|
| Create run | Same key → same run | `AgentRunService::create` (unique key + race-catch) |
| Runtime | Retry-safe, no side effects | propose-only |
| Decision | One decision per `agent_run_id`; replay/different-action → same outcome | `priorDecision` short-circuit |
| Approval | One pending per run | `approvals_active_pending_unique` + firstOrCreate |
| Execution | Once per `(run, slug)` | `ExecutionRequestService` unique claim key |
| Handoff | Replay → existing row (200); one active per tuple | `createIdempotent` + `handoffs_active_logical_unique` |
| Policy | One active per `(agent, action, env)` | `policies_active_unique` |

**At-least-once safe end-to-end**, with no reliance on caller retry discipline.

## 6. Security guarantees

- API key only in the request header; never logged; no `Log::` of secrets; HTTP-error path throws the status code only (no body).
- No prompts and no model responses persisted (only token counts/model/latency).
- Runtime fails closed (preconditions → 404/422/409; invalid output → 422; provider error → 502).
- PII minimization: explicit allowlist excludes email/phone/applicant name/social URLs.
- All `/api/v1/ai/*` writes behind `McpAuth:full`; knowledge refresh/rollback behind admin permissions; every create/decision/supersede/rollback/runtime-run audited with correlation IDs; sanitized DB errors (409/413/422).

## 7. Production prerequisites (must be satisfied before live activation)

1. **Product sign-off:** confirm `move_to_review = always_allow` (auto-executes in `prod`) — G8.
2. **Canonical second-stage key:** repo seeds `prescreen`; doc prefers `pre_screen` — decide and, if changing, run the additive rename.
3. **Production active-duplicate pre-check:** before applying the guard migrations to production, prove **0** pre-existing active duplicates in `handoffs` / `agent_action_policies` / `agent_approval_requests` (supersede the known `id 3` handoff for startup 551 first).
4. **Operational setup:** dedicated Make service account/token + rotation; alert destinations; incident owner + incident-log location; Anthropic no-training/data-retention account setting.
5. **Controlled test policy:** written safe-testing rules + marked safe test records (staging if provisioned; otherwise controlled `prod` per §5.9).

## 8. Remaining external work (Make.com only — no backend change)

Build the scenarios per `MAKE_NUMU_ORCHESTRATION_CONTRACT.md`:
`trigger → POST /agent-runs → GET knowledge/meta → POST /agent-runtime/run →
POST /agent-runs/{run}/decision → (POST /handoffs on move_to_review) →
GET /handoffs/latest`. Use a dedicated `numu:full` token; map `production↔prod` /
`staging↔test`; honor the per-endpoint retry rules (decision = check-before-retry
via `GET /agent-runs/{run}`; run-create/runtime/handoff = safe to retry). Make must
not call the model, assemble prompts, validate output, decide policy, or execute.

## 9. Operational checklist before production

- [ ] Product owner confirms `move_to_review` policy (always_allow vs needs_approval).
- [ ] Decide `prescreen` vs `pre_screen`; apply additive rename if changing.
- [ ] Verify production `handoffs`/`agent_action_policies`/`agent_approval_requests` have 0 active duplicates; supersede `id 3` if present.
- [ ] Apply migrations in order (see §11); confirm each guard column + unique index.
- [ ] Set `ANTHROPIC_API_KEY` (+ optional per-agent model) in backend env/secret store; confirm no-training setting on the Anthropic account.
- [ ] Provision a dedicated Make service account + `numu:full` token; document rotation.
- [ ] Seed agents/policies (`AgentsSeeder`) and load + activate the `intake_triage` knowledge package for the target environment.
- [ ] Run one controlled dev/staging E2E (per `E2E_VALIDATION_REPORT.md`) and audit-verify.
- [ ] Configure alert destination + incident owner + incident-log location.
- [ ] Run the AUTO-025 evaluation fixture through the Runtime (training parity) before broad rollout.

## 10. Rollback strategy

- **Knowledge:** pointer-only `AgentKnowledgeLoader::activateVersion` (admin route, perm `ai.knowledge.rollback`); no content is destroyed.
- **Agent runtime state:** `AgentRuntimeService::pause`/`pauseAll` (enforced in `AgentDecisionService::pauseBlock`) halts new decisions immediately.
- **Migrations:** every guard migration has a working `down()` (drops the unique index then the generated column); reversible with no data loss.
- **Decision/handoff/policy data:** immutable + versioned — supersede/pointer moves, never destructive; a bad policy is reverted via `setPolicy()` (new version).
- **Deploy rollback:** revert the branch / redeploy the prior build (operational).

## 11. Recommended deployment sequence

```
1. Merge/deploy feat/agent-knowledge-cache to the target environment (code + config).
2. Pre-flight: prove 0 active duplicates in handoffs / policies / approvals.
   (Supersede the known id 3 handoff first if present.)
3. Run migrations in order:
   2026_07_02_150000, 2026_07_02_160000,
   2026_07_04_120000, 2026_07_05_120000, 2026_07_05_130000, 2026_07_05_140000.
   Verify each guard column + unique index applied.
4. Seed agents + policies; load + activate the intake_triage knowledge package.
5. Configure secrets (ANTHROPIC_API_KEY, no-training setting) + Make service token.
6. Controlled E2E in dev/staging (or marked prod test record): create run → runtime
   (live) → decision (safe env) → handoff; audit-verify; roll back test data.
7. Product sign-off on move_to_review policy; confirm prescreen/pre_screen.
8. Wire Make scenarios (external) per the orchestration contract.
9. Limited intake_triage activation on marked records; monitor audit + alerts.
10. Broad rollout only after step 9 is clean; defer AUTO-027+ multi-agent until then.
```

---

## Final verification (at handover)

- ✅ No uncommitted code changes; working tree clean (`git status` clean on `feat/agent-knowledge-cache`).
- ✅ All Phase-5 commits present on `feat/agent-knowledge-cache` (13 since baseline `a9fa600`, listed in §3).
- ✅ No `TODO` / `FIXME` / `HACK` / `XXX` markers in the managed-agents implementation (`app/Services/Agents/**`, `app/Http/Controllers/Api/V1/Ai/Agents/**`, phase migrations, `tests/Feature/Agents/**`).
- ✅ 79 Agents feature tests passing.
- Not pushed; no PR — per instruction.

**This is the final Phase 5 handover package.**
