# Final Phase 5 Consistency Audit (repository-based)

> Read-only final audit. No implementation, no architecture change, no commit.
> Every finding is independently re-verified against the live code on branch
> `feat/agent-knowledge-cache`. The repository is the source of truth.

**Audit basis:** re-scan of routes, controllers, services, models, migrations,
tests + re-read of all Phase-5 reports. **Suite:** 79 Agents feature tests passing.
**Verdict legend:** VERIFIED · PARTIALLY VERIFIED · NOT VERIFIED · FAILED.

---

## Architecture

| Invariant | Verdict | Evidence |
|---|---|---|
| Exactly one business execution pipeline | **VERIFIED** | Single execution engine `AgentActionExecutor::executeForRun` (`app/Services/Agents/AgentActionExecutor.php:48`). It has exactly two legitimate callers — `AgentDecisionService::executeNow` (`:135`, always_allow) and `ApprovalExecutionService` (`:101`, after a human approves). One engine, two triggers of the **same** pipeline; no third/parallel executor exists. |
| Runtime is propose-only | **VERIFIED** | `RuntimeRunner::run` returns `'executed' => false`; the only write is a traceability `save()` (`RuntimeRunner.php:209`). |
| Claude never executes business actions | **VERIFIED** | Claude is invoked only inside `AnthropicModelProvider::complete`; its output is validated (`StructuredOutputValidator`, rejects execution-claim keys) and returned as a proposal — it cannot reach `AgentActionExecutor`. |
| Make contains no business logic | **PARTIALLY VERIFIED** | The repository proves all business logic (policy/approval/expected-state/execution) is Numu-internal and Make would only call REST endpoints (`MAKE_NUMU_ORCHESTRATION_CONTRACT.md`). The actual Make scenarios are **external and not yet built**, so they cannot be repo-verified. |
| Numu remains the sole business authority | **VERIFIED** | `AgentDecisionService` + `ActionPolicyService`/`ApprovalService`/`ExpectedStateValidator`/`ExecutionRequestService`/`AgentActionExecutor` are all Numu-internal; the runtime has no path into them. |
| No parallel execution path exists | **VERIFIED** | Grep of `AgentActionExecutor`/`executeForRun` shows the two callers above only; the Runtime namespace contains none. |

## Runtime

| Invariant | Verdict | Evidence |
|---|---|---|
| Runtime writes only traceability | **VERIFIED** | Only mutating call in `app/Services/Agents/Runtime/*` is `RuntimeRunner.php:209` `->save()` (the 12 traceability columns). |
| Runtime cannot create approvals | **VERIFIED** | No `AgentApprovalRequest::create` in the Runtime namespace — only `AgentApprovalRequest::query()` (read) in `RuntimeContextBuilder:138`. |
| Runtime cannot create executions | **VERIFIED** | No `ExecutionRequest`/`claim` reference in the Runtime namespace. |
| Runtime cannot mutate startup state | **VERIFIED** | `RuntimeRunner.php:66` is `Startup::find` (read); no `Startup` `save`/`update` anywhere in the namespace. |
| Runtime cannot create handoffs | **VERIFIED** | Only `HandoffService::latest` (read) is used (`RuntimeContextBuilder:66`); no `createIdempotent`/`create`/`supersede`. |

## Decision pipeline

| Invariant | Verdict | Evidence |
|---|---|---|
| Decision endpoint is fully idempotent | **VERIFIED** | `AgentDecisionService::decide` short-circuits via `priorDecision` (`:60` guard, `:275` impl) on an already-decided run; `DecisionIdempotencyTest` (6) + rolled-back real-DB proof (`51d9bb8`). |
| Approval creation is idempotent | **VERIFIED** | `ApprovalService::createFromRun` = `pendingFor` firstOrCreate + `UniqueConstraintViolationException` race-catch; DB guard `approvals_active_pending_unique` (`2026_07_05_140000`). |
| Execution remains idempotent | **VERIFIED** | `ExecutionRequestService::claim` atomic on unique `idempotency_key` (`agent-run-{id}:{slug}`); native action/notifications never repeat; `ManagedAgentsInvariantsTest::execution_request_claim_is_idempotent`. |
| Replays cannot create side effects | **VERIFIED** | Real-DB proof: replay + different-action on a decided run → same outcome, `no_execution_triggered=YES`, one `agent.decision`/`agent.approval_created` audit only. |

## Handoffs

| Invariant | Verdict | Evidence |
|---|---|---|
| Handoff creation is idempotent | **VERIFIED** | `HandoffService::createIdempotent`/`findActiveByLogicalTuple` + DB guard `handoffs_active_logical_unique` (`2026_07_05_130000`); `HandoffHardeningTest` (200 on replay). |
| Replay protection works | **VERIFIED** | Replay (even different body) returns the existing row; concurrent race caught → existing row; rolled-back real-DB proof (`8e892f9`). |
| Supersede/correction logic remains valid | **VERIFIED** | `HandoffService::supersede` (create-as-superseded → promote) keeps a single active head; re-supersede → 409; immutability guard on `Handoff`; tests green. |
| Audit behavior is correct | **VERIFIED** | `agent.handoff_created` on create, `agent.handoff_superseded` on supersede, **no** audit on an idempotent duplicate hit (proven). |

## Knowledge

| Invariant | Verdict | Evidence |
|---|---|---|
| Runtime loads only from the Knowledge Cache | **VERIFIED** | `RuntimeRunner` injects `AgentKnowledgeReader` (`:34`) and reads `->active($agentKey,$environment)` (`:72`) — a DB query on `agent_knowledge_cache`. |
| Runtime never loads directly from SharePoint | **VERIFIED** | No `KnowledgeSource`/`LocalFolderSource`/`GraphSharePointSource`/`->fetch()` in the Runtime namespace (only docblock comments stating "never SharePoint"). |
| Registry ordering remains authoritative | **VERIFIED** | `AgentKnowledgeLoader` sorts files by `load_order` and stores the ordered list; `RuntimePromptBuilder` concatenates the stored `files_json` in load order; `AgentKnowledgeCacheTest::load_order_follows_registry_not_disk_or_alpha`. |

## Environment handling

| Invariant | Verdict | Evidence |
|---|---|---|
| Current split remains fail-safe | **VERIFIED** | Decision validates `in:dev,test,prod` (`AgentRunController.php:30`); runtime/loader validate against `config('agent_knowledge.environments')` = `dev|staging|production` (`RuntimeRunner.php:221`). A value valid in one namespace is rejected by the other. |
| No silent cross-environment execution is possible | **VERIFIED** | A wrong env yields `422` (decision) or `409 knowledge_unavailable` (runtime) — never a mis-scoped action. |

## Security

| Invariant | Verdict | Evidence |
|---|---|---|
| No API keys logged | **VERIFIED** | `AnthropicModelProvider` places the key only in the `x-api-key` header (`:39`); no `Log::` calls; the HTTP-error path throws the status code only (no body); the transport-catch (`:59`) includes the transport error message (no headers/key). |
| No prompts persisted | **VERIFIED** | `RuntimeRunner` writes only token counts/model/latency; no prompt text stored (grep: none). |
| No model responses persisted | **VERIFIED** | The raw model text is validated then discarded; only the sanitized proposal is returned (not stored on the run). |
| Runtime fails closed | **VERIFIED** | Preconditions before the model call (404/422/409); invalid output after one repair → `422 invalid_model_output`; provider error → `502` — deterministic, no partial state. |
| PII minimization remains intact | **VERIFIED** | `RuntimeContextBuilder::startup` allowlist excludes `email`/`phone_number`/`applicant_full_name`/`first_name_*`/`linkedin_url`. |

---

## Production readiness

### Remaining blockers (must resolve before LIVE production activation — none are backend-code defects)
1. **Make orchestration scenarios** — external; not built. Required for automation (the manual/REST path works today).
2. **Product sign-off that `move_to_review = always_allow`** — an intake `move_to_review` decision auto-executes in `prod`; confirm before any live `prod` decision (G8).
3. **Production active-duplicate pre-check** — before applying the guard migrations (`handoffs_active_logical_unique`, `approvals_active_pending_unique`, `policies_active_unique`) to production, prove 0 pre-existing active duplicates (the known `id 3` handoff for startup 551 must be superseded first).
4. **Operational setup** — dedicated Make service account/token + rotation, alert destinations, incident owner + incident-log location, provider no-training/data-retention setting.

### Remaining optional improvements (non-blocking)
1. Environment reconciliation onto `dev/test/prod` (fail-safe today; `ENVIRONMENT_RECONCILIATION_IMPLEMENTATION_PLAN.md`).
2. `schema_version` + `retention_class` allow-lists (currently free strings).
3. Runtime context expansion (notes/documents, full handoff history).
4. A live *repaired-output* runtime run (repair path is unit-tested, not live-exercised).
5. Runtime-owned Handoff trigger / production push event (poll-only today).
6. Least-privilege scopes; read/validation-failure audit rows.
7. AUTO-025 evaluation-fixture run through the Runtime.

### Assumptions still requiring human/product approval
1. `move_to_review = always_allow` is the intended intake policy.
2. Canonical second-stage key: repo seeds **`prescreen`**; doc prefers **`pre_screen`** — pick one.
3. Whether a **staging** environment is provisioned.
4. Anthropic account **no-training / data-retention** configuration.
5. Operational incident owner + rollback/pause authority.

---

# FINAL VERDICT

## PHASE_5_COMPLETE_WITH_KNOWN_LIMITATIONS

Every architectural, runtime, decision, handoff, knowledge, environment, and
security invariant is **VERIFIED** against the live code, with one
**PARTIALLY VERIFIED** ("Make contains no business logic" — the backend enforces
authority, but the external Make scenarios are unbuilt and cannot be repo-verified).
The backend implementation is complete and validated end-to-end (live model +
idempotent at every hop). It is **not** unqualified "PRODUCTION_READY" only because
of external Make wiring, pending product/operational approvals, and the
production-data pre-checks required before the guard migrations apply there — none
of which are backend-code defects.

## Metrics

- **Blocking issues (backend-code defects / failed invariants):** 0
- **Blocking prerequisites for live go-live (non-code):** 4 (Make wiring, `move_to_review` sign-off, prod active-duplicate pre-check, operational setup)
- **Non-blocking issues (optional improvements):** 7
- **Architectural risks:** 2 (dual environment vocabulary — fail-safe; `prescreen` vs `pre_screen` naming — unenforced)
- **Production risks:** 4 (prod active-duplicate before migration; `move_to_review` auto-execute; provider no-training unset; operational incident/rotation undefined)
- **Confidence score:** **92 / 100** (backend implementation + invariants: very high; deductions for external Make wiring, NOT-VERIFIED operational items, and pending product approvals)

---

*No implementation, architecture change, or commit was performed. This document is the final Phase-5 consistency audit only.*
