# Final Repository Answers — Unified Developer Request (v02, 2026-07-04)

> Strict, repository-grounded answers to every question in the Unified Developer
> Request, using the actual state of branch `feat/agent-knowledge-cache` and all
> Phase-5 artifacts. **The repository is the single source of truth.** Where a
> question cannot be proven from the repository it is marked
> `NOT VERIFIED IN REPOSITORY`. No implementation was performed to produce this
> document.

**Branch:** `feat/agent-knowledge-cache` (not pushed, no PR).
**Phase-5 commits:** `29f916a` (Runtime) · `47621d4`, `e6fbfc7`, `4e9c236` (docs) · `dde6349` (G8) · `8e892f9` (Handoff idempotency) · `2aa01b7`, `83ea211` (docs) · `51d9bb8` (Decision idempotency) · `eaed88d` (E2E+env docs) · Phase-5 completion report. Baseline `a9fa600` (registry refactor + Handoff P1); `4b4854f`/`33d8e49` (Knowledge Cache).
**Test suite:** 79 Agents feature tests passing.

---

# Question 1 — Section 6: Backend Repository Verification

## Question
Verify the real backend repository: path/URL, framework, current branch (and whether `feat/agent-knowledge-cache` exists), setup, test command, code locations, migrations (incl. `2026_07_02_150000_augment_knowledge_traceability`), risks.

## Repository Evidence
- file: `composer.json`, `artisan`, `routes/api.php`, `app/`, `config/`, `database/migrations/`
- git remote: `origin = https://dev.azure.com/Lun-Dev/numu/_git/numu`
- branch: `feat/agent-knowledge-cache` (HEAD reachable from the commits above)
- migration: `database/migrations/2026_07_02_150000_augment_knowledge_traceability.php` (present)
- report: `docs/managed-agents/REPOSITORY_QUESTION_ANSWERS.md`

## Verified Answer
- **Repository:** Azure DevOps `Lun-Dev/numu/_git/numu`; local dev checkout at `C:\xampp\htdocs\projects\lun\numu angels`.
- **Framework/language:** Laravel `^12` (running 12.58.0), PHP `^8.2` (8.2.12); DB **MariaDB 10.4.32**; queue+cache = `database` driver.
- **Branch:** `feat/agent-knowledge-cache` **exists** (created this phase from `main`; it did not exist before — `REPOSITORY_QUESTION_ANSWERS.md` correction).
- **Test command:** `php artisan test` (Agents: `php artisan test tests/Feature/Agents`).
- **Code locations:** `app/Services/Agents/Knowledge/*` (`AgentKnowledgeRegistry`, `AgentKnowledgeLoader`, `AgentKnowledgeReader`, `KnowledgeValidator`, `LocalFolderSource`, `GraphSharePointSource`); `app/Services/Agents/Runtime/*` (Runtime); `app/Models/{AgentRun,AgentKnowledgeCache,Handoff}.php`; `app/Http/Controllers/Api/V1/Ai/Agents/*`; `app/Services/Agents/{AgentDecisionService,ActionPolicyService,ApprovalService,ExpectedStateValidator,AgentActionExecutor,HandoffService,ExecutionRequestService,AgentRunService}.php`; envelope `AgentApiController`; routes `routes/api.php`.
- **Migration present:** yes (see evidence).

## Confidence
HIGH (code/branch/migration). LOW that this exact tree deploys to `dashboard.numuangels.net` — no deploy config in repo (`.env` `APP_URL=http://localhost`).

## Reconciliation Notes
Doc §3.3 says "real backend not verified" — now verified as this repo. Doc §6 asks "does branch exist" — it did not before Phase 5; created during it.

## Production Impact
None — verification only. Deployment-target confirmation is an operational item.

---

# Question 2 — 7.1 Runtime activation checklist (Training-Pack / Core-Skill approval metadata)

## Question
Confirm whether a runtime activation checklist verifies Training Pack / Core Skill approval metadata (per 5.3) before loading the runtime package.

## Repository Evidence
- class: `App\Services\Agents\Knowledge\AgentKnowledgeRegistry`, `AgentKnowledgeLoader`, `KnowledgeValidator`
- report: `REPOSITORY_QUESTION_ANSWERS.md` (Q7.1)

## Verified Answer
`NOT VERIFIED IN REPOSITORY` — **no such check exists.** The loader validates file presence/order/required flags/content safety and checksums; it reads **no** "Approved v1.0" approval-metadata flag before activation. Approval status is not a field the code consults.

## Confidence
HIGH (feature is absent).

## Reconciliation Notes
Decision 5.3 approves NUMU_CORE_SKILL.md v1.0; the backend does not gate activation on approval metadata. If required, it would be new (out of scope this phase).

## Production Impact
Low — package composition/integrity is enforced by the registry + checksums; human approval is an external governance step.

---

# Question 3 — 7.2 Official API routes + authoritative enums/limits

## Question
Confirm official production routes and authoritative enums/limits/validation for Handoff/AgentRun/decision/approval/Knowledge/Runtime.

## Repository Evidence
- file: `routes/api.php`; controllers under `app/Http/Controllers/Api/V1/Ai/Agents/`
- models: `AgentRun`, `AgentActionPolicy`, `AgentApprovalRequest`, `ExecutionRequest`, `AgentRuntimeConfig`
- report: `MAKE_NUMU_ORCHESTRATION_CONTRACT.md` (§H/§I), `REPOSITORY_QUESTION_ANSWERS.md` (§H/§I)

## Verified Answer
**Routes (all `/api/v1/ai`, `McpAuth:full` + `ai.cb` + `throttle:600,1`):** `POST agent-runs`, `GET agent-runs`, `GET agent-runs/{run}`, `PATCH agent-runs/{run}/analysis-result`, `POST agent-runs/{run}/decision`, `POST agent-runs/{run}/execution-results`, `POST agent-runs/{run}/cancel`; `POST agent-runtime/run`; `GET agents/{agentKey}/knowledge` (+ `/runtime`, `/versions`, `/meta`, `/{version}`); `POST handoffs`, `GET handoffs/latest`, `GET handoffs`, `GET handoffs/{handoff}`, `POST handoffs/{handoff}/supersede`; approvals `GET/POST`. Knowledge **refresh/activate** are admin web routes (`routes/web.php`, `permission:ai.knowledge.refresh|rollback,admin`), not `/api/v1/ai`.

**Enums (verified from model constants):** AgentRun analysis `queued|running|completed|failed|cancelled`; execution `not_started|executing|executed|blocked|cancelled|failed`; policy `always_allow|needs_approval|blocked`; approval status `pending|approved|rejected|cancelled|expired`, risk `none|expected|high`; execution-request `processing|succeeded|failed|cancelled|duplicate|deferred`; runtime provider `anthropic|fake`, default model `claude-sonnet-5`. `action_slug` for intake_triage = `move_to_review|reject|stop` (`config/agent_runtime.php`); generic `action_slug` = free string ≤120.

**Limits:** `schema_version` string ≤40 (**no allow-list**); `retention_class` string ≤40 default `standard` (**no allow-list**); handoff JSON fields `nullable|array` (**no field-size cap** — DB errors mapped 409/413/422); knowledge per-file `max_file_size` = 1 MB.

## Confidence
HIGH (routes + code enums). Startup group/status/action identifiers = `NOT VERIFIED IN REPOSITORY` as fixed enums — they are DB rows (`Group`, `LabelOption`), not code constants.

## Reconciliation Notes
Doc §5.2 default `/api/v1/ai` is **confirmed correct**. Doc §9 expects "invalid schema_version/retention_class → reject" — the repo does **not** reject (free strings); **repository is authoritative** (see Conflicts).

## Production Impact
Medium — callers must not rely on backend rejecting unknown `schema_version`/`retention_class`; the fields are stored as-is.

---

# Question 4 — 7.3 / 9. Canonical second-stage agent key (pre_screen vs prescreen)

## Question
Does the backend enforce/seed `prescreen`? Is aliasing/migration needed? Where is the canonical value enforced?

## Repository Evidence
- file: `database/seeders/AgentsSeeder.php:40` (`'agent_key' => 'prescreen'`)
- file: `config/agent_knowledge_registry.json` (only `intake_triage` enabled)
- controller: `HandoffController::store` (`from_agent`/`to_agent`/`stage` = `string|max:80`, no enum)

## Verified Answer
```
Official key:        prescreen  (seeded agent_key in AgentsSeeder.php:40 — NO underscore)
Compatibility notes: from_agent/to_agent/stage on handoffs are FREE strings (<=80);
                     the P1 semantic check only requires from_agent == source_run.agent_key
                     and stage == source_run.agent_key. There is NO enum enforcing
                     pre_screen vs prescreen anywhere.
Migration needed:    To adopt pre_screen as canonical, rename the seeded agent_key
                     'prescreen' -> 'pre_screen' in AgentsSeeder + any agent_runtime_config
                     / agent_action_policies rows keyed on it. Handoff fields need no
                     migration (free strings). No production intake_triage->pre_screen
                     handoff flow is wired yet, so no hard conflict exists today.
Alias behavior:      None in code.
Where to enforce:    Currently nowhere (free strings). If enforced, it would be the
                     seeded agent_key + optionally handoff validation.
```

## Confidence
HIGH — the seeded value `prescreen` is proven; the absence of enforcement is proven.

## Reconciliation Notes
**CONFLICT:** Decision 5.1 prefers `pre_screen`; the backend **seeds `prescreen`**. Repository is authoritative on current state: it is `prescreen`, unenforced on handoff fields. Aligning to `pre_screen` is an unimplemented, additive change.

## Production Impact
Medium — until reconciled, a handoff's `to_agent` could be either spelling; downstream consumers must not assume one. Only `intake_triage` is registry-enabled, so no live second-stage runtime exists yet.

---

# Question 5 — 7.4 Handoff Supersede / Correction contract

## Question
Provide the supersede/correction contract, or formally defer.

## Repository Evidence
- route: `POST /api/v1/ai/handoffs/{handoff}/supersede`
- method: `HandoffController::supersede`, `HandoffService::supersede`
- model: `Handoff` (immutability guard; `superseded_by`)
- test: `HandoffHardeningTest::{test_supersede_live_handoff_links_chain, test_superseding_an_already_superseded_handoff_returns_409, test_supersede_promotes_correction_and_latest_returns_it}`
- commit: `a9fa600` (P1), `8e892f9` (idempotency-compatible reorder)

## Verified Answer
**Option A — IMPLEMENTED.** `POST /api/v1/ai/handoffs/{id}/supersede` (`McpAuth:full`). Creates a NEW immutable handoff inheriting `startup_id|investor_id, source_run_id, from_agent, to_agent, stage, schema_version, retention_class` + merged new payload; sets `old.superseded_by = new.id`, `new.superseded_by = null`; the correction becomes the active head; `GET latest` returns it. **Repeated supersede of an already-superseded handoff → 409 `conflict`** (locked re-check). Original is immutable (model guard; no `updated_at`; delete throws). Audit: `agent.handoff_superseded`. No branching chains (single active head enforced by `handoffs_active_logical_unique`).

## Confidence
HIGH.

## Reconciliation Notes
Decision 5.5 (deferred) is now **resolved to Option A**. The preferred behavior in 7.4 matches the implementation exactly.

## Production Impact
Low — correction is safe, immutable, deterministic.

---

# Question 6 — 7.5 Handoff Backend Idempotency

## Question
A1 key / A2 tuple / B append-only + caller dedup?

## Repository Evidence
- migration: `2026_07_05_130000_add_handoff_logical_idempotency_guard.php`
- method: `HandoffService::createIdempotent`, `findActiveByLogicalTuple`; `HandoffController::store` (200/201)
- test: `HandoffHardeningTest::{test_duplicate_logical_handoff_is_idempotent_200, test_create_after_supersede_returns_the_active_correction, test_duplicate_active_logical_handoff_insert_is_blocked_by_guard}`
- commit: `8e892f9`; report: `HANDOFF_GAP_ANALYSIS.md`

## Verified Answer
**Option A2 — IMPLEMENTED (idempotent-200 variant).** Logical tuple = `(startup_id|investor_id, source_run_id, from_agent, to_agent, stage, schema_version)`. A replayed CREATE returns the existing active handoff with **HTTP 200**, no new row, no second audit, no supersede. Enforced by app-level `createIdempotent` + an active-scoped DB unique guard (`handoffs_active_logical_unique`, a STORED generated column that holds the tuple key only while `superseded_by IS NULL`). Concurrent race → the losing insert is caught and resolved to the existing row. Acceptance criteria (no new row, no second `handoff_created`, stable `GET latest`) verified.

## Confidence
HIGH.

## Reconciliation Notes
Doc §3.2 "backend idempotency absent (duplicate created id 3)" is now **superseded** — idempotency is implemented. The "temporary pre-POST duplicate check" rule (5.6) is now optional (backend enforces).

## Production Impact
Low — duplicate delivery is safe/deterministic.

---

# Question 7 — 7.6 Production Handoff Trigger

## Question
Runtime-owned / Make fallback / Hybrid; trigger conditions; flow.

## Repository Evidence
- service: `AgentRunService` (sets `completed_at`, emits **no** event); grep of `app/Services/Agents` finds no `Event::`/`dispatch`/webhook
- route: `GET /api/v1/ai/agent-runs` (pollable)
- report: `REPOSITORY_QUESTION_ANSWERS.md` (Q13), `MAKE_NUMU_ORCHESTRATION_CONTRACT.md` (§7)

## Verified Answer
**No automated production Handoff trigger exists in the backend.** There is no webhook/event/queue on AgentRun completion — only a **pollable** `GET /agent-runs`. Handoffs are created solely by an explicit authenticated `POST /handoffs` (now idempotent). Therefore, per Decision 5.7, the current viable owner is **Make (temporary fallback)** calling `POST /handoffs` after a completed run — the REST write path is confirmed, and backend idempotency now backs the duplicate pre-check. A **Runtime-owned** trigger would be new work (not implemented).

## Confidence
HIGH (trigger not implemented; poll-only proven).

## Reconciliation Notes
Decision 5.7 default "Runtime/backend owns creation" is **not yet implemented**; the only wired path is Make → `POST /handoffs`. The §10 gate "Runtime-to-live-Handoff persistence must wait" remains: no runtime code persists handoffs.

## Production Impact
Medium — full automation requires either building the backend trigger or approving Make as the handoff creator (idempotency makes this safe).

---

# Question 8 — 7.7 Handoff id 3 impact + cleanup

## Question
Does the duplicate affect `GET latest` / downstream; safest cleanup path.

## Repository Evidence
- method: `HandoffService::latest` (`whereNull(superseded_by)->orderByDesc(sequence_number)`)
- data: dev `handoffs` table is EMPTY (0 rows — verified during Handoff idempotency pre-flight)

## Verified Answer
- **GET latest behavior (by code):** returns the highest-`sequence_number` non-superseded row, so a later duplicate (higher sequence) **would** be returned instead of the earlier one — matching the reported observation for startup 551.
- **Downstream:** a consumer reading `latest` would see the duplicate; the new **idempotency guard prevents any NEW such duplicate**.
- **Safest cleanup path (now that supersede exists):** `POST /handoffs/{original}/supersede` (or point the duplicate’s `superseded_by`) to remove the duplicate from `latest`, preserving history.
- The specific rows `id 2`/`id 3` (startup 551) are **`NOT VERIFIED IN REPOSITORY`** — they are production/Data-Store data; this dev DB has 0 handoffs.

## Confidence
HIGH (behavior). LOW on the specific production rows (not in repo).

## Reconciliation Notes
Decision 5.4 keeps id 3 untouched — consistent; the cleanup mechanism (supersede) is now available. On any environment holding an active logical duplicate, it must be superseded **before** the idempotency migration can apply (documented in the migration).

## Production Impact
Medium — the production `handoffs` table must be checked for active logical duplicates before deploying `2026_07_05_130000`.

---

# Question 9 — 7.8 PII / Anthropic implementation

## Question
PII minimization + excluded fields; prompt/response logging; provider data-retention/no-training; audit metadata store.

## Repository Evidence
- method: `RuntimeContextBuilder::startup` (allowlist)
- class: `AnthropicModelProvider` (no key logging; error sanitized), `RuntimeRunner` (no prompt/response persistence)
- migration: `2026_07_04_120000_add_runtime_metrics_to_agent_runs` (+ `2026_07_02_150000`)
- report: `LIVE_RUNTIME_SMOKE_TEST_REPORT.md` (§13)

## Verified Answer
- **PII minimization:** an explicit allowlist — includes `id, name, name_en, website, group, status, sector, product/investment stage, asking_fund_sar, offered_equity_pct, min_ticket_size_sar, total_raised_sar, runway_months`; **excludes** `email, phone_number, applicant_full_name, first_name_*, linkedin_url`.
- **Prompt/response logging:** **not persisted** — the runtime stores no prompt text or raw model output; only token counts/model/latency.
- **Audit metadata stored:** `knowledge_package_id, knowledge_version, knowledge_checksum, knowledge_registry_version, knowledge_loaded_at, source_stale, prompt_contract_version, runtime_provider, runtime_model, runtime_latency_ms, runtime_token_input, runtime_token_output` + `activity_logs` `agent.runtime_run`.
- **Provider data-retention / no-training config:** `NOT VERIFIED IN REPOSITORY` — no such setting exists in code/config.

## Confidence
HIGH (minimization + logging). NOT VERIFIED (provider retention setting).

## Reconciliation Notes
Decision 5.8 fully satisfied except the provider no-training account setting (operational, outside the repo).

## Production Impact
Low-Medium — the no-training/data-retention setting is an Anthropic account configuration to confirm operationally.

---

# Question 10 — 7.9 Environment availability

## Question
Which environments exist; staging; env-specific knowledge; env representation.

## Repository Evidence
- config: `config/agent_knowledge.php` (`environments=[dev,staging,production]`, default `production`); `config/agents.php` (`AI_AGENT_ENV` default `prod`)
- validation: `AgentRunController::decision` (`in:dev,test,prod`)
- report: `ENVIRONMENT_RECONCILIATION_ANALYSIS.md`, `ENVIRONMENT_RECONCILIATION_IMPLEMENTATION_PLAN.md`

## Verified Answer
- **Two env namespaces (code-verified):** Knowledge/Loader/Runtime = `dev|staging|production` (default `production`); Decision/Policy/Agents = `dev|test|prod` (default `prod`).
- **Env-specific knowledge:** yes — `agent_knowledge_cache.environment` column; independent version sequence + active pointer per `(agent_key, environment)`.
- **Env representation:** request param (`?environment=` / `environment` body) + DB column.
- **Physical staging deployment:** `NOT VERIFIED IN REPOSITORY`.

## Confidence
HIGH (modeling). NOT VERIFIED (deployment/staging existence).

## Reconciliation Notes
**CONFLICT (documented):** two vocabularies. It is **fail-safe** (wrong value 422/409, never silent cross-env). Reconciliation plan = adopt `dev/test/prod`, post-integration (`ENVIRONMENT_RECONCILIATION_IMPLEMENTATION_PLAN.md`). Repository is authoritative on the current split.

## Production Impact
Medium — Make must map `production↔prod`, `staging↔test` per call until reconciliation.

---

# Question 11 — 7.10 Secret storage & rotation

## Question
Where secrets live; rotation; which stay in Make; access.

## Repository Evidence
- config: `config/services.php` (`services.anthropic.api_key = env(ANTHROPIC_API_KEY)`), `config/microsoft_bookings.php` (`MS_GRAPH_*`)
- class: `AnthropicModelProvider`, `GraphTokenProvider`

## Verified Answer
- **Storage:** environment variables read via `config()` — `ANTHROPIC_API_KEY`, `MS_GRAPH_*`. No secrets committed in the repo; never returned by APIs; provider errors sanitized to `HTTP <code>`.
- **Rotation / ownership / which stay in Make / access:** `NOT VERIFIED IN REPOSITORY` — operational; no rotation code or secret-manager integration exists.

## Confidence
HIGH (storage mechanism). NOT VERIFIED (rotation/ownership).

## Reconciliation Notes
Decision 5.10 satisfied for "backend secrets in env vars, not Make." The rest is operational.

## Production Impact
Low (mechanism sound); rotation policy is an operational must-define.

---

# Question 12 — 7.11 Rollback / incident process

## Question
Rollback mechanisms; who executes + pause; alert destinations; incident log.

## Repository Evidence
- method: `AgentKnowledgeLoader::activateVersion` (knowledge rollback); `AgentRuntimeService::{pause,pauseAll}` (agent pause); migrations’ `down()`
- routes: `POST admin/agents/{agent}/knowledge/{environment}/{version}/activate`, `agents/{agent}/pause`

## Verified Answer
- **Knowledge rollback:** pointer-only `activateVersion` via the admin route (perm `ai.knowledge.rollback`).
- **Agent pause:** `AgentRuntimeService::pause/pauseAll` (TD-003), enforced in `AgentDecisionService::pauseBlock`.
- **Guard migrations:** all reversible via `down()`.
- **Deploy/runtime rollback, alert destinations, incident owner, incident-log location:** `NOT VERIFIED IN REPOSITORY` (operational).

## Confidence
HIGH (knowledge rollback + pause). NOT VERIFIED (deploy/alerts/incident log).

## Reconciliation Notes
Decision 5.11 ownership is operational; the technical rollback mechanisms exist in code.

## Production Impact
Low technically; the incident runbook is an operational deliverable.

---

# Question 13 — 7.12 Write endpoint verification policy

## Question
Concrete safe-testing rules per endpoint class.

## Repository Evidence
- middleware: `McpAuth:full` on all `/api/v1/ai/*` writes; audit via `AgentAuditLogger` → `activity_logs`
- report: `MAKE_NUMU_ORCHESTRATION_CONTRACT.md` (§2/§3)

## Verified Answer
`NOT VERIFIED IN REPOSITORY` as an encoded policy. The repo **enforces** auth/scope/permission on writes and **audits** them, and idempotency now makes replays safe; but there is no committed "safe test-record" policy artifact. The controlled-testing method **used** in Phase 5 (rolled-back transactions + `test` decision env + live-Claude smoke/E2E) is documented in `LIVE_RUNTIME_SMOKE_TEST_REPORT.md` and `E2E_VALIDATION_REPORT.md`.

## Confidence
HIGH (auth/audit enforcement). NOT VERIFIED (formal written policy).

## Reconciliation Notes
Decision 5.12 is an operational policy; the technical guardrails (auth, audit, idempotency, rollback) are in place.

## Production Impact
Medium — a written safe-testing policy should be authored before broad production writes.

---

# Question 14 — 7.13 MCP vs REST coverage matrix

## Question
Provide the MCP vs REST coverage matrix.

## Repository Evidence
- file: `app/Services/Ai/Mcp/McpToolRegistry.php` (tool names); `routes/api.php`
- report: `REPOSITORY_QUESTION_ANSWERS.md` (§J)

## Verified Answer
The MCP connector exposes **no** Managed-Agents tools (only `investor_*`, `startup_*`, `committee_*`, `demo_*`, `meeting_*`, `note_*`, `tag_*`, `label_option_*`, `activity_log_*`, `notification_*`, `task_*`, `startup_member_*`, `startup_file_*`). Managed-Agents ops are **REST-only** (`McpAuth:full`).

| Area | MCP | REST | Write | Status | Recommended | Auth |
|---|---|---|---|---|---|---|
| AgentRun create/read/list | ❌ | ✅ | ✅ | committed | REST | `McpAuth:full` |
| Decision submit | ❌ | ✅ | ✅ | committed | REST | `McpAuth:full` |
| Approvals read/decide | ❌ | ✅ | ✅ | committed | REST | `McpAuth:full` |
| Handoff create/read/list/supersede | ❌ | ✅ | ✅ | committed | REST | `McpAuth:full` |
| Knowledge read/runtime/versions/meta | ❌ | ✅ | read | committed | REST | `McpAuth:full` |
| Knowledge refresh/activate/rollback | ❌ | ✅ (admin web) | ✅ | committed | Admin UI | `permission:*,admin` |
| Runtime run | ❌ | ✅ | proposal-only | committed | REST | `McpAuth:full` |
| Startup/entity lookup | ✅ | ✅ | read | committed | MCP (read) / REST | `McpAuth:read`/`full` |

Known gap: no `handoff_*`/`agent_run_*`/`knowledge_*` MCP tools (intentional — machine-to-machine REST surface).

## Confidence
HIGH.

## Reconciliation Notes
Decision 5.13 (REST for writes, MCP for read) matches: writes = REST; entity reads available via MCP.

## Production Impact
Low — Make uses REST for the agent flow.

---

# Question 15 — Section 9: Required error responses & audit

## Question
Confirm the deterministic error responses + audit/security requirements.

## Repository Evidence
- class: `AgentApiController::{guard,fail,mapQueryException}`; `HandoffController::validateSourceRun`; `AgentRuntimeRunController`; `McpAuth`
- tests: `HandoffHardeningTest`, `AgentRuntimeRunTest`, `DecisionIdempotencyTest`

## Verified Answer
| Case | Repository behavior |
|---|---|
| Invalid token | ✅ `401` (`McpAuth::reject`) |
| Missing scope | ✅ `403 insufficient_scope` |
| Invalid `source_run_id` | ✅ `422` (`exists` + semantic) |
| Run ∉ startup / not completed / from_agent≠agent_key | ✅ reject (`422`, `validateSourceRun`) |
| Duplicate Handoff | ✅ **idempotent `200`** (existing row) |
| Payload too large | ✅ `413 payload_too_large` (`mapQueryException`) |
| **Invalid `schema_version`** | ❌ **NOT rejected** (free string) |
| **Invalid `retention_class`** | ❌ **NOT rejected** (free string) |
| Supersede already superseded | ✅ `409` |
| Supersede target not found | ✅ `404` |
| Supersede creates branch | ✅ prevented (`handoffs_active_logical_unique`) |
| No active knowledge | ✅ `409 knowledge_unavailable` (fail-closed) |
| AgentRun/startup mismatch (runtime) | ✅ `422 precondition_failed` |
| Invalid model output | ✅ `422 invalid_model_output` (fail-closed after 1 repair) |
| Provider timeout/error | ✅ `502 provider_error` |
| Permission denied (knowledge refresh/activate) | ✅ reject (`permission:*,admin`) |

**Audit/security:** tokens never logged; secrets never returned; handoff create/supersede + knowledge load/rollback + runtime run + decision emit `activity_logs` events with actor + correlation IDs; idempotent replays add no duplicate audit.

## Confidence
HIGH.

## Reconciliation Notes
**CONFLICT:** §9 expects `schema_version`/`retention_class` invalid → reject; the repo does **not** enforce allow-lists (`HANDOFF_GAP_ANALYSIS.md` §4/§5). **Repository is authoritative** — these two rows are NOT met (optional hardening).

## Production Impact
Low-Medium — callers cannot rely on backend rejecting unknown `schema_version`/`retention_class`.

---

# Question 16 — Section 8: Backend implementation status (8.1–8.9)

## Question
Confirm the required backend implementation (Knowledge Cache, Loader, Knowledge APIs, Runtime, Structured Validator, Traceability, Anthropic Adapter, Startup Context, Make Integration).

## Repository Evidence
See the Phase-5 Coverage Matrix below and commits `29f916a`, `a9fa600`, `dde6349`, `8e892f9`, `51d9bb8`.

## Verified Answer
| § | Requirement | Status | Evidence |
|---|---|---|---|
| 8.1 | Knowledge Cache (immutable, active pointer, rollback, refresh audit, checksum no-op, versioned) | ✅ Implemented — **single table** `agent_knowledge_cache` (equivalent naming) | `AgentKnowledgeCache`, migrations `2026_07_02_130000/140000/150000` |
| 8.2 | Registry-driven Loader (ordered, validated, checksummed, no hardcoding) | ✅ | `AgentKnowledgeRegistry`, `AgentKnowledgeLoader`; `AgentKnowledgeCacheTest` |
| 8.3 | Knowledge APIs (metadata/runtime split, immutable versions, permissioned refresh/rollback) | ✅ | `AgentKnowledgeController`; `routes/api.php` + `routes/web.php` |
| 8.4 | Agent Runtime endpoint (fail-closed, propose-only) | ✅ available; **live-verified**, repair path unit-tested | `AgentRuntimeRunController`, `RuntimeRunner`; `29f916a`; `LIVE_RUNTIME_SMOKE_TEST_REPORT.md` |
| 8.5 | Structured output validator (fail-closed + 1 repair, no execution claims) | ✅ | `StructuredOutputValidator`; `AgentRuntimeRunTest` |
| 8.6 | AgentRun traceability (8+ fields) | ✅ all present | `AgentRun`; migrations `2026_07_02_150000` + `2026_07_04_120000` |
| 8.7 | Anthropic adapter (config model/timeout/retry, error mapping, no secret logging) | ✅ | `AnthropicModelProvider`, `config/agent_runtime.php`, `config/services.php` |
| 8.8 | Startup context (core fields, PII-min, no doc parsing in milestone 1) | ✅ (notes/documents deferred) | `RuntimeContextBuilder` |
| 8.9 | Make integration (Runtime-first; Make orchestration-only) | ✅ backend ready; Make wiring **external, not built** | `MAKE_NUMU_ORCHESTRATION_CONTRACT.md`, `E2E_VALIDATION_REPORT.md` |

## Confidence
HIGH.

## Reconciliation Notes
**CONFLICT:** §8.1 lists four tables (`agent_knowledge_packages/_package_files/_active_versions/_refresh_runs`); the repo uses **one** `agent_knowledge_cache` with an ordered `files_json` list. §8.1 explicitly allows "equivalent Numu naming"; repository is authoritative.

## Production Impact
None — all lifecycle/rules are satisfied by the single-table design.

---

# Question 17 — Section 11: Minimal tests before closeout

## Question
Confirm the Loader/Cache, Runtime, Integration, and Evaluation tests.

## Repository Evidence
- tests: `tests/Feature/Agents/{AgentKnowledgeCacheTest,AgentRuntimeRunTest,HandoffHardeningTest,DecisionIdempotencyTest,ManagedAgentsInvariantsTest}.php`
- reports: `LIVE_RUNTIME_SMOKE_TEST_REPORT.md`, `E2E_VALIDATION_REPORT.md`

## Verified Answer
- **Loader/Cache:** ✅ `AgentKnowledgeCacheTest` (27) — 7-file load, ordering, missing/empty/duplicate/traversal/wildcard rejects, keep-active-on-failure, checksum no-op, rollback, metadata vs runtime split.
- **Runtime:** ✅ `AgentRuntimeRunTest` (14) — success, no-active-knowledge (409), ownership/agent-key mismatch, invalid output fail-closed, schema repair, traceability, runtime-only creates no decision/note/handoff/execution, provider→502, token-sum-on-repair, decisions-in-context.
- **Integration:** ✅ controlled E2E (live) run + invariants; `DecisionIdempotencyTest` (6); `HandoffHardeningTest` (17). One controlled dev E2E audit-verified (`E2E_VALIDATION_REPORT.md`); duplicate handoff attempt skipped (200).
- **Evaluation (`TRAINING_EVALUATION_CASES.md` fixture through Runtime):** ❌ **NOT executed** — `NOT VERIFIED IN REPOSITORY` (no `05 Testing` fixture in-repo; requires the SharePoint fixture).

## Confidence
HIGH (Loader/Runtime/Integration). NOT VERIFIED (AUTO-025 evaluation fixture).

## Reconciliation Notes
The AUTO-025 evaluation-fixture run is the one §11 item not yet done (external fixture).

## Production Impact
Low — behavior is validated; the evaluation fixture is a training-parity confirmation, best run before broad rollout.

---

# Definition of Done (Section 13) — verification

| # | DoD item | Status | Evidence |
|---|---|---|---|
| 1 | Real backend repo verified | ✅ | Q1 |
| 2 | Registry-driven loading works | ✅ | `AgentKnowledgeLoader`; `AgentKnowledgeCacheTest` |
| 3 | 7-file `intake_triage` package loads/activates | ✅ | registry JSON; `AgentKnowledgeCacheTest`; E2E |
| 4 | Cache versioned + immutable | ✅ | `AgentKnowledgeCache` guard |
| 5 | Rollback works | ✅ | `AgentKnowledgeLoader::activateVersion` |
| 6 | Active package survives refresh failure | ✅ | `AgentKnowledgeLoader::refresh` (keep-active) |
| 7 | Runtime endpoint exists | ✅ | `AgentRuntimeRunController` |
| 8 | Runtime reads active knowledge | ✅ | `RuntimeRunner` (cache only) |
| 9 | Runtime retrieves startup context | ✅ | `RuntimeContextBuilder` |
| 10 | Runtime calls approved adapter (live) | ✅ | `AnthropicModelProvider`; smoke test |
| 11 | Structured validator works | ✅ | `StructuredOutputValidator` |
| 12 | AgentRun traceability stamped | ✅ | `RuntimeRunner::stampTraceability` |
| 13 | Output cannot bypass pipeline | ✅ | propose-only; only `decide()` executes |
| 14 | Make calls Runtime only as orchestration | ✅ (contract) | `MAKE_NUMU_ORCHESTRATION_CONTRACT.md` |
| 15 | Handoff duplicate behavior deterministic + documented | ✅ | idempotent-200; `HANDOFF_GAP_ANALYSIS.md` |
| 16 | Supersede implemented + verified | ✅ | `HandoffService::supersede`; tests |
| 17 | Production Handoff trigger ownership confirmed | ⚠️ Make-fallback only (Runtime-owned NOT built) | Q7 |
| 18 | Handoff flow dup pre-check OR backend idempotency | ✅ backend idempotency exists | Q6 |
| 19 | Handoff flow verifies read-back | ✅ (E2E `GET latest`) | `E2E_VALIDATION_REPORT.md` |
| 20 | Write endpoint testing policy confirmed | ⚠️ enforcement yes; written policy NOT VERIFIED | Q13 |
| 21 | Official enums + limits confirmed | ✅ (allow-lists absent) | Q3 |
| 22 | MCP vs REST path confirmed | ✅ REST | Q14 |
| 23 | PII/env/secrets/rollback/incident details | ⚠️ partial — provider no-training + rotation + incident log NOT VERIFIED | Q9–Q12 |
| 24 | Handoff id 3 impact + cleanup confirmed | ✅ behavior; ⚠️ specific rows not in repo | Q8 |
| 25 | Canonical `pre_screen` enforced consistently | ❌ backend seeds `prescreen`, unenforced | Q4 |

---

# Final Architecture Validation

**Definitively implemented (repository-proven):** Knowledge Cache (immutable, per-env, rollback, audit) · registry-driven Loader · Knowledge read APIs (metadata/runtime split) · Agent Runtime endpoint (propose-only, live-verified, fail-closed + schema repair) · full AgentRun traceability · Anthropic adapter · PII-minimized context builder · Structured output validator · Handoff P1 + supersede + idempotency (200) · G8 active-policy uniqueness · Decision idempotency (B+D) · approval deduplication (one pending per run) · execution idempotency · controlled live E2E. 79 tests green.

**Deferred (in-repo, non-blocking):** environment reconciliation; `schema_version`/`retention_class` allow-lists; runtime context expansion (notes/documents, full handoff history); a live repaired-output run; runtime activation approval-metadata check.

**External (Make only):** all Make scenarios (trigger → create run → knowledge meta → runtime → decision → handoff), per the orchestration contract — no backend change required.

**Optional future:** production push trigger (poll-only today); least-privilege scopes; expanded read/validation-failure audit; multi-agent rollout (AUTO-027+); AUTO-025 evaluation-fixture run.

# Repository vs Documentation Conflicts

| Conflict | Document | Repository (authoritative) |
|---|---|---|
| Second-stage key | `pre_screen` (5.1) | **seeds `prescreen`**, unenforced on handoff fields |
| Knowledge tables | four tables (8.1) | **single** `agent_knowledge_cache` (+ ordered `files_json`) |
| Package | 4-file / `04 Knowledge` (stale) | **7-file `04 Prompts`** registry package |
| Env vocabulary | one set implied | **two**: `dev/staging/production` (knowledge) vs `dev/test/prod` (decision) |
| `schema_version`/`retention_class` | reject invalid (§9) | **not rejected** (free strings; allow-lists absent) |
| Handoff idempotency | absent (3.2) | **implemented** (idempotent-200) |
| Supersede | deferred (5.5) | **implemented** (Option A) |
| Runtime model default | (unstated) | runtime `claude-sonnet-5`; service `claude-sonnet-4-6` (env); config fallback `claude-opus-4-7` |
| Handoff trigger owner | Runtime-owned default (5.7) | **not built**; Make-fallback is the only wired path |
| `agent_runs.stage` | referenced | **does not exist** (stage identity = `agent_key`) |

# Phase 5 Coverage Matrix

| Request area | Answered by (feature / commit / report) |
|---|---|
| §6 repo verification | `REPOSITORY_QUESTION_ANSWERS.md` |
| 7.1 activation checklist | Loader/Validator (absent) |
| 7.2 routes/enums | routes + model constants; orchestration contract |
| 7.3 pre_screen | `AgentsSeeder.php:40` |
| 7.4 supersede | `HandoffService::supersede` · `a9fa600`/`8e892f9` |
| 7.5 idempotency | Handoff idempotency · `8e892f9` · `HANDOFF_GAP_ANALYSIS.md` |
| 7.6 trigger | `AgentRunService` (no event) · `MAKE_NUMU_ORCHESTRATION_CONTRACT.md` |
| 7.7 id 3 | `HandoffService::latest` · `HANDOFF_GAP_ANALYSIS.md` |
| 7.8 PII/Anthropic | `RuntimeContextBuilder`/`AnthropicModelProvider` · `LIVE_RUNTIME_SMOKE_TEST_REPORT.md` |
| 7.9 environments | `ENVIRONMENT_RECONCILIATION_ANALYSIS.md` |
| 7.10 secrets | `config/services.php` |
| 7.11 rollback | `AgentKnowledgeLoader::activateVersion` / `AgentRuntimeService::pause` |
| 7.12 write policy | `McpAuth` + audit (policy doc absent) |
| 7.13 MCP vs REST | `McpToolRegistry` · `REPOSITORY_QUESTION_ANSWERS.md` §J |
| §8 implementation | Q16 matrix |
| §9 errors | Q15 · `AgentApiController::mapQueryException` |
| §11 tests | Q17 · 79-test suite |
| decision idempotency | `AgentDecisionService` B+D · `51d9bb8` · `DECISION_IDEMPOTENCY_ANALYSIS.md` |
| G8 policy uniqueness | `dde6349` · `G8_INTAKE_TRIAGE_POLICY_ANALYSIS.md` |
| E2E | `E2E_VALIDATION_REPORT.md` |

# Open Questions (cannot be proven from the repository)

1. Whether this tree deploys to `dashboard.numuangels.net` (no deploy config).
2. Whether a **staging** environment physically exists.
3. Anthropic **data-retention / no-training** account settings.
4. **Secret rotation** cadence + per-secret access ownership.
5. **Alert destinations, incident owner, incident-log location.**
6. Production `handoffs` rows `id 2`/`id 3` (startup 551) — production/Data-Store data, not in this repo.
7. Whether the production `handoffs`/`agent_action_policies`/`agent_approval_requests` tables have pre-existing active duplicates (must be checked before the guard migrations apply there).
8. Authoritative **startup group/status/action** value lists (DB-driven, not code enums).
9. AUTO-025 **evaluation-fixture** runtime results (external `TRAINING_EVALUATION_CASES.md` / `05 Testing`).
10. Whether the product owner has confirmed `move_to_review = always_allow` for live `prod` decisions (G8).
