# Controlled End-to-End Validation Report (repository-based)

> Permanent record of the controlled end-to-end validation of the full
> Make ↔ Numu orchestration chain, driven through the real controllers with one
> live Anthropic call. Non-destructive: all DB work ran inside a transaction that
> was **rolled back** (0 rows persisted). The API key was never printed.

**Date:** 2026-07-05 · **Branch:** `feat/agent-knowledge-cache` (post `51d9bb8`).

```text
Backend is sufficient for Make orchestration with NO additional backend code.
Every hop of the documented chain works and is idempotent, validated end-to-end
against a real Anthropic model while preserving the propose-only architecture.
```

## Backend sufficiency (proven from the repository)

`php artisan route:list` confirms every endpoint the orchestration contract
requires is registered and wired (all behind `McpAuth:full`):

| Step | Route | Handler |
|---|---|---|
| Create run | `POST /api/v1/ai/agent-runs` | `AgentRunController@store` |
| State check | `GET /api/v1/ai/agent-runs/{run}` | `@show` |
| Knowledge meta | `GET …/agents/{agentKey}/knowledge/meta` | `AgentKnowledgeController@meta` |
| Runtime | `POST /api/v1/ai/agent-runtime/run` | `AgentRuntimeRunController@run` |
| Decision | `POST /api/v1/ai/agent-runs/{run}/decision` | `@decision` |
| Handoff | `POST /api/v1/ai/handoffs` (+ `/latest`, `/{id}/supersede`) | `HandoffController` |
| Approvals | `GET/POST /api/v1/ai/approvals/{id}/{decide,approve,reject}` | `ApprovalApiController` |

No new route, controller, service, or migration is needed for Make to orchestrate.

## Execution method (controlled, non-destructive)

Driven through the real controllers inside a single `DB::beginTransaction()` …
`DB::rollBack()`. The decision was submitted in the **`test`** environment (no
seeded policy → deterministic `blocked`), so **no native action, notification, or
external side effect** could fire. One live `claude-sonnet-5` call was made for
the runtime step.

## Results — every hop validated

| Step | Result |
|---|---|
| 1. Create run | ✅ `201`; replay with the same `idempotency_key` → **same run** (idempotent) |
| 2. Knowledge meta | ✅ `version=v1`, `checksum=e2e-checksum` |
| 3. Runtime (LIVE) | ✅ `200`; `model=claude-sonnet-5`, `action=stop`, `confidence=0.75`, `runtime_calls=1`, **`executed:false`** |
| 4. Decision (test env) | ✅ `decision=blocked`, `executed=false`; **replay → same outcome** (`idempotent_replay=true`) |
| 5. Handoff | ✅ `201` create; **replay (different body) → `200` same row** (idempotent) |
| 6. `GET /handoffs/latest` | ✅ resolves to the created handoff |
| Side-effect audit | ✅ **0 approvals**, **0 execution requests**, **exactly 1 handoff** |
| Persistence | ✅ rolled back — **0 rows persisted** |

## Idempotency proven at every hop
- **Run create** — `idempotency_key` → replay returns the same run.
- **Runtime** — side-effect-free; safe to retry.
- **Decision** — `agent_run_id` one-decision-per-run boundary (Options B + D); replay returns the same outcome, no re-branch, no re-execution.
- **Handoff** — logical-tuple dedup; replay returns the existing row (200).

## Conclusion
The full chain — create run → knowledge meta → runtime (live model) → decision →
handoff — operates correctly and idempotently end-to-end, with the propose-only
architecture intact (`executed:false`; submission only via the decision endpoint;
no execution triggered in the safe test path). **The backend is end-to-end ready
for Make orchestration under at-least-once delivery.** The remaining work is
external Make scenario wiring (no backend change) and, optionally, the
environment reconciliation (see `ENVIRONMENT_RECONCILIATION_IMPLEMENTATION_PLAN.md`).
