# Agent Runtime — Live Claude Smoke Test Report

> Permanent repository artifact recording the first controlled live validation of
> the Agent Runtime (`POST /api/v1/ai/agent-runtime/run`) against a real Anthropic
> model, using the production runtime path and real Knowledge Cache data.

**Date:** 2026-07-05 · **Branch:** `feat/agent-knowledge-cache` · **Runtime commit under test:** `29f916a`
**Executed by:** a controlled dev harness inside a single DB transaction that was **rolled back** (0 rows persisted). The Anthropic call was real. The API key was never printed or logged.

```text
The Agent Runtime has now been validated against a real Anthropic model
using the production runtime path and real Knowledge Cache data,
while preserving the propose-only architecture.
```

---

## 1. Runtime flow exercised

The exact production path (`App\Services\Agents\Runtime\RuntimeRunner::run()`):

```
validate request (agent_key, startup_id, agent_run_id, environment)
 → load AgentRun + fail-closed preconditions (ownership / agent_key / non-terminal)
 → load Startup
 → AgentKnowledgeReader::active(agent_key, environment)      [DB Knowledge Cache ONLY]
 → RuntimeContextBuilder::build()                            [reuse existing models, PII-min]
 → RuntimePromptBuilder::build()                             [active package ordered files]
 → RuntimeModelProvider::complete()  →  AnthropicModelProvider  →  Claude Messages API (LIVE)
 → StructuredOutputValidator::validate()                    [fail-closed; 1 repair if needed]
 → stamp traceability on the AgentRun                        [metadata columns only]
 → return PROPOSAL (executed:false)
 ── STOP ── proposal submitted separately to POST /agent-runs/{run}/decision
```

Provider resolution was the real one: `config('agent_runtime.provider') = anthropic` → `AnthropicModelProvider`, reusing `config('services.anthropic.*')` credentials.

## 2. Request payload

```json
{ "agent_key": "intake_triage", "startup_id": 479, "agent_run_id": 36, "environment": "production" }
```
(`startup_id`/`agent_run_id` were transient rows created inside the rolled-back transaction.)

## 3. Knowledge source verification (Cache only, never SharePoint)

- A knowledge package was written to the **`agent_knowledge_cache` table** (DB) with a unique sentinel `SMOKE-SENTINEL-KNOWLEDGE-7Q42` embedded in a file body.
- The assembled prompt **contained the sentinel** → `knowledge_sentinel_in_prompt = YES`, proving the content came from the DB cache.
- Code path confirms exclusivity: `RuntimeRunner` reads knowledge only via `AgentKnowledgeReader::active()` (a query on `agent_knowledge_cache`). It never constructs or calls `KnowledgeSource` / `LocalFolderSource` / `GraphSharePointSource`. **No SharePoint/Graph request is possible on the runtime path.**

## 4. Prompt statistics

| Metric | Value |
|---|---|
| System block | 2130 bytes |
| User block | 1314 bytes |
| Total prompt | 3444 bytes |
| Package files (in load order) | 2 |

## 5. Model, token usage, latency

| Metric | Value |
|---|---|
| Model | **`claude-sonnet-5`** (available + callable on the account) |
| Provider | `anthropic` |
| Input tokens | 1375 |
| Output tokens | 435 |
| Runtime latency | 7767 ms (provider call) |
| Wall-clock | 8321 ms |
| Provider calls | 1 |

## 6. Structured validation result

- Claude returned valid JSON on the **first** attempt → validation passed with **no repair** (`output_repaired = no`, `runtime_calls = 1`).
- Validated proposal:

| Field | Value |
|---|---|
| `proposed_action` | `stop` (∈ allowed `[move_to_review, reject, stop]`) |
| `confidence` | `0.72` |
| `reasoning_summary` | "Key fields (sector, status, product/investment stage) are missing … impossible to confidently assess sector fit…" |
| `analysis` | "The startup record … is largely incomplete: sector, status, product_stage, investment_stage … are all null" |
| `proposed_note` | present |
| `proposed_handoff` | null |
| `executed` | false |

The model correctly picked an **allowed** action and gave a sensible triage rationale for the deliberately sparse fixture.

## 7. Repair path result

Not triggered on this run (first output valid). The repair path is unit-tested separately (`AgentRuntimeRunTest::test_invalid_output_is_repaired_then_accepted` and `…twice_fails_closed_422`) and, per the token-accounting fix, sums usage across both calls.

## 8. Traceability fields stamped on the AgentRun

| Field | Value |
|---|---|
| `knowledge_package_id` | 38 |
| `knowledge_version` | v1 |
| `knowledge_checksum` | smoke-bundle-checksum |
| `knowledge_registry_version` | 1.0 |
| `knowledge_loaded_at` | 2026-07-05 02:11:19 UTC |
| `source_stale` | false |
| `prompt_contract_version` | v1 |
| `runtime_provider` | anthropic |
| `runtime_model` | claude-sonnet-5 |
| `runtime_latency_ms` | 7767 |
| `runtime_token_input` | 1375 |
| `runtime_token_output` | 435 |

## 9. Audit events

The runtime step wrote exactly one `activity_logs` row: **`agent.runtime_run`**.

## 10. Database mutations

- **Runtime step:** the only columns changed on `agent_runs` were the **12 traceability fields + `updated_at`**. No decision/execution/analysis columns changed.
- Side-effect counters after the runtime step: **approvals 0, handoffs 0, execution_requests 0**; `analysis_status` still `queued`; `proposed_action` still `null`.
- **Net persisted:** none (transaction rolled back).

## 11. Pipeline submission result

The proposal was submitted through the **existing** decision entry (`AgentDecisionService::decide()`), in the **`test`** environment (chosen deliberately — see §13):

| Metric | Value |
|---|---|
| Decision environment | `test` |
| Resolved policy | `blocked` (source `default`) |
| Decision | `blocked` |
| Executed | false |
| Run after decision | `analysis_status = completed`, `proposed_action = stop`, `execution_status = blocked` |
| Approvals / executions created | 0 / 0 |

The proposal flowed through the real policy-resolution → block branch with no execution, confirming the propose-only → decision-endpoint handoff works end-to-end.

## 12. Rollback guarantees

All fixtures (admin actor reuse, temp startup, temp AgentRun, temp active knowledge package) and all writes (runtime stamp + decision block record) were created inside `DB::beginTransaction()` and reverted with `DB::rollBack()`. **Zero rows persisted.** The only non-transactional effect was the real (billed) Anthropic API call, which was the intended validation.

## 13. Security observations

- **API key**: read from `config('services.anthropic.api_key')`; never printed, logged, or echoed. Provider errors are sanitized to `HTTP <code>` (no key, no body).
- **Safe decision environment**: `intake_triage` has seeded **`prod`** policies including `move_to_review → always_allow`. To avoid any native action / notification / external side effect (which a DB rollback cannot undo), the decision was submitted in **`test`** (no seeded policy → `blocked`), with an explicit pre-resolve guard that would have skipped `decide()` had it resolved to `always_allow`.
- **Queue safety**: queue driver is `database`; any dispatched job would have been inserted into `jobs` inside the transaction and rolled back (never run). The block branch dispatches nothing anyway.
- **PII**: the context builder's allowlist excludes email/phone/applicant name/social URLs; no PII was sent to the model.
- **No prompt/response persistence**: the runtime stores only token counts/model/latency; the prompt and raw model text are not written to the DB.

## 14. Production-readiness conclusions

**Validated (now closed):**
- `claude-sonnet-5` is available and callable on the account.
- The production runtime path loads knowledge exclusively from the Knowledge Cache.
- Prompt assembly, live structured-output validation, and traceability stamping work against a real model.
- The runtime performs exactly one AgentRun traceability write and mutates no decision/approval/execution/startup/handoff/workflow state.
- The proposal submits cleanly into the existing `AgentDecisionService::decide()` pipeline.

**Still open (not blockers for a controlled rollout; tracked in the Phase 5 plan):**
- A production trigger (Make orchestration) to call the runtime then the decision endpoint (no automated trigger exists in-repo).
- Environment-namespace reconciliation (`production` for knowledge/runtime vs `prod` for decision/policy) — fail-safe today, clarity improvement.
- The seeded `intake_triage` policy set contains duplicate/conflicting `move_to_review` rows (`always_allow` and `needs_approval`) in `prod` — an operational data-cleanup item, independent of the runtime.
- Repair path and multi-turn cost behaviour verified by unit tests, not yet by a live repaired run.

**Conclusion:** the Agent Runtime is functionally production-ready on its own path; remaining items are integration/orchestration and data-hygiene, not runtime correctness.
