# Managed Agents — Incident Register & Minimum Incident Plan

> Operational runbook + living incident log for the Managed Agents runtime
> (runtime / decision / handoff). This is the **human incident record** — it is
> deliberately separate from `activity_logs` (which auto-records *successful*
> changes). Every runtime/decision/handoff failure handled during a test or in
> production gets **one row** in the register below.
>
> **Scope note:** this is documentation only. It changes no runtime, decision,
> approval, execution, or handoff behaviour. No auto-alerting is built by this
> file — see "Known gaps".

---

## 1. Confirmed minimum plan (the three decisions)

| Decision | Confirmed value | Meaning |
|---|---|---|
| **Alert destination** | **In-platform (pull-based)** | Failures are seen *inside the dashboard* — there is **no auto-push** (no Teams/Slack/email). The on-call watcher keeps the Managed-Agents console open during the test window. |
| **Incident owner** | **Platform-admin role** | Any holder of the `agents.manage` permission is the incident owner. For each test window, **one** such person is named as the active watcher. |
| **Incident log** | **This file** (`docs/managed-agents/INCIDENT_REGISTER.md`) | Every incident is logged as a row in §6. |

**Do not deploy to production or wire live Make production scenarios until Fahad approves rollout.**

---

## 2. Where a failure surfaces (technical reality)

| Failure type | Where it appears | Persisted in DB? | Auto-alert? |
|---|---|---|---|
| **Runtime** (`POST /api/v1/ai/agent-runtime/run`) | HTTP error to Make (422 / 409 / 502) + `storage/logs/laravel.log` | No (runtime saves traceability only on success) | ❌ |
| **Decision** (`POST /api/v1/ai/agent-runs/{run}/decision`) | HTTP error + `agent_runs.error_code` / `error_message` + run detail page | Partial (error on the run; successful events in `activity_logs`) | ❌ |
| **Handoff** | HTTP 409 / error + `activity_logs` on success | Yes (successful events) | ❌ |
| **Any successful change** | `activity_logs` (entity activity pages) | Yes | ❌ |
| **Make (orchestrator)** | Make scenario run history | — | Make can self-notify |

Because nothing pushes, **monitoring is pull**: someone must look.

---

## 3. Where to LOOK (in-platform surfaces)

- **Managed Agents console:** `admin/agents` (index) → per-agent `admin/agents/{agent}` → **run detail** `admin/agents/runs/{run}`.
- **Approvals console:** `admin/approvals` (pending items that need a human decision).
- **Activity / audit:** the per-entity activity pages (`.../activity/{kind}/{logId}`, `kind = activity|notification|audit`) backed by `activity_logs`.
- **Server logs:** `storage/logs/laravel.log` (exceptions, runtime provider errors).
- **Make:** the scenario's own run history.

All of the above require **`agents.manage`** (admin guard) for the agent console.

---

## 4. How to STOP the runtime (manual kill switch)

Owner = any `agents.manage` holder. **Test this button before the test starts.**

- **Pause everything:** `admin/agents` → **Pause-all** button (API mirror: `POST /api/v1/ai/agents/pause-all`).
- **Pause one agent:** `admin/agents/{agent}/pause` (API: `POST /api/v1/ai/agents/{agentKey}/pause`).
- **Re-enable:** `admin/agents/enable-all` or `admin/agents/{agent}/enable`.

Pausing sets `agent_runtime_config` status; the runtime refuses new work for a paused agent.

---

## 5. How to ROLL BACK a bad change

- Every AI write is audited with **before/after values**, so it is reversible.
- Connector rollback capability exists at **single / session / task / date-range** scope (permission `ai.rollback`).
- Notes are **never deleted** — an edit restores the prior body; a create is append-only.

---

## 6. Severity guide

| Sev | Definition | First action |
|---|---|---|
| **S1** | Wrong state change reached a **real** startup/investor, or repeated auto-execution | **Pause-all immediately**, then roll back, then log |
| **S2** | Failure blocks the flow but no bad write (502/422/409 loops) | Pause the affected agent, retry once, log |
| **S3** | Cosmetic / transient / single non-repeating error | Log only; monitor |

---

## 7. How to log an incident (procedure)

When a failure is handled:
1. **Contain first** (Pause-all / per-agent pause if S1/S2), then **roll back** if a bad write happened.
2. Add **one row** to §8 with: timestamp, run id, type, severity, what happened, action taken, owner, status.
3. If it needs follow-up code/config work, open an Azure DevOps work item and paste its link in the **Ref** column.
4. Close the row when resolved (status = `closed`).

---

## 8. Incident register (living log)

> Newest first. Keep one row per incident. `run id` = the `agent_runs.id`.
> Example row is illustrative — delete it once the first real row is added.

| # | Date/Time (Asia/Riyadh) | Run id | Type | Sev | What happened | Action taken | Owner | Status | Ref |
|---|---|---|---|---|---|---|---|---|---|
| _ex_ | 2026-07-06 14:20 | 312 | runtime | S2 | Provider returned 502 (timeout) on `intake_triage` | Paused `intake_triage`, retried once — succeeded | platform-admin (watcher) | closed | — |
|   |   |   |   |   |   |   |   |   |   |

---

## 9. Known gaps / limitations (be honest before the test)

1. **No auto-alert.** Nothing pushes on failure — an "in-platform alert" today means *the watcher is actively looking at the console*. A real in-platform failures widget/counter is a **small additive build**, to be done only **after Fahad approves**.
2. **Runtime failures aren't persisted** — a failed `agent-runtime/run` shows only as an HTTP error to Make + a `laravel.log` line. The watcher must also check Make's run history.
3. **Metering/alerting** on loader failures, stale knowledge, validation failures, latency/cost is **not built** (flagged non-blocking in `FINAL_PHASE_5_AUDIT.md`).

---

## 10. Minimum checklist — before a controlled-prod test

- [ ] Watcher named (a `agents.manage` holder) and console open for the window.
- [ ] Pause-all button tested and working for that account.
- [ ] Test uses **marked test records only** + a temporary `needs_approval` policy override (nothing auto-executes).
- [ ] This register is open and the watcher knows to add a row on any failure.
- [ ] Rollback path (audit / `ai.rollback`) confirmed available.

*Once all five are checked, the minimum monitoring/incident bar for a controlled-prod test is met.*
