# Scheduling on cPanel / shared hosting — minimum setup

The system aims for **zero manual cPanel cron entries**, but PHP cannot
launch persistent processes from inside itself on shared hosting. We have
to compromise on **exactly one cron entry**, and in exchange:

* Everything else (sync flush, reap, snapshot warm, daily reconciliation)
  is auto-driven by Laravel's scheduler.
* If even that one entry silently breaks (a known shared-hosting failure
  mode), the **`SchedulerFallback` middleware** catches it: the next
  admin request runs `schedule:run` inline (after the response is sent)
  to keep the system caught up.
* The dashboard's **System Health panel** at `/sync/v2` shows the live
  status of every component. If anything goes red, the panel tells you
  what's stale and how stale.

---

## The minimum cron entry (recommended)

Add **one** line to cPanel → Cron Jobs (replace `USER` and PHP path with
the values from `which php` on your server):

```
* * * * * cd /home/USER/numu_angels && /usr/local/bin/ea-php82 artisan schedule:run >> /dev/null 2>&1
```

That's it. `schedule:run` then drives:

| Schedule | Task | Why |
|---|---|---|
| every minute | `scheduler-heartbeat` (closure) | Heartbeats `system_health` so the dashboard knows the scheduler is alive |
| every 30s    | `sync:flush`                  | Mirrors Redis/cache counters → DB row |
| every 1m     | `sync:reap`                   | Reclaims pages whose worker died |
| every 5m     | `sync:snapshot:warm`          | Refreshes `/sync/snapshot` counts cache |
| daily 04:30  | `sync:reconcile`              | Soft-deletes Monday-removed records |

You can confirm by visiting `/sync/v2` — the **System Health** panel will
show all ticks within their expected windows.

---

## Queue workers — also need cron entries

Workers can't be persistent processes either. The pattern: per-minute cron
that runs the worker for ~55s with `--stop-when-empty`, then exits and waits
for the next cron tick.

```
* * * * * cd /home/USER/numu_angels && /usr/local/bin/ea-php82 artisan queue:work --queue=sync-control --stop-when-empty --max-time=55 --tries=3 --memory=128 >> /home/USER/numu_angels/storage/logs/queue-control.log 2>&1
* * * * * cd /home/USER/numu_angels && /usr/local/bin/ea-php82 artisan queue:work --queue=sync-fetch   --stop-when-empty --max-time=55 --tries=4 --memory=256 >> /home/USER/numu_angels/storage/logs/queue-fetch.log   2>&1
* * * * * cd /home/USER/numu_angels && /usr/local/bin/ea-php82 artisan queue:work --queue=sync-process --stop-when-empty --max-time=55 --tries=1 --memory=192 >> /home/USER/numu_angels/storage/logs/queue-process.log 2>&1
* * * * * cd /home/USER/numu_angels && /usr/local/bin/ea-php82 artisan queue:work --queue=sync-files   --stop-when-empty --max-time=55 --tries=3 --memory=192 >> /home/USER/numu_angels/storage/logs/queue-files.log   2>&1
```

If your host limits cron entries: collapse the four lines into one
`--queue=sync-control,sync-fetch,sync-process,sync-files` worker — slower
but functional.

If your host allows persistent processes (Supervisor / systemd / etc),
prefer `queue:work --tries=4 --timeout=240 --memory=256` as a long-running
service. No `--stop-when-empty`.

---

## Notifications queue + retry sweeper (scheduler-driven — no extra cron)

The notification system fans every automation trigger out into **one
job per recipient** on a dedicated `notifications` queue. Both the
worker and the delayed-retry sweeper are registered in
`routes/console.php`, so the single `schedule:run` cron above already
drives them — **no additional cron entry is required**:

| Schedule | Task | Why |
|---|---|---|
| every minute | `queue:work --queue=notifications --stop-when-empty --max-time=55 --timeout=60 --tries=3` | Drains per-recipient send jobs, isolated from sync/ai queues |
| every 5m     | `notifications:retry-failed` | Walks the 5m → 30m → 2h retry ladder (max 3) for failed/undelivered rows, off the send path |

**Critical invariant — `--timeout (60s) < retry_after (90s)`.** The
per-recipient `DispatchAutomationJob` sets `$timeout = 60`, kept below
the database connection's `retry_after` (90s, `config/queue.php`). A
stuck job is therefore killed before the queue re-releases it, so the
same recipient is never processed twice concurrently. Never raise the
job timeout to ≥ 90 without also raising `retry_after`.

If you prefer an explicit cron entry (e.g. to give notifications its own
log file or more frequent ticks than the scheduler), add:

```
* * * * * cd /home/USER/numu_angels && /usr/local/bin/ea-php82 artisan queue:work --queue=notifications --stop-when-empty --max-time=55 --timeout=60 --tries=3 --memory=192 >> /home/USER/numu_angels/storage/logs/queue-notifications.log 2>&1
```

Observability: `php artisan notifications:report` (all triggers) or
`--automation=ID` (one automation: recipients, per-channel delivery,
retry posture, success rate, dispatch rollup).

---

## Self-healing fallback (no extra setup needed)

`SchedulerFallback` middleware runs on every admin route. After the
response is sent, it checks `system_health[scheduler].last_seen_at`. If
> 3 minutes old (3× the schedule interval), it runs `Artisan::call('schedule:run')`
inline, protected by a Cache lock so concurrent admins don't double-fire.

Effect: even if your cron entry breaks, the system catches itself up the
next time anyone visits the admin panel. The cron remains the primary
mechanism (it fires when nobody is logged in), but you won't get silently
stuck for hours either way.

---

## Diagnostics

| Question | How to answer |
|---|---|
| Is the scheduler firing? | `GET /admin/sync/health` — check `components[scheduler].is_stale` |
| When did sync:flush last run? | Same JSON, `components[task:sync:flush].age_seconds` |
| Did my cron entry succeed? | Tail `/home/USER/numu_angels/storage/logs/queue-*.log` |
| Is a queue stuck? | `components[queue:sync-fetch].is_stale = true` means no jobs ran in 10+ min |

For external monitoring (Pingdom, UptimeRobot, etc.) point at:
```
GET https://yourdomain.com/admin/sync/health
```
The endpoint returns `200 OK` with `{ "status": "healthy" | "degraded" }`
and a JSON breakdown. Set the alert to "alert when status != healthy".
