# Fleet: replica floors, monitor, pause, restart Runtime copies of middleware workers, WORM archives, and **tree nodes** are declared in `packages/verae-fleet/fleet.json` (how many) and `packages/verae-fleet/services/*.json` (how each one runs). ```text operator --HTTP 127.0.0.1:3850--> fleet control | spawn workers with /health tree-node min=3 keepFloor | archive-worm min=3 | paused ≠ available | unhealthy → restart v replicas on loopback health ports ``` - **list** — `node src/cli.js list` - **monitor** — `serve` probes `/health` and reconciles - **restart** — crash or 503 → same instance id respawned - **pause / off** — pause does not count toward `min`; keepFloor starts another tree-node. `stop` on a service disables it. - **tree-node floor** — `fleet.json` `tree-node.min` (default 3). Do not drop this without changing the spec; bulk-summary leaf queries need several bloom-filtered nodes. ## Extra machines Define hosts in `packages/verae-fleet/machines.json` (or **Add machine** on the monitor). Each machine has `capacity` and `roles` it may run. The supervisor places new replicas on the **least-loaded** eligible host. - `kind: local` — spawn on the control plane. - `kind: agent` — HTTP to `http://host:agentPort` (`node src/agent.js` on that box). Do not publish NATS. Raising `tree-node.max` and adding machines increases bulk-summary lookup capacity. ## Message RTT (planning) The monitor samples `POST /message` on each healthy replica and shows **min / avg / p50 / p90** milliseconds of message processing (service row, instance, and machine). Use p90 when sizing extra tree nodes. Zapier cloud is not spawned. NATS is monitored, not bound publicly.