master-zapier-plan-draft/packages/overview/04-uptime.md
George Lambert a32c91475c
Some checks are pending
offline / test (push) Waiting to run
Add overview repo and verae-nats-process expansion template
High-level system map with diagrams, TOC, and a docs index. Template
worker shows how to add a new verae.* address for search, storage, or
unplanned functions without teaching Zapier NATS.
2026-09-11 13:50:37 -04:00

1.6 KiB

4. How the system maintains uptime

Uptime is NATS durability + fleet replica floors + more than one machine, not a single always-on Zapier connection.

Replica floors (verae-fleet)

packages/verae-fleet/fleet.json sets min / max / keepFloor per service. Available means running, healthy, and not paused.

Service Default min keepFloor
tree-node 3 yes
archive-worm 3 yes
archive-aggregator, job-poller, webhook-deliver 1 yes
zappier-edge, middleware-http 1 yes

If a tree node is paused, crashes, or fails /health, fleet starts another copy until three are available. Monitor: http://127.0.0.1:3850/ (green / yellow / red rows).

Restart and pause

  • Unhealthy /health → same instance id restarted.
  • Pause does not count toward min.
  • stop on a service disables keepFloor for that service.

Spread across machines

machines.json lists hosts (local, ns1 = marchon@70.88.205.138 with ~/.ssh/id_ed25519, optional lan-134). New replicas go to the least-loaded eligible host. Remote spawn/health/kill is SSH; workers bind loopback on the remote box.

JetStream

Work queues (jobs.watch, webhooks.deliver) replay if a consumer dies. Event stream (jobs.events) lets waiters and webhook routers catch up. A 3-node cluster keeps the stream if one NATS server is down.

What Zapier sees

HTTPS 202 jobId, wait JSON, or REST Hook. Timeouts return pending + jobId so the Timestamp Completed trigger can finish the job. Zapier retries are safe: the same SHA-256 returns the original seal (existing: true).