master-zapier-plan-draft/packages/verae-fleet
George Lambert a694f246fb
Some checks are pending
offline / test (push) Waiting to run
Detach SSH-spawned workers so remote start returns instead of hanging
Remote node is launched in a bash -c background subshell. Live check
on NS1: spawn, /health, kill as marchon with ~/.ssh/id_ed25519.
2026-09-11 13:37:08 -04:00
..
public Control remote fleet workers over SSH with user, host, and key path 2026-09-11 13:31:21 -04:00
services Add fleet control: service configs, replica floors, monitor, pause/restart 2026-09-11 13:09:55 -04:00
src Detach SSH-spawned workers so remote start returns instead of hanging 2026-09-11 13:37:08 -04:00
test Control remote fleet workers over SSH with user, host, and key path 2026-09-11 13:31:21 -04:00
.gitignore Control remote fleet workers over SSH with user, host, and key path 2026-09-11 13:31:21 -04:00
fleet.json Add fleet control: service configs, replica floors, monitor, pause/restart 2026-09-11 13:09:55 -04:00
machines.json Control remote fleet workers over SSH with user, host, and key path 2026-09-11 13:31:21 -04:00
machines.secrets.json.example Control remote fleet workers over SSH with user, host, and key path 2026-09-11 13:31:21 -04:00
NATS.md Add fleet control: service configs, replica floors, monitor, pause/restart 2026-09-11 13:09:55 -04:00
package.json Add fleet control: service configs, replica floors, monitor, pause/restart 2026-09-11 13:09:55 -04:00
README.md Control remote fleet workers over SSH with user, host, and key path 2026-09-11 13:31:21 -04:00
SERVICES.md Spread fleet replicas across extra machines and show message RTT percentiles 2026-09-11 13:25:05 -04:00
SUMMARY.md Add fleet control: service configs, replica floors, monitor, pause/restart 2026-09-11 13:09:55 -04:00

verae-fleet

Operator control plane for Verae Time × Zapier runtime services:

  • a list of every service and its config file
  • a central replica spec (fleet.json) — how many copies must be up
  • a monitor of active / paused / unhealthy replicas
  • restart when a copy is offline or fails /health
  • on / off / pause / resume per service or per instance
  • keepFloor on tree nodes so at least min copies stay available (paused copies do not count)
cd packages/verae-fleet
npm test
node src/cli.js list
node src/cli.js serve    # http://127.0.0.1:3850/

Against a running daemon:

node src/cli.js status
node src/cli.js pause tree-node-0      # floor starts another tree-node
node src/cli.js resume tree-node-0
node src/cli.js restart tree-node-1
node src/cli.js stop webhook-deliver   # disable that service
node src/cli.js start webhook-deliver

Add capacity in machines.json or the monitor Add machine form (kind=ssh, user, host, identity file path). New replicas land on the least-loaded eligible host.

node src/cli.js ssh-check ns1    # marchon@70.88.205.138 with ~/.ssh/id_ed25519

Private keys stay on disk (~/.ssh/id_ed25519); git stores only the path. Optional overrides: machines.secrets.json (gitignored).

The UI shows message-processing min / avg / p50 / p90 RTT per service, instance, and machine.

Zapier cloud apps are listed but not spawned. NATS on NS1 is monitored only (loopback :4222, never a public bind).

Clone: ssh://git@git.georgelambert.org:2223/marchon/verae-fleet.git