master-zapier-plan-draft/packages/verae-keep/README.md
George Lambert 7a5e25639e
Some checks are pending
offline / test (push) Waiting to run
Add Proxmox worker CT, off-box watch, backup, and tagged deploy.
NS1 is the Proxmox host. verae-proxmox creates LXC 510 (verae-px-worker
10.10.10.20 on vmbr1) with a private NATS proxy on 10.10.10.1:4222.
verae-uptime GET-watches public doors; verae-backup snapshots SQLite and
worm/tree data; verae-deploy does host-deps + checkout + npm ci.
Fleet overlays/ns1 are checked in (start.sh no longer rewrites JSON).
User systemd + linger for keep and fleet survive reboot.
2026-09-11 23:35:44 -04:00

1.4 KiB
Raw Blame History

verae-keep

Host keep-alive for Verae workers. Restarts a unit that crashed unless the operator console paused or stopped it.

Forgejo: https://git.georgelambert.org/marchon/verae-keep

Two processes:

Process Port Job
node src/keep.js :3860 Probe units, spawn if intent is run
node src/watch.js :3861 Restart keep if /health fails
scripts/guard.sh Restart watch if watch exits

Intent is run | pause | stop. Sources, newest wins:

  1. POST /intent/:id (admin / API)
  2. ~/verae-fleet-runtime/data/<instance>/intent.json (fleet writes this on start/pause/stop)
  3. live FLEET_URL/api/status when set
  4. unit defaultIntent (run)

Pause: leave the process down or POST /pause. Stop: SIGTERM and do not spawn. Run: spawn if health fails.

NS1 (70.88.205.138)

KEEP_UNITS=$HOME/verae-keep/units.ns1.json KEEP_STATE=$HOME/verae-keep/data \
  nohup bash scripts/guard.sh >/tmp/verae-keep-guard.out 2>&1 &

Or user systemd (survives reboot; do not also nohup guard.sh):

bash scripts/install-systemd.sh
systemctl --user start verae-keep-guard

Proxmox guest extra copies: KEEP_UNITS=units.px-worker.json.

Units: NATS (observe only), job-poller, webhook-deliver, archive-aggregator, archive-worm ×3, tree-node ×3.

API

  • GET /health GET /status
  • POST /intent/:id {"state":"pause|stop|run"}
  • POST /reconcile