Add Proxmox worker CT, off-box watch, backup, and tagged deploy.
Some checks are pending
offline / test (push) Waiting to run

NS1 is the Proxmox host. verae-proxmox creates LXC 510 (verae-px-worker
10.10.10.20 on vmbr1) with a private NATS proxy on 10.10.10.1:4222.
verae-uptime GET-watches public doors; verae-backup snapshots SQLite and
worm/tree data; verae-deploy does host-deps + checkout + npm ci.
Fleet overlays/ns1 are checked in (start.sh no longer rewrites JSON).
User systemd + linger for keep and fleet survive reboot.
This commit is contained in:
George Lambert 2026-09-11 23:35:44 -04:00
parent 2d51d7a0dd
commit 7a5e25639e
69 changed files with 1087 additions and 118 deletions

View file

@ -23,7 +23,11 @@ If a tree node is paused, crashes, or fails `/health`, fleet **starts another co
## Spread across machines
`machines.json` lists hosts (`local`, `ns1` = `marchon@70.88.205.138` with `~/.ssh/id_ed25519`, optional `lan-134`). New replicas go to the **least-loaded** eligible host. Remote spawn/health/kill is SSH; workers bind loopback on the remote box.
`machines.json` lists hosts (`local`, `ns1` = `marchon@70.88.205.138` with `~/.ssh/id_ed25519`, optional `lan-134`, **`px-worker`** = LXC 510 `10.10.10.20` on `vmbr1`). New replicas go to the **least-loaded** eligible host. Remote spawn/health/kill is SSH; workers bind loopback on the remote box.
NS1 **is** the Proxmox host. Extra worm/tree copies on `px-worker` mean a host-process crash does not take every bloom/tree replica. A second **chassis** is still the next step for disk/PSU failure.
Off-box probe: [verae-uptime](https://git.georgelambert.org/marchon/verae-uptime) (`node src/watch.js --once`). Backup: [verae-backup](https://git.georgelambert.org/marchon/verae-backup). Tagged upgrade: [verae-deploy](https://git.georgelambert.org/marchon/verae-deploy).
## JetStream