Snapshot of zapier-decisions (optimal NATS config study)

This commit is contained in:
George Lambert 2026-09-12 02:06:06 -04:00
commit 13f1663238
42 changed files with 6600 additions and 0 deletions

View file

@ -0,0 +1,853 @@
<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml" lang="" xml:lang="">
<head>
<meta charset="utf-8" />
<meta name="generator" content="pandoc" />
<meta name="viewport" content="width=device-width, initial-scale=1.0, user-scalable=yes" />
<title>NATS optimal configuration study</title>
<style>
html {
color: #1a1a1a;
background-color: #fdfdfd;
}
body {
margin: 0 auto;
max-width: 36em;
padding-left: 50px;
padding-right: 50px;
padding-top: 50px;
padding-bottom: 50px;
hyphens: auto;
overflow-wrap: break-word;
text-rendering: optimizeLegibility;
font-kerning: normal;
}
@media (max-width: 600px) {
body {
font-size: 0.9em;
padding: 12px;
}
h1 {
font-size: 1.8em;
}
}
@media print {
html {
background-color: white;
}
body {
background-color: transparent;
color: black;
font-size: 12pt;
}
p, h2, h3 {
orphans: 3;
widows: 3;
}
h2, h3, h4 {
page-break-after: avoid;
}
}
p {
margin: 1em 0;
}
a {
color: #1a1a1a;
}
a:visited {
color: #1a1a1a;
}
img {
max-width: 100%;
}
svg {
height: auto;
max-width: 100%;
}
h1, h2, h3, h4, h5, h6 {
margin-top: 1.4em;
}
h5, h6 {
font-size: 1em;
font-style: italic;
}
h6 {
font-weight: normal;
}
ol, ul {
padding-left: 1.7em;
margin-top: 1em;
}
li > ol, li > ul {
margin-top: 0;
}
blockquote {
margin: 1em 0 1em 1.7em;
padding-left: 1em;
border-left: 2px solid #e6e6e6;
color: #606060;
}
code {
font-family: Menlo, Monaco, Consolas, 'Lucida Console', monospace;
font-size: 85%;
margin: 0;
hyphens: manual;
}
pre {
margin: 1em 0;
overflow: auto;
}
pre code {
padding: 0;
overflow: visible;
overflow-wrap: normal;
}
.sourceCode {
background-color: transparent;
overflow: visible;
}
hr {
background-color: #1a1a1a;
border: none;
height: 1px;
margin: 1em 0;
}
table {
margin: 1em 0;
border-collapse: collapse;
width: 100%;
overflow-x: auto;
display: block;
font-variant-numeric: lining-nums tabular-nums;
}
table caption {
margin-bottom: 0.75em;
}
tbody {
margin-top: 0.5em;
border-top: 1px solid #1a1a1a;
border-bottom: 1px solid #1a1a1a;
}
th {
border-top: 1px solid #1a1a1a;
padding: 0.25em 0.5em 0.25em 0.5em;
}
td {
padding: 0.125em 0.5em 0.25em 0.5em;
}
header {
margin-bottom: 4em;
text-align: center;
}
#TOC li {
list-style: none;
}
#TOC ul {
padding-left: 1.3em;
}
#TOC > ul {
padding-left: 0;
}
#TOC a:not(:hover) {
text-decoration: none;
}
code{white-space: pre-wrap;}
span.smallcaps{font-variant: small-caps;}
div.columns{display: flex; gap: min(4vw, 1.5em);}
div.column{flex: auto; overflow-x: auto;}
div.hanging-indent{margin-left: 1.5em; text-indent: -1.5em;}
/* The extra [class] is a hack that increases specificity enough to
override a similar rule in reveal.js */
ul.task-list[class]{list-style: none;}
ul.task-list li input[type="checkbox"] {
font-size: inherit;
width: 0.8em;
margin: 0 0.8em 0.2em -1.6em;
vertical-align: middle;
}
.display.math{display: block; text-align: center; margin: 0.5rem auto;}
</style>
<style>/* Colored print + screen stylesheet for zapier.georgelambert.org */
:root {
--ink: #171a26;
--muted: #5b6178;
--line: #d9dce8;
--bg: #f4f5fb;
--paper: #ffffff;
--accent: #4f46e5;
--accent-deep: #312e81;
--accent-soft: #eef0fe;
--ok: #047857;
--warn: #8a5a00;
--code-bg: #1b1f33;
--code-fg: #e8ecff;
}
html { background: var(--bg); }
body {
margin: 0 auto;
padding: 1.5rem 1.25rem 3rem;
max-width: 48rem;
font: 15px/1.55 -apple-system, "Segoe UI", Georgia, serif;
color: var(--ink);
background: var(--paper);
}
.doc-banner {
background: linear-gradient(160deg, #312e81 0%, #4f46e5 60%, #7c74f0 100%);
color: #eef0fe;
margin: -1.5rem -1.25rem 1.5rem;
padding: 1.1rem 1.25rem 1rem;
}
.doc-banner a { color: #fff; }
.doc-banner .kicker {
letter-spacing: 0.12em;
text-transform: uppercase;
font: 700 10px system-ui, sans-serif;
opacity: 0.8;
}
.doc-banner h1 { margin: 0.25rem 0 0; font-size: 1.45rem; color: #fff; }
h1, h2, h3, h4 { color: var(--accent-deep); page-break-after: avoid; }
h1 { font-size: 1.7rem; }
h2 {
font-size: 1.2rem;
border-bottom: 2px solid var(--accent);
padding-bottom: 0.2rem;
margin-top: 1.6rem;
}
h3 { font-size: 1.05rem; color: var(--accent); }
a { color: var(--accent); }
p, li { orphans: 3; widows: 3; }
code {
font-family: ui-monospace, Menlo, Consolas, monospace;
font-size: 0.86em;
background: var(--accent-soft);
color: var(--accent-deep);
padding: 0.08em 0.28em;
border-radius: 4px;
}
pre, div.sourceCode, div.sourceCode pre {
background: var(--code-bg) !important;
color: var(--code-fg) !important;
padding: 0.85rem 1rem;
border-radius: 10px;
overflow: auto;
font-size: 0.78rem;
line-height: 1.4;
page-break-inside: avoid;
}
pre code { background: transparent; color: inherit; padding: 0; }
#title-block-header, header#title-block-header, h1.title { display: none; }
.doc-banner + h1 { display: none; }
table {
border-collapse: collapse;
width: 100%;
margin: 0.8rem 0 1.2rem;
font-size: 0.9rem;
page-break-inside: avoid;
}
th, td { border: 1px solid var(--line); padding: 0.38rem 0.55rem; text-align: left; vertical-align: top; }
th {
background: var(--accent);
color: #fff;
font: 650 12px system-ui, sans-serif;
}
tr:nth-child(even) td { background: var(--accent-soft); }
blockquote {
margin: 1rem 0;
padding: 0.4rem 0.9rem;
border-left: 4px solid var(--accent);
background: var(--accent-soft);
color: var(--accent-deep);
}
img { max-width: 100%; height: auto; border-radius: 8px; page-break-inside: avoid; }
hr { border: 0; border-top: 1px solid var(--line); }
ul, ol { padding-left: 1.25rem; }
nav.site { font: 13px system-ui, sans-serif; margin-bottom: 0.4rem; }
.source-path { font: 11px ui-monospace, Menlo, monospace; color: var(--muted); }
@page {
size: letter;
margin: 0.65in 0.7in 0.8in 0.7in;
@top-left {
content: "Verae Time × Zapier";
font: 700 8pt system-ui, sans-serif;
color: #4f46e5;
}
@top-right {
content: "zapier.georgelambert.org";
font: 8pt system-ui, sans-serif;
color: #6b7186;
}
@bottom-center {
content: counter(page) " / " counter(pages);
font: 8pt system-ui, sans-serif;
color: #6b7186;
}
}
@media print {
html, body { background: #fff; max-width: none; padding: 0; }
.doc-banner { margin: 0 0 1rem; border-radius: 8px; -webkit-print-color-adjust: exact; print-color-adjust: exact; }
a { text-decoration: none; }
th, tr:nth-child(even) td, pre, blockquote, code { -webkit-print-color-adjust: exact; print-color-adjust: exact; }
}
</style>
</head>
<body>
<div class="doc-banner"><nav class="site"><a href="/">zapier.georgelambert.org</a></nav><div class="kicker">Verae Time × Zapier · progress report</div><h1>NATS optimal configuration study</h1><div class="source-path">packages/zapier-decisions/reports/optimal-config/REPORT.md</div></div>
<header id="title-block-header">
<h1 class="title">NATS optimal configuration study</h1>
</header>
<p><strong>Progress report — optimal configuration study</strong> ·
<code>20260912T055851Z</code> (UTC) · all code on
<strong>NS1.GEORGELAMBERT.ORG</strong> (<code>70.88.205.138</code>)</p>
<p>This document folds every ladder we have run (1-core ZFS,
NS1-orchestrated, tmpfs maximize, and this exhaustive 8c/16G
<strong>ZFS</strong> factorial) plus UDP / MQTT / reconnect probes. It
recommends a lab config and a <strong>three-box HP DL360 Gen10</strong>
projection. veth/10G was not changed.</p>
<hr />
<h2 id="verdict-read-this-first">1. Verdict (read this first)</h2>
<p><strong>Keep NATS + JetStream.</strong> Do not replace the fabric
with MQTT, UDP, or a custom persistent-socket protocol for Verae
jobs/events/archive. Those are either slower, less durable, or already
what NATS is.</p>
<p><strong>Lab (NS1, one host, three LXC) — optimal now</strong></p>
<table>
<colgroup>
<col style="width: 25%" />
<col style="width: 28%" />
<col style="width: 31%" />
<col style="width: 15%" />
</colgroup>
<thead>
<tr class="header">
<th>Stream</th>
<th>Storage</th>
<th>Replicas</th>
<th>Why</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><code>ZAPIER_JOBS</code>, <code>ZAPIER_WEBHOOKS</code>,
<code>VERAE_ARCHIVE</code></td>
<td><strong>file</strong> (ZFS)</td>
<td><strong>3</strong></td>
<td>Survive a nats LXC death; archive must persist</td>
</tr>
<tr class="even">
<td><code>ZAPIER_EVENTS</code></td>
<td><strong>memory</strong></td>
<td><strong>3</strong></td>
<td>Waiters are latency-sensitive; events rebuild from job status</td>
</tr>
<tr class="odd">
<td><code>ZAPIER_USAGE</code></td>
<td>file</td>
<td>3</td>
<td>Telemetry, limits + max-age</td>
</tr>
</tbody>
</table>
<p>Keep <strong>8 cores / 16 GiB / <code>max_mem: 8G</code></strong> on
510513 (already live). Do <strong>not</strong> leave JetStream on
tmpfs. Do <strong>not</strong> drop product streams to r=1. Reuse
<strong>one NATS connection per process</strong> (already true in
middleware); never connect-per-message.</p>
<p><strong>Metal (3× DL360 Gen10) — optimal later</strong></p>
<p>Same stream table. File store on <strong>local NVMe/M.2</strong>, not
a shared SAN. Cluster + client on <strong>10GbE</strong> (or 25GbE if
you already have it). Dual Gold Xeon is surplus CPU for this workload;
816 cores dedicated to <code>nats-server</code> is enough. Expected JS
file r=3: <strong>~4080k</strong> 128 B pubs/s (about
<strong>36×</strong> this labs 8c ZFS 1p, <strong>24×</strong> tmpfs
1p) — bounded by <strong>10GbE replica RTT</strong>, not by Xeon clocks.
Core NATS will sit in the <strong>13M msgs/s</strong> band until the
NIC saturates (~9 Gbit/s ≈ 89M × 128 B theoretical; CPU and client will
hit first).</p>
<hr />
<h2 id="what-we-actually-ran-this-exhaustive-pass">2. What we actually
ran (this exhaustive pass)</h2>
<p>Live cluster during this run: LXC 510513 <strong>8 cores / 16
GiB</strong>, JetStream <strong>on ZFS</strong> (tmpfs from the maximize
study was already unmounted). Extra factorial: file/memory × replicas
1/3, 4 KiB file r=3, reconnect-per-message ping, UDP echo 510→511, MQTT
QoS0 against nats-a <code>:1883</code>. Product streams were not the
bench target.</p>
<h3 id="cross-study-history">2.1 Cross-study history</h3>
<table style="width:100%;">
<colgroup>
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
<col style="width: 16%" />
</colgroup>
<thead>
<tr class="header">
<th>Study</th>
<th>Env</th>
<th>Core 1p pub</th>
<th>JS file r=3 1p</th>
<th>JS mem r=3 4p</th>
<th>Ping p99</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><code>20260912T051237Z</code></td>
<td>1c/1G ZFS (NS1 orch.)</td>
<td>502,502</td>
<td>7,393</td>
<td></td>
<td>1.377ms</td>
</tr>
<tr class="even">
<td><code>20260912T053120Z</code></td>
<td>8c/16G tmpfs + mem extra</td>
<td>599,004</td>
<td>17,388</td>
<td>36,355</td>
<td>0.684ms</td>
</tr>
<tr class="odd">
<td><code>20260912T055851Z</code></td>
<td>8c/16G ZFS exhaustive <code>20260912T055851Z</code></td>
<td>662,227</td>
<td>14,330</td>
<td>37,736</td>
<td>1.140ms</td>
</tr>
</tbody>
</table>
<figure>
<img src="charts-optimal/history-js1p.png"
alt="JS 1p file r=3 history" />
<figcaption aria-hidden="true">JS 1p file r=3 history</figcaption>
</figure>
<h3 id="this-run-jetstream-factorial">2.2 This run — JetStream
factorial</h3>
<table>
<thead>
<tr class="header">
<th>Run</th>
<th>What</th>
<th>Pub msgs/s</th>
<th>Pub MB/s</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><code>js-file-1p-20k-128-r1</code></td>
<td>file r=1 1p 128 B</td>
<td>18,888</td>
<td>2.31</td>
</tr>
<tr class="even">
<td><code>js-file-4p-50k-128-r1</code></td>
<td>file r=1 4p 128 B</td>
<td>24,560</td>
<td>3.00</td>
</tr>
<tr class="odd">
<td><code>js-1p-20k-128-r3</code></td>
<td>file r=3 1p 128 B</td>
<td>14,330</td>
<td>1.75</td>
</tr>
<tr class="even">
<td><code>js-4p-50k-128-r3</code></td>
<td>file r=3 4p 128 B</td>
<td>19,232</td>
<td>2.35</td>
</tr>
<tr class="odd">
<td><code>js-4p-20k-1k-r3</code></td>
<td>file r=3 4p 1 KiB</td>
<td>15,197</td>
<td>14.84</td>
</tr>
<tr class="even">
<td><code>js-file-1p-20k-4k-r3</code></td>
<td>file r=3 1p 4 KiB</td>
<td>8,673</td>
<td>33.88</td>
</tr>
<tr class="odd">
<td><code>js-mem-1p-20k-128-r1</code></td>
<td>memory r=1 1p 128 B</td>
<td>29,972</td>
<td>3.66</td>
</tr>
<tr class="even">
<td><code>js-mem-4p-50k-128-r1</code></td>
<td>memory r=1 4p 128 B</td>
<td>64,923</td>
<td>7.93</td>
</tr>
<tr class="odd">
<td><code>js-mem-1p-20k-128-r3</code></td>
<td>memory r=3 1p 128 B</td>
<td>20,188</td>
<td>2.46</td>
</tr>
<tr class="even">
<td><code>js-mem-4p-50k-128-r3</code></td>
<td>memory r=3 4p 128 B</td>
<td>37,736</td>
<td>4.61</td>
</tr>
<tr class="odd">
<td><code>js-mem-4p-20k-1k-r3</code></td>
<td>memory r=3 4p 1 KiB</td>
<td>33,916</td>
<td>33.12</td>
</tr>
</tbody>
</table>
<p>Replica <strong>1 vs 3</strong> on this stand (file 1p 128 B): r=1 is
18,888 vs r=3 14,330 (1.32× if r=3 is the slower one). Memory r=1 1p
29,972 vs memory r=3 20,188.</p>
<figure>
<img src="charts-optimal/replicas.png" alt="Replica cost" />
<figcaption aria-hidden="true">Replica cost</figcaption>
</figure>
<h3 id="delay-reconnect-tax-udp-mqtt">2.3 Delay, reconnect tax, UDP,
MQTT</h3>
<table>
<colgroup>
<col style="width: 29%" />
<col style="width: 33%" />
<col style="width: 37%" />
</colgroup>
<thead>
<tr class="header">
<th>Probe</th>
<th>Result</th>
<th>Meaning</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>NATS ping (persistent sockets) p50 / p99</td>
<td>0.456ms / 1.140ms</td>
<td>Quiet hop with a long-lived TCP conn</td>
</tr>
<tr class="even">
<td>NATS <strong>reconnect-per-message</strong> p50 / p99</td>
<td>0.503ms / 1.750ms</td>
<td>TCP+NATS handshake on every pub — this is the tax to avoid</td>
</tr>
<tr class="odd">
<td>UDP echo 510→511 p99</td>
<td>0.363ms</td>
<td>Raw datagram ceiling on the same veth (no NATS)</td>
</tr>
<tr class="even">
<td>MQTT QoS0 5k×128 B</td>
<td>44862 pubs/s</td>
<td>nats-server MQTT gateway on <code>:1883</code></td>
</tr>
</tbody>
</table>
<p>Core 1p1s 128 B this run: 662,227 pub msgs/s. Flood delay is still
backlog/consume_rate, not RTT.</p>
<hr />
<h2 id="alternative-transports-why-we-are-not-switching-the-fabric">3.
Alternative transports (why we are not switching the fabric)</h2>
<p>NATS already <strong>is</strong> persistent TCP sockets with a tiny
binary protocol, automatic reconnect, and optional JetStream durability.
“Reduce connection overhead” is a <strong>client</strong> discipline:
hold the connection. The reconnect probe exists to prove that opening a
socket per job would dominate ping RTT.</p>
<table>
<colgroup>
<col style="width: 7%" />
<col style="width: 44%" />
<col style="width: 32%" />
<col style="width: 15%" />
</colgroup>
<thead>
<tr class="header">
<th>Idea</th>
<th>Fit for Verae jobs/events/archive</th>
<th>Throughput vs NATS core</th>
<th>Durability</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td><strong>NATS core pub/sub</strong></td>
<td>Fan-out, request-reply (<code>verae.billing.*</code>)</td>
<td>Highest we measured (~0.52M msgs/s)</td>
<td>None</td>
</tr>
<tr class="even">
<td><strong>NATS JetStream file r=3</strong></td>
<td>Jobs, webhooks, archive</td>
<td>~823k on this lab; see metal projection</td>
<td>Disk + 1-node loss</td>
</tr>
<tr class="odd">
<td><strong>NATS JetStream memory r=3</strong></td>
<td>Events mailbox</td>
<td>~2236k on this lab</td>
<td>RAM + 1-node loss; <strong>empty on full restart</strong></td>
</tr>
<tr class="even">
<td><strong>MQTT</strong> (NATS gateway or Mosquitto)</td>
<td>IoT endpoints that already speak MQTT</td>
<td>This probe: 44862 pubs/s QoS0 — typically <strong>well
below</strong> NATS core; QoS1 ≈ JetStream-ish with more chatter</td>
<td>QoS1/2 session state; not our WORM model</td>
</tr>
<tr class="odd">
<td><strong>UDP</strong></td>
<td>Telemetry that may drop</td>
<td>RTT 0.363ms p99 — fastest hop, <strong>no</strong> reliability, no
cluster, no auth</td>
<td>None</td>
</tr>
<tr class="even">
<td><strong>Custom persistent sockets / HTTP long-poll</strong></td>
<td>Worse NATS</td>
<td>You would re-implement reconnect, flow control, and fan-out</td>
<td>DIY</td>
</tr>
<tr class="odd">
<td><strong>WebSocket</strong></td>
<td>Browsers only</td>
<td>Extra framing; NATS already has WS for UIs, not for middleware</td>
<td>Same as core/JS behind it</td>
</tr>
<tr class="even">
<td><strong>QUIC / WebTransport</strong></td>
<td>Lossy WAN / browsers</td>
<td>NATS QUIC is not the lab path; 10GbE LAN does not need it</td>
<td>Same</td>
</tr>
<tr class="odd">
<td><strong>Kafka / Redis streams</strong></td>
<td>Heavy log replay</td>
<td>Higher ops cost; not on <code>vmbr1</code> today</td>
<td>Yes, heavier</td>
</tr>
</tbody>
</table>
<p><strong>MQTT:</strong> NATS documents MQTT as an <em>enabling</em>
gateway for existing IoT, and prefers NATS end-to-end for greenfield.
Zapier cloud never talks NATS or MQTT; it talks HTTPS. Putting MQTT in
the middle of timestamp jobs adds protocol translation and QoS timers
without helping <code>jobId → events</code>. Use MQTT only if a device
already cannot speak NATS.</p>
<p><strong>UDP:</strong> Fine as a <em>measurement</em> of veth RTT.
Unusable as the job fabric (no ack, no replica, no flow control). NATS
ping is already within a small multiple of UDP on this bridge.</p>
<p><strong>Persistence sockets:</strong> Middleware and keep already
keep <code>NATS_URL</code> connections open. Optimal: one connection (or
a small pool) per process, <code>max_reconnect</code>, jitter, no
<code>connect()</code> in the per-job path. The reconnect ladder is the
anti-pattern.</p>
<hr />
<h2 id="optimal-configurations">4. Optimal configurations</h2>
<h3 id="ns1-lab-now">4.1 NS1 lab (now)</h3>
<ol type="1">
<li><strong>Leave 8 cores / 16 GiB</strong> on nats-a/b/c and the
worker. Host has 40 cores / 377 GiB; this is cheap.</li>
<li><strong><code>max_mem: 8G</code></strong> stays. Required for memory
streams.</li>
<li><strong>File r=3 on ZFS</strong> for jobs/webhooks/archive. tmpfs
doubled JS 1p (7.4k→17k) but <strong>loses the stream on reboot</strong>
— unacceptable for archive.</li>
<li><strong>Memory r=3 for <code>ZAPIER_EVENTS</code></strong> if we
accept “all three nats CTs reboot ⇒ in-flight waiters fall back to HTTP
poll.” That matches the designed wait path
(<code>GET /api/status/{jobId}</code>).</li>
<li><strong>r=1 only for throwaway benches</strong>, never product
streams. Replica=3 is the point of three guests.</li>
<li><strong>veth on vmbr1, no fake 10G NICs.</strong> Already 10000Mb/s;
JS does not fill it.</li>
<li><strong>Pin cpusets</strong> later if keep/fleet steal; not required
to beat these numbers.</li>
<li>Clients: persistent NATS connections; pull consumers with bounded
<code>max_ack_pending</code> for webhooks.</li>
</ol>
<h3 id="three-hp-dl360-gen10-projection-not-measured">4.2 Three HP DL360
Gen10 (projection — not measured)</h3>
<p>Assumed bill of materials (state it in the buy):</p>
<table>
<colgroup>
<col style="width: 36%" />
<col style="width: 63%" />
</colgroup>
<thead>
<tr class="header">
<th>Piece</th>
<th>Assumption</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Chassis</td>
<td>3× DL360 Gen10 1U</td>
</tr>
<tr class="even">
<td>CPU</td>
<td>Dual 2nd-gen Xeon <strong>Gold</strong> (e.g. 6226R 16c or 6248 20c
<strong>3240 cores/box</strong>)</td>
</tr>
<tr class="odd">
<td>Memory</td>
<td>DDR4-2933, <strong>192384 GiB</strong>/box (612×32 GiB); NATS will
not use most of it</td>
</tr>
<tr class="even">
<td>Storage</td>
<td><strong>NVMe M.2 or U.2</strong> for
<code>/var/lib/nats/jetstream</code> (XFS or ext4, <strong>not</strong>
shared ZFS over the network). RAID1 of two NVMe if you want disk HA
<em>inside</em> a box</td>
</tr>
<tr class="odd">
<td>Network</td>
<td><strong>10GbE</strong> (FlexibleLOM or PCIe); dedicated VLAN for
<code>:4222</code>+<code>:6222</code>. Do not share with public
<code>vmbr0</code> traffic</td>
</tr>
<tr class="even">
<td>OS</td>
<td>Debian/Ubuntu bare metal, <code>nats-server</code> systemd, same
<code>nats.conf</code> as lab (bind private IP only)</td>
</tr>
</tbody>
</table>
<p><strong>What changes vs NS1 LXC</strong></p>
<table>
<colgroup>
<col style="width: 15%" />
<col style="width: 20%" />
<col style="width: 18%" />
<col style="width: 45%" />
</colgroup>
<thead>
<tr class="header">
<th>Factor</th>
<th>NS1 today</th>
<th>3× DL360</th>
<th>Effect on JS file r=3</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Failure domain</td>
<td>1 Proxmox host</td>
<td>3 chassis, 3 NVMe, 3 NICs</td>
<td>r=3 <strong>means</strong> something</td>
</tr>
<tr class="even">
<td>Disk</td>
<td>Shared ZFS SSD2</td>
<td>Local NVMe fsync ~50150 µs</td>
<td>Big win vs ZFS; similar to tmpfs for sequential 128 B</td>
</tr>
<tr class="odd">
<td>Replica path</td>
<td>veth/bridge (~µstens of µs)</td>
<td>10GbE RTT typically <strong>50200 µs</strong></td>
<td><strong>Slower than same-host tmpfs</strong>, faster than a bad
SAN</td>
</tr>
<tr class="even">
<td>CPU</td>
<td>8 of 40 shared</td>
<td>3240 dedicated Gold cores</td>
<td>Headroom for many clients, not 10× JS</td>
</tr>
<tr class="odd">
<td>NIC</td>
<td>software 10G veth, already ~5 Gbit/s core</td>
<td>real 10GbE ~9 Gbit/s TCP</td>
<td>Core NATS can grow; JS r=3 stays replica-bound</td>
</tr>
</tbody>
</table>
<p><strong>Projected bands</strong> (128 B, 3-node cluster, dedicated
10GbE, local NVMe, 8+ cores pinned to nats-server):</p>
<table>
<colgroup>
<col style="width: 16%" />
<col style="width: 34%" />
<col style="width: 29%" />
<col style="width: 19%" />
</colgroup>
<thead>
<tr class="header">
<th>Workload</th>
<th>NS1 measured (best)</th>
<th>DL360 projection</th>
<th>Confidence</th>
</tr>
</thead>
<tbody>
<tr class="odd">
<td>Core pub/sub 1p</td>
<td>0.50.8M</td>
<td><strong>0.82M</strong></td>
<td>Medium — NIC + syscall, plenty of CPU</td>
</tr>
<tr class="even">
<td>Core 4p4s 1 KiB</td>
<td>~0.60.7M (~0.6 GB/s)</td>
<td><strong>~1M msgs/s / ~1 GB/s</strong> approaching 10GbE</td>
<td>Medium</td>
</tr>
<tr class="odd">
<td>JS file r=1</td>
<td>this run r=1</td>
<td><strong>80200k</strong> pubs/s</td>
<td>Medium — NVMe + no replica wait</td>
</tr>
<tr class="even">
<td>JS file r=3</td>
<td>723k (ZFS/tmpfs)</td>
<td><strong>4080k</strong> pubs/s</td>
<td>Medium-low — replica RTT dominates; 3 NVMe still help vs shared
ZFS</td>
</tr>
<tr class="odd">
<td>JS memory r=3</td>
<td>2236k</td>
<td><strong>50100k</strong></td>
<td>Medium-low — RAM + 10GbE ack</td>
</tr>
<tr class="even">
<td>Ping p99</td>
<td>0.71.4 ms</td>
<td><strong>0.20.6 ms</strong></td>
<td>Medium — real NIC but no Proxmox tax</td>
</tr>
</tbody>
</table>
<p>These are <strong>not</strong> DL360 measurements. Scale from: (a)
our replica-1 vs replica-3 ratio once this runs r=1 numbers exist, (b)
tmpfs vs ZFS ratio (2.35× on 1p), (c) Synadia/nats bench async file r=1
~100400k on NVMe loopback, derated for 10GbE RTT.</p>
<p><strong>Buy notes:</strong> M.2 via Dual uFF / enablement kit; put
JetStream on NVMe <strong>directly</strong>, not behind a RAID
controller write-through unless you measure. 1GbE onboard is a trap —
use 10GbE for <code>:6222</code>. Dual Gold is for isolation (nats vs
worm/tree vs OS), not because JS needs 56 cores.</p>
<hr />
<h2 id="what-we-are-not-doing">5. What we are not doing</h2>
<ul>
<li>MQTT as the Zapier or middleware transport.</li>
<li>UDP for jobs.</li>
<li>Emulated 10G fiber NICs on LXC.</li>
<li>tmpfs as the production store.</li>
<li>r=1 for product streams.</li>
<li>Connect-per-job.</li>
</ul>
<p>Re-run exhaustive: <code>bash scripts/exhaustive-ns1-study.sh</code>
on NS1.</p>
</body>
</html>

View file

@ -0,0 +1,160 @@
**Progress report — optimal configuration study** · `20260912T055851Z` (UTC) · all code on **NS1.GEORGELAMBERT.ORG** (`70.88.205.138`)
This document folds every ladder we have run (1-core ZFS, NS1-orchestrated, tmpfs maximize, and this exhaustive 8c/16G **ZFS** factorial) plus UDP / MQTT / reconnect probes. It recommends a lab config and a **three-box HP DL360 Gen10** projection. veth/10G was not changed.
---
## 1. Verdict (read this first)
**Keep NATS + JetStream.** Do not replace the fabric with MQTT, UDP, or a custom persistent-socket protocol for Verae jobs/events/archive. Those are either slower, less durable, or already what NATS is.
**Lab (NS1, one host, three LXC) — optimal now**
| Stream | Storage | Replicas | Why |
|--------|---------|----------|-----|
| `ZAPIER_JOBS`, `ZAPIER_WEBHOOKS`, `VERAE_ARCHIVE` | **file** (ZFS) | **3** | Survive a nats LXC death; archive must persist |
| `ZAPIER_EVENTS` | **memory** | **3** | Waiters are latency-sensitive; events rebuild from job status |
| `ZAPIER_USAGE` | file | 3 | Telemetry, limits + max-age |
Keep **8 cores / 16 GiB / `max_mem: 8G`** on 510513 (already live). Do **not** leave JetStream on tmpfs. Do **not** drop product streams to r=1. Reuse **one NATS connection per process** (already true in middleware); never connect-per-message.
**Metal (3× DL360 Gen10) — optimal later**
Same stream table. File store on **local NVMe/M.2**, not a shared SAN. Cluster + client on **10GbE** (or 25GbE if you already have it). Dual Gold Xeon is surplus CPU for this workload; 816 cores dedicated to `nats-server` is enough. Expected JS file r=3: **~4080k** 128 B pubs/s (about **36×** this labs 8c ZFS 1p, **24×** tmpfs 1p) — bounded by **10GbE replica RTT**, not by Xeon clocks. Core NATS will sit in the **13M msgs/s** band until the NIC saturates (~9 Gbit/s ≈ 89M × 128 B theoretical; CPU and client will hit first).
---
## 2. What we actually ran (this exhaustive pass)
Live cluster during this run: LXC 510513 **8 cores / 16 GiB**, JetStream **on ZFS** (tmpfs from the maximize study was already unmounted). Extra factorial: file/memory × replicas 1/3, 4 KiB file r=3, reconnect-per-message ping, UDP echo 510→511, MQTT QoS0 against nats-a `:1883`. Product streams were not the bench target.
### 2.1 Cross-study history
| Study | Env | Core 1p pub | JS file r=3 1p | JS mem r=3 4p | Ping p99 |
| --- | --- | --- | --- | --- | --- |
| `20260912T051237Z` | 1c/1G ZFS (NS1 orch.) | 502,502 | 7,393 | — | 1.377ms |
| `20260912T053120Z` | 8c/16G tmpfs + mem extra | 599,004 | 17,388 | 36,355 | 0.684ms |
| `20260912T055851Z` | 8c/16G ZFS exhaustive `20260912T055851Z` | 662,227 | 14,330 | 37,736 | 1.140ms |
![JS 1p file r=3 history](charts-optimal/history-js1p.png)
### 2.2 This run — JetStream factorial
| Run | What | Pub msgs/s | Pub MB/s |
| --- | --- | --- | --- |
| `js-file-1p-20k-128-r1` | file r=1 1p 128 B | 18,888 | 2.31 |
| `js-file-4p-50k-128-r1` | file r=1 4p 128 B | 24,560 | 3.00 |
| `js-1p-20k-128-r3` | file r=3 1p 128 B | 14,330 | 1.75 |
| `js-4p-50k-128-r3` | file r=3 4p 128 B | 19,232 | 2.35 |
| `js-4p-20k-1k-r3` | file r=3 4p 1 KiB | 15,197 | 14.84 |
| `js-file-1p-20k-4k-r3` | file r=3 1p 4 KiB | 8,673 | 33.88 |
| `js-mem-1p-20k-128-r1` | memory r=1 1p 128 B | 29,972 | 3.66 |
| `js-mem-4p-50k-128-r1` | memory r=1 4p 128 B | 64,923 | 7.93 |
| `js-mem-1p-20k-128-r3` | memory r=3 1p 128 B | 20,188 | 2.46 |
| `js-mem-4p-50k-128-r3` | memory r=3 4p 128 B | 37,736 | 4.61 |
| `js-mem-4p-20k-1k-r3` | memory r=3 4p 1 KiB | 33,916 | 33.12 |
Replica **1 vs 3** on this stand (file 1p 128 B): r=1 is 18,888 vs r=3 14,330 (1.32× if r=3 is the slower one). Memory r=1 1p 29,972 vs memory r=3 20,188.
![Replica cost](charts-optimal/replicas.png)
### 2.3 Delay, reconnect tax, UDP, MQTT
| Probe | Result | Meaning |
|-------|--------|---------|
| NATS ping (persistent sockets) p50 / p99 | 0.456ms / 1.140ms | Quiet hop with a long-lived TCP conn |
| NATS **reconnect-per-message** p50 / p99 | 0.503ms / 1.750ms | TCP+NATS handshake on every pub — this is the tax to avoid |
| UDP echo 510→511 p99 | 0.363ms | Raw datagram ceiling on the same veth (no NATS) |
| MQTT QoS0 5k×128 B | 44862 pubs/s | nats-server MQTT gateway on `:1883` |
Core 1p1s 128 B this run: 662,227 pub msgs/s. Flood delay is still backlog/consume_rate, not RTT.
---
## 3. Alternative transports (why we are not switching the fabric)
NATS already **is** persistent TCP sockets with a tiny binary protocol, automatic reconnect, and optional JetStream durability. “Reduce connection overhead” is a **client** discipline: hold the connection. The reconnect probe exists to prove that opening a socket per job would dominate ping RTT.
| Idea | Fit for Verae jobs/events/archive | Throughput vs NATS core | Durability |
|------|-----------------------------------|-------------------------|------------|
| **NATS core pub/sub** | Fan-out, request-reply (`verae.billing.*`) | Highest we measured (~0.52M msgs/s) | None |
| **NATS JetStream file r=3** | Jobs, webhooks, archive | ~823k on this lab; see metal projection | Disk + 1-node loss |
| **NATS JetStream memory r=3** | Events mailbox | ~2236k on this lab | RAM + 1-node loss; **empty on full restart** |
| **MQTT** (NATS gateway or Mosquitto) | IoT endpoints that already speak MQTT | This probe: 44862 pubs/s QoS0 — typically **well below** NATS core; QoS1 ≈ JetStream-ish with more chatter | QoS1/2 session state; not our WORM model |
| **UDP** | Telemetry that may drop | RTT 0.363ms p99 — fastest hop, **no** reliability, no cluster, no auth | None |
| **Custom persistent sockets / HTTP long-poll** | Worse NATS | You would re-implement reconnect, flow control, and fan-out | DIY |
| **WebSocket** | Browsers only | Extra framing; NATS already has WS for UIs, not for middleware | Same as core/JS behind it |
| **QUIC / WebTransport** | Lossy WAN / browsers | NATS QUIC is not the lab path; 10GbE LAN does not need it | Same |
| **Kafka / Redis streams** | Heavy log replay | Higher ops cost; not on `vmbr1` today | Yes, heavier |
**MQTT:** NATS documents MQTT as an *enabling* gateway for existing IoT, and prefers NATS end-to-end for greenfield. Zapier cloud never talks NATS or MQTT; it talks HTTPS. Putting MQTT in the middle of timestamp jobs adds protocol translation and QoS timers without helping `jobId → events`. Use MQTT only if a device already cannot speak NATS.
**UDP:** Fine as a *measurement* of veth RTT. Unusable as the job fabric (no ack, no replica, no flow control). NATS ping is already within a small multiple of UDP on this bridge.
**Persistence sockets:** Middleware and keep already keep `NATS_URL` connections open. Optimal: one connection (or a small pool) per process, `max_reconnect`, jitter, no `connect()` in the per-job path. The reconnect ladder is the anti-pattern.
---
## 4. Optimal configurations
### 4.1 NS1 lab (now)
1. **Leave 8 cores / 16 GiB** on nats-a/b/c and the worker. Host has 40 cores / 377 GiB; this is cheap.
2. **`max_mem: 8G`** stays. Required for memory streams.
3. **File r=3 on ZFS** for jobs/webhooks/archive. tmpfs doubled JS 1p (7.4k→17k) but **loses the stream on reboot** — unacceptable for archive.
4. **Memory r=3 for `ZAPIER_EVENTS`** if we accept “all three nats CTs reboot ⇒ in-flight waiters fall back to HTTP poll.” That matches the designed wait path (`GET /api/status/{jobId}`).
5. **r=1 only for throwaway benches**, never product streams. Replica=3 is the point of three guests.
6. **veth on vmbr1, no fake 10G NICs.** Already 10000Mb/s; JS does not fill it.
7. **Pin cpusets** later if keep/fleet steal; not required to beat these numbers.
8. Clients: persistent NATS connections; pull consumers with bounded `max_ack_pending` for webhooks.
### 4.2 Three HP DL360 Gen10 (projection — not measured)
Assumed bill of materials (state it in the buy):
| Piece | Assumption |
|-------|------------|
| Chassis | 3× DL360 Gen10 1U |
| CPU | Dual 2nd-gen Xeon **Gold** (e.g. 6226R 16c or 6248 20c — **3240 cores/box**) |
| Memory | DDR4-2933, **192384 GiB**/box (612×32 GiB); NATS will not use most of it |
| Storage | **NVMe M.2 or U.2** for `/var/lib/nats/jetstream` (XFS or ext4, **not** shared ZFS over the network). RAID1 of two NVMe if you want disk HA *inside* a box |
| Network | **10GbE** (FlexibleLOM or PCIe); dedicated VLAN for `:4222`+`:6222`. Do not share with public `vmbr0` traffic |
| OS | Debian/Ubuntu bare metal, `nats-server` systemd, same `nats.conf` as lab (bind private IP only) |
**What changes vs NS1 LXC**
| Factor | NS1 today | 3× DL360 | Effect on JS file r=3 |
|--------|-----------|----------|------------------------|
| Failure domain | 1 Proxmox host | 3 chassis, 3 NVMe, 3 NICs | r=3 **means** something |
| Disk | Shared ZFS SSD2 | Local NVMe fsync ~50150 µs | Big win vs ZFS; similar to tmpfs for sequential 128 B |
| Replica path | veth/bridge (~µstens of µs) | 10GbE RTT typically **50200 µs** | **Slower than same-host tmpfs**, faster than a bad SAN |
| CPU | 8 of 40 shared | 3240 dedicated Gold cores | Headroom for many clients, not 10× JS |
| NIC | software 10G veth, already ~5 Gbit/s core | real 10GbE ~9 Gbit/s TCP | Core NATS can grow; JS r=3 stays replica-bound |
**Projected bands** (128 B, 3-node cluster, dedicated 10GbE, local NVMe, 8+ cores pinned to nats-server):
| Workload | NS1 measured (best) | DL360 projection | Confidence |
|----------|---------------------|------------------|------------|
| Core pub/sub 1p | 0.50.8M | **0.82M** | Medium — NIC + syscall, plenty of CPU |
| Core 4p4s 1 KiB | ~0.60.7M (~0.6 GB/s) | **~1M msgs/s / ~1 GB/s** approaching 10GbE | Medium |
| JS file r=1 | this run r=1 | **80200k** pubs/s | Medium — NVMe + no replica wait |
| JS file r=3 | 723k (ZFS/tmpfs) | **4080k** pubs/s | Medium-low — replica RTT dominates; 3 NVMe still help vs shared ZFS |
| JS memory r=3 | 2236k | **50100k** | Medium-low — RAM + 10GbE ack |
| Ping p99 | 0.71.4 ms | **0.20.6 ms** | Medium — real NIC but no Proxmox tax |
These are **not** DL360 measurements. Scale from: (a) our replica-1 vs replica-3 ratio once this runs r=1 numbers exist, (b) tmpfs vs ZFS ratio (2.35× on 1p), (c) Synadia/nats bench async file r=1 ~100400k on NVMe loopback, derated for 10GbE RTT.
**Buy notes:** M.2 via Dual uFF / enablement kit; put JetStream on NVMe **directly**, not behind a RAID controller write-through unless you measure. 1GbE onboard is a trap — use 10GbE for `:6222`. Dual Gold is for isolation (nats vs worm/tree vs OS), not because JS needs 56 cores.
---
## 5. What we are not doing
- MQTT as the Zapier or middleware transport.
- UDP for jobs.
- Emulated 10G fiber NICs on LXC.
- tmpfs as the production store.
- r=1 for product streams.
- Connect-per-job.
Re-run exhaustive: `bash scripts/exhaustive-ns1-study.sh` on NS1.

Binary file not shown.

Binary file not shown.

After

Width:  |  Height:  |  Size: 25 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 27 KiB