The FIFO’s state (the request holding the turn, the requests waiting) and the write counters are held only in memory,
for the life of the process. Nothing is saved, and nothing is rebuilt: a new process starts with an empty FIFO, which
is right, since every client waiting before a crash lost its connection too.
The Sequence Position counter and the store ID are rebuilt at startup, from the database, because they have to match
what’s on disk.
On SIGINT or SIGTERM, the server shuts down in order:
The HTTP server stops taking new connections (http.Server.Shutdown, capped at 10 seconds).
At the same moment, the
write FIFO closes: requests still waiting get 503 ShuttingDown, and so does any later one. A request already holding the turn finishes.
The HTTP server finishes the requests still in flight.
The store closes last, releasing its connections and the .lock file.
Why the FIFO closes first. A request waiting in it only ends once it gets its turn, so the HTTP server would
otherwise wait for it.
A panic in a request handler is caught by the HTTP server, without crashing the process. The handler’s deferred
rollback and the release of its turn still run during the unwind, so no turn stays taken. A write’s reserved Sequence
Positions are given back the same way.
A process crash (an unrecovered panic, SIGKILL, an out-of-memory kill) takes the whole in-memory state with it,
so nothing is left to leak.
A crash during a write leaves an uncommitted WAL transaction, discarded the next time a connection opens. Neither
the write’s events nor its projections ever become visible, in whole or in part. The counter, read back from the
database, matches.
A SQLite error that suggests the file itself may be damaged (an I/O error, detected corruption, failure to open the
database file) is fatal: the server logs it and drives the same ordered shutdown as a signal.
A temporary SQLite error (a busy lock during a WAL checkpoint, say) isn’t fatal: it fails that one request, and
nothing it would have written is kept.
Any error while opening the store at startup is fatal.
Why. No per-error recovery logic: a clean restart is cheap and safe given the in-memory state above, and /health
and a process supervisor are already set up to catch it.