Audit retention and pruning
The audit log is an HMAC chain, so removing rows from it is not a plain
DELETE. Every removal the brain performs is recorded in the authenticated
chain state as a prune boundary, and that record is what lets
z4j audit verify tell retention apart from a truncation. This page covers
the three ways rows leave the table (the periodic sweep, z4j audit prune
in soft mode, and an epoch cut in hard mode), the retention windows that
drive them, and what verification says afterwards.
How a prune works
Section titled “How a prune works”Rows are only ever removed as a contiguous prefix, oldest first. For each
batch the brain verifies every row of the prefix against the chain, checks
that the first row follows the current boundary and that the row after the
prefix links to its last row, deletes exactly those rows, and advances the
signed prune boundary in audit_chain_state to the newest row it removed,
all in one transaction. The next row ever appended links to that boundary,
and the verifier accepts it as the chain's anchor.
The consequence worth knowing: the prune stops at the first row that retention still keeps, whatever lies beyond it. A row the policy would remove is retained while an older row before it is kept. The chain can describe one boundary, not a set of holes.
When that happens with expired rows waiting behind the retained row, the
sweep logs one WARNING per pass naming the row that holds the line: its
class and age, the moment its class window keeps it until, and how many
expired rows are retained behind it. A long class window is a decision to
keep everything written after its oldest retained row; the warning is the
size of that decision, not a fault. The dry run below prints the same row
under stops at.
The periodic sweep runs on Z4J_AUDIT_RETENTION_SWEEP_INTERVAL_SECONDS,
removes at most Z4J_AUDIT_RETENTION_SWEEP_BATCH_SIZE rows per statement
and Z4J_AUDIT_RETENTION_SWEEP_MAX_PER_PASS per pass, and commits each
batch, so an interrupted pass leaves a signed, consistent state and the next
pass continues. Only one process sweeps at a time: on PostgreSQL the sweep
takes an advisory lock, on SQLite the writer lock serialises it.
Retention windows
Section titled “Retention windows”Z4J_AUDIT_RETENTION_DAYS (default 90) is the window for every row. Set it
to your real obligation before the first sweep runs: rows past the window
are deleted, not archived.
There is no ceiling. An obligation can run past ten years, so a window that long is accepted; the brain logs one WARNING at startup naming each window above 3650 days, because a value that large is more often a typo than a policy, and the table grows for the whole window.
Retention by action class
Section titled “Retention by action class”Z4J_AUDIT_RETENTION_BY_CLASS gives a class of actions its own window. It
is a JSON object mapping an action class to days; a class not listed uses
the global window. The class of an action is the first dotted segment of its
name: auth for auth.login, command for
command.issue.requeue_dead_letter, dead_letters for dead_letters.list,
audit for the chain's own markers.
Z4J_AUDIT_RETENTION_DAYS=90Z4J_AUDIT_RETENTION_BY_CLASS='{"auth": 365, "command": 30}'Validation is strict and happens when settings load: keys must look like an
action class (lower-case letters, digits and underscores, no dots), values
must be whole positive numbers of days. A key that names a whole action
(auth.login) is rejected rather than silently matching nothing. A class
window may be longer or shorter than the global one, and the prefix rule
above decides what each buys:
- Longer (
auth: 365above a global 90): everyauthrow is kept for a year, and so is every row written after the oldestauthrow still inside that year, whatever its class. The chain holds at that row. If you need a long window for one class, expect the whole trail from that class's oldest retained row onward to stay. - Shorter (
command: 30under a global 90): acommandrow is removed once it is both past 30 days and older than every retained row before it. In a trail where classes interleave, that means the shorter window takes effect only on the rows at the very oldest end, before the oldest row any other window keeps. It never punches a hole in the middle of the chain.
The sweep and z4j audit prune apply exactly the same rule, and the dry
run below shows where the prefix stops and why.
Settings are read from the environment when the brain starts. There is no write-back: nothing in the dashboard or the API changes a retention window, and changing one is a restart.
z4j audit prune
Section titled “z4j audit prune”The operator-driven form of the same prune, for the cases the periodic sweep does not cover: pruning now rather than on the next tick, pruning to an explicit date, confirming what a policy change would remove before it runs, and cutting an epoch.
# What the configured windows would remove right now. Changes nothing.z4j audit prune
# Remove it, under the authenticated boundary.z4j audit prune --apply
# Prune every class before an explicit moment instead of the windows.z4j audit prune --apply --before 2026-01-31T00:00:00Z
# Epoch cut: once the generation is fully pruned, start a fresh signed one.z4j audit prune --hard --apply --before 2026-06-30T00:00:00ZIt requires the audit-chain key and an authenticated chain state, so it runs with the brain's settings, not with a read-only database role.
Soft mode (the default) is the authenticated prefix prune described
above, driven by hand. Rows older than their class cutoff, up to the first
row retention keeps, are verified, removed and recorded under the signed
boundary. The chain stays in the same generation; the next row links to the
boundary; verify reports the chain clean with the prune on record.
Hard mode (--hard) is an epoch cut. After the soft prune, the fully
pruned generation is replaced by a fresh one: one signed
audit.chain_generation_reset row becomes the new genesis and the chain
state carries no prune boundary at all, so nothing in the live chain refers
to the old history. It refuses, before changing anything, while any active
row younger than the cutoff remains (an epoch cut never removes a row the
cutoff keeps; pass --before later than the newest row you are prepared to
lose) and while frozen legacy rows exist (export them first with
z4j audit export-and-delete-frozen). Every head exported from the old
generation reports UNPROVABLE afterwards; export a new one.
Those preconditions are checked on the preview, which runs before the
command takes its leases, and a live brain can append a row in between.
The reset therefore checks every active row against the same cutoffs
again, inside its own transaction with the chain locks held, and refuses
with exit 1 naming the row if one younger than its cutoff has appeared.
The soft prune's batches stay committed under the signed boundary, the
generation is not reset, and no audit.prune row is written. Run hard
mode on a quiet brain (a maintenance window, or with the agents stopped)
and treat the refusal as safe to rerun.
| Flag | Effect |
|---|---|
| (none) | Dry run. Prints the mode, the cutoffs, the generation, the rows that would be removed by class, the row the prefix stops at and why, the boundary that would be recorded, and the rows that remain. Changes nothing. |
--apply |
Execute. |
--hard |
Epoch cut after the prune, with the preconditions above. |
--before <ISO-8601> |
Use one explicit cutoff for every class instead of the configured windows. A timezone offset or Z is required and the value must not be in the future. |
What it writes and what it holds
Section titled “What it writes and what it holds”On --apply the command writes one audit.prune row about itself, in the
retained chain (soft) or in the new generation after the reset marker
(hard). Its metadata records the mode, the cutoff and its source (the
configured windows or --before), the per-class cutoffs, the rows removed
and their classes, the generation, and the boundary row's id and
row_hmac; in hard mode also the new generation and the reset marker.
While it works it holds the scheduled verifier's leader lease and the
periodic sweep's advisory lock, so neither can walk or prune the same rows
underneath it. If the verifier is mid-walk the command refuses and says so;
if the sweep is mid-pass it refuses likewise rather than reporting nothing
to prune. On PostgreSQL both are advisory locks; on SQLite the single writer
serialises everything, and another connection holding the writer lock past
the busy timeout is the same refusal (database is locked, exit 1), not a
traceback.
Each batch is its own signed transaction and the audit.prune row is
written last, once every batch is through. A run interrupted in between (a
killed process, a dropped connection) leaves the boundary at the last
committed batch, the chain verifying clean as it stands, and no
audit.prune row. Run the command again: it continues from the boundary,
removes whatever is left, and records itself, with the row count of that
rerun.
Exit codes
Section titled “Exit codes”| Code | Meaning |
|---|---|
0 |
Done, dry run complete, or nothing to prune. |
1 |
Refused: no audit-chain key, chain state missing or not authenticating under the configured keys, the row count or a link in the prefix disagreeing with the signed state, a held lease or SQLite writer lock, a hard-mode precondition on the preview, or a row appended before the reset that the cutoff keeps. The message names the reason and ends with what was changed, which is nothing unless batches had already committed. |
2 |
Settings failed to load, or --before is not an aware, past ISO-8601 timestamp. |
A refusal mid-prune (REFUSING mid-prune) means batches already committed
are signed and consistent and nothing after them was touched; run
z4j audit verify, resolve the finding, and run the prune again.
What verification reports afterwards
Section titled “What verification reports afterwards”After a soft prune z4j audit verify reports the chain clean. The pruned
range is on record in the signed boundary, and --known-head tells the
story exactly:
- a head you exported at the boundary row reports
PRUNE_MATCH(CURRENT_PRUNE_MATCHif nothing has been appended since); - a head from deeper inside the pruned range reports
UNPROVABLEwith no other finding: the anchor is gone, the chain is not broken. Record heads more often than your shortest window and read that result as "the anchor aged out; confirm why", not as tampering by itself.
After an epoch cut every head from the old generation reports UNPROVABLE
for the same reason, and the new generation's genesis has no boundary.
Export a new head after any prune, soft or hard. UNPROVABLE is exit 1
from z4j audit verify --known-head, so a scheduled check that still
anchors on a head from inside the pruned range fails from the next run on;
the command's success line says so, and the
scheduled head export
writes a fresh anchor on its next cadence.
A deletion that did not go through the prune, a row removed by hand or by a
job that bypassed the audit service, still reports as broken: the physical
row count disagrees with the signed count, and the row after the gap fails
its link check. z4j audit prune refuses on that chain rather than
blessing the gap with a new boundary; it only ever advances the boundary
over rows it verified itself.
Related
Section titled “Related”- Audit log for what the log records and how the chain is built.
- HMAC audit chain for verification, the threat model and the off-box head export.
- Settings and environment variables for the retention knobs.
- CLI reference for
audit prunebeside the otherauditsubcommands.