Agent long-poll transport
The long-poll transport is the HTTPS fallback for agents that cannot hold a
WebSocket, typically because a corporate proxy strips Upgrade headers. It
uses two routes under /api/v1/agent/ and the same bearer token, the same v2
frame types and the same per-frame HMAC envelope as the WebSocket; the frame
formats are on the WebSocket protocol page. Both
routes share the frame router with the WebSocket gateway, so projection,
audit and notification behaviour is the same once a frame is accepted.
When an agent uses it
Section titled “When an agent uses it”The agent's transport setting (Z4J_TRANSPORT) accepts auto, ws or
longpoll and defaults to auto. auto is a synonym for ws: the runtime
builds the WebSocket transport and never falls back to long-poll on its own.
Only transport=longpoll selects this transport.
Long-poll has no hello handshake to discover the agent's identity, so the
configuration must also carry agent_id (Z4J_AGENT_ID), the UUID shown when
the agent was minted. The configuration is rejected at load time when
transport=longpoll and agent_id is missing or not a valid UUID. The
transport also refuses a plain http:// brain URL unless dev_mode is set,
because the bearer token would travel in cleartext on every request. See the
environment variables reference.
How it differs from the WebSocket
Section titled “How it differs from the WebSocket”- There is no
hello/hello_ackexchange. Ahelloorhello_ackframe uploaded to/agent/eventsis dropped and counted as accepted so the agent does not loop on it; it carries no HMAC and never counts as liveness. - No
agent_workersrows are registered and the hello-derived worker metadata is not persisted, so the process inventory that the WebSocket builds is absent for long-poll agents. - Liveness is refreshed only when an upload contains at least one signed,
non-handshake frame that passes HMAC verification. That refresh bumps
last_seen_atand promotes an offline agent toonline; it happens even when the frame's downstream dispatch fails. An idle or command-only agent can therefore age toofflineon the dashboard while it is still polling successfully. - Delivery uses repeated HTTP requests instead of an open socket, so command latency and connection visibility differ.
What a long-poll agent cannot do
Section titled “What a long-poll agent cannot do”The brain learns an agent's engines and capabilities only from a WebSocket
hello. A long-poll agent sends none, so its row keeps the empty
engine_adapters, scheduler_adapters and capabilities it was minted with,
and both the brain and the dashboard treat that inventory as unreported rather
than as "no engines". The consequences, by name:
- Scheduler fires buffer. The brain routes a
schedule.firecommand, and replays a buffered one, only to an online agent whose row lists the schedule's engine. A long-poll agent never lists one, so every z4j-scheduler fire for its project lands inbufferedand stays there until a WebSocket agent of the same project that advertises the schedule's engine connects; if none does before the buffer time runs out, the fire ends asbuffer_expired. The agent showsonlinethe whole time, so the symptom is visible on the schedule fire history, not on the agents page. - Scheduler-adapter controls are not routed to it. The controls on the
per-engine scheduler adapters pick an agent by its
reported
scheduler_adapters, which a long-poll agent never reports. - Task commands reach it only when the request names it.
retry_task,cancel_taskandrequeue_dead_letterare admitted for an unreported inventory when the request carries the agent'sagent_id; the agent then checks each command against the retry contracts it states on that poll (X-Z4J-Retry-Contracts) and refuses an engine it has not loaded. Wherever a target is chosen by advertised engine, a WebSocket agent that advertised the engine wins whenever one is online. - Dead-letter listing and the dashboard prefer a WebSocket agent. The
brain admits a long-poll agent as the fallback target for the dead-letter
listing, but the dashboard cannot tell which engines it runs: the agents
table shows
-in the engines column, the dead-letters page offers only the engines some reported inventory advertiseslist_dead_lettersfor (so a project served by long-poll agents alone has no engine to pick there), and task commands pick it last.
What it still does: signed event delivery with the same projection, audit and
notification behaviour as the WebSocket; liveness (online after any
verified non-handshake upload); the agent-side disk buffer with at-least-once
delivery; and delivery on the next poll of any command addressed to it.
Use the WebSocket wherever the network allows it. Leave the agent's
transport setting (Z4J_TRANSPORT) at its default auto, which selects the
WebSocket, and choose longpoll only where a proxy strips Upgrade headers.
A project whose schedules fire through z4j-scheduler needs at least one
WebSocket agent that advertises each engine those schedules run on.
Authentication and session state
Section titled “Authentication and session state”Every request carries Authorization: Bearer <agent token>, the token minted
by POST /projects/{slug}/agents. Verification hashes the
token against every accepted master secret (Z4J_SECRET and
Z4J_PREVIOUS_SECRETS). A missing, malformed or unknown token is 401
invalid agent token; so is a token revoked while a request is in flight,
including during a long-poll wait. When Z4J_AGENT_IP_ALLOWLIST is set, the
source address is checked before the token is read, on both routes and on the
connect probe, as the WebSocket gateway checks it before reading a hello: a
request from outside the list is 403 ip_denied whatever token it carries
and whatever state its project is in. An archived project is refused after
the token authenticates (403 project_inactive), and the refusal writes an
agent.auth.project_inactive audit row naming the agent, the project, the
resolved address and the route path, the row the WebSocket gateway writes for
a hello from an archived project. Every request is refused, but the row is
written once per agent per ten minutes (the longest step of the agent's
authentication backoff): an agent's first refusal writes its row at once, and
an agent that keeps calling inside that window leaves one row per window, not
one per request. The auth.ip_denied row of a refused address is written the
same way, once per address per ten minutes, while
z4j_auth_ip_denied_total{surface="agent"} counts every refusal. The
WebSocket gateway shares the record: an agent or an address refused on a
hello and on these routes inside one window leaves one row across both, and
the gateway's own agent.auth.bearer_failed row is written once per address
per ten minutes in the same way.
Frames are signed with the per-project key derived from the current master
secret alone, bound to the agent UUID, the project UUID and the session
nonce. The brain keeps one signer and verifier pair per
(agent_id, session nonce), so a reconnecting agent gets fresh sequence
state and a forged frame can only poison the nonce it arrived under.
| Header | Direction | Meaning |
|---|---|---|
X-Z4J-Session-Nonce |
request | Fresh random value the agent generates on every connect and sends on every request. The brain keys its signer and verifier state by it and binds it into the HMAC envelope. A request without it falls into one shared legacy session per agent with no nonce binding. |
X-Z4J-Runtime-Features |
request | Comma-separated runtime feature names, recorded on the agent row for operator observability. Read only on the non-claiming probe (max_frames=0); at most 64 names of 64 characters. |
X-Z4J-Retry-Contracts |
request | engine=1 pairs naming the engines whose retry contract the loaded adapters support. Checked on every claiming poll; a missing or malformed header means no retry_task or bulk_retry command is delivered to that request. |
X-Z4J-Agent-Id, X-Z4J-Project-Id |
response | The canonical agent and project UUIDs the agent must sign with. The agent configuration holds the project slug, and the envelope binds the project UUID, so the agent reads these from the connect probe before building its signer. |
X-Z4J-LongPoll-Worker |
response | PID of the brain worker process that answered. The session registry is process-local, so in a multi-worker brain pin each agent to one worker at the load balancer. |
The registry is bounded: 4096 sessions in total and 16 per agent, and an idle
session expires after five minutes. When a new session cannot be admitted, the
brain evicts only an idle session of the same agent; if it has none, the
request gets 503 with Retry-After: 1. A signature failure replaces the
session with an invalidation marker. While the agent has a marker, every
request under any of its nonces is 409; only the non-claiming connect probe
(max_frames=0) under a fresh nonce is admitted, and it retires the markers.
The marker itself expires with the idle timeout.
Both routes sit behind the per-IP bucket they share with the WebSocket
connect: 600 requests per minute, 429 beyond that with the retry delay in
the detail. See rate limits.
Upload frames
Section titled “Upload frames”POST /api/v1/agent/eventsAuthorization: Bearer <agent token>X-Z4J-Session-Nonce: <nonce>Content-Type: application/json{"frames": ["<signed v2 frame as a JSON string>", "..."]}frames holds 1 to 500 serialised, signed frames of at most 1 MiB each
(422 otherwise); the accepted types are event_batch, heartbeat,
agent_status, command_ack, command_result, registry_delta and
error. Each frame is
verified with the session's verifier and then dispatched through the shared
frame router. The response is FrameUploadResponse:
{"accepted": 3, "rejected": 0, "errors": [], "error_code": null}The 200 response is the acknowledgement; there is no event_batch_ack
frame on this transport. The counts follow the dispatch verdict per frame:
| Counted as | Frames |
|---|---|
accepted |
Stored durably, or dropped permanently (re-sending would fail identically), or unparseable, or an unsigned handshake frame. The agent should confirm and delete these. |
rejected |
A transient failure (the agent should re-send), a protocol-version skew (authentic, re-sent so it lands on a matching replica), or a dispatch crash. |
errors carries at most the first ten messages. error_code is
scheduler_upgrade_required when a frame needs a newer scheduler adapter
than the agent runs; that frame counts as rejected. A frame that fails
signature verification rejects it and every frame after it in the batch,
invalidates the session, and the response is 409 rather than 200. A frame
whose dispatch reports the agent as revoked ends the batch with 401.
The agent confirms the batch only when the accepted count equals the frame
count it sent; on any other count it re-sends the whole batch after a backoff, and
the brain's idempotent event identity absorbs the duplicates. A network drop
after the brain committed but before the agent read the response produces the
same replay.
Poll for commands
Section titled “Poll for commands”GET /api/v1/agent/commands?wait=30&max_frames=50Authorization: Bearer <agent token>X-Z4J-Session-Nonce: <nonce>X-Z4J-Retry-Contracts: celery=1| Parameter | Default | Range | Meaning |
|---|---|---|---|
wait |
30 |
0 to 60 | Seconds to hold the response open when nothing is pending. |
max_frames |
50 |
0 to 500 | Maximum commands returned. 0 is the non-claiming connect probe. |
The response is CommandPullResponse:
{"frames": ["<signed v2 command frame as a JSON string>"]}Each entry is a command frame freshly signed by the session's signer, so the
agent verifies it through the same code path it uses on the WebSocket. An
empty list means no command was delivered: the probe, a wait that expired, or
a concurrent poll that won the row. The agent treats all three the same and
polls again.
Connect probe
Section titled “Connect probe”max_frames=0 claims nothing. The agent calls it with wait=0 on connect to
check the token and reachability and to read the identity headers it will
sign with; the brain records the request's X-Z4J-Runtime-Features header on
the agent row at this point. It is also the
only request that clears a 409 invalidation, and it must arrive under a
fresh nonce. If the brain answers the probe with 422, the agent retries it
with max_frames=1.
Wait semantics
Section titled “Wait semantics”A claiming poll first selects what is already deliverable. If nothing is and
wait is above zero, the brain re-checks the commands table every 250 ms
until wait seconds have passed, and returns on the first pass that finds
work. The agent's liveness is re-validated before every pass and before an
empty response, so a revoke that commits during the wait ends the request
with 401 instead of a normal empty poll. Once a poll has claimed at least
one command, a revoke detected before the next claim does not discard that
work: the already-claimed frames are returned with 200, and only a poll
that has claimed nothing gets 401.
Selection, in issued_at order and bounded by max_frames:
- Durable bulk-retry children are claimed first, through the bulk-retry coordinator, when the retry-contracts header covers their engine.
- Commands in
pendingare claimed one row per transaction; the claim is a conditional update, so of two concurrent polls only one receives a given command. retry_taskandbulk_retrycommands are delivered only whenX-Z4J-Retry-Contractslists the engine they target.- Commands already in
dispatchedare re-sent only inside the redispatch window below, only for the redeliverable actionsschedule.fire,schedule.trigger_now,schedule.trigger_now.via_scheduler,cancel_task,reconcile_task,schedule.resync,schedule.enableandschedule.disable. Every other action is at-most-once once dispatched; an unknown delivery outcome surfaces as a command timeout rather than a possible double side effect. A schedule delivery that carries a protocol marker is re-selected only by the session that claimed it.
Each command is claimed before it is signed and appended to the response. If
signing fails after the claim, the command is marked failed with
longpoll sign failed, so the problem is visible instead of waiting out the
timeout. The exception is a schedule delivery carrying a protocol marker: its
claim is immutable, so a signing failure there is only logged and the row
stays claimed. Each frame carries timeout_seconds from
Z4J_COMMAND_TIMEOUT_SECONDS.
Redispatch window
Section titled “Redispatch window”If the HTTP response carrying a command never reaches the agent, the row sits
in dispatched with nothing to deliver it. The long-poll path re-selects such
rows for a bounded window, and re-sends the same row at most once per
interval:
| Setting | Default | Range | Effect |
|---|---|---|---|
Z4J_AGENT_LONGPOLL_REDISPATCH_SECONDS |
60 |
1 to 3600 | How long after dispatched_at a redeliverable command is still treated as a possibly lost delivery and re-selected by a poll. The command dispatcher's re-issue guard uses the same window to tell an in-flight row from an orphaned one. |
Z4J_AGENT_LONGPOLL_REDISPATCH_MIN_INTERVAL_SECONDS |
10 |
1 to 3600 | Minimum gap between successive re-sends of the same still-dispatched command. The re-send is a lease: it fires only when the last send is at least this old. |
Settings load refuses a minimum interval that is not strictly below the
redispatch window, and a Z4J_COMMAND_TIMEOUT_SECONDS below the minimum
interval, because either leaves no window in which a dropped command can be
re-sent.
Both cutoffs are recomputed on every pass of the wait loop, so a row that becomes eligible during the wait is picked up on the next pass rather than on the next request.
Agent-side behaviour
Section titled “Agent-side behaviour”The z4j-bare runtime generates a new session nonce on every connect, sends
the runtime-features and retry-contracts headers on every request, and polls
with wait set to its poll_wait_seconds (default 30, clamped to 1 to 60)
and max_frames=50. A command frame that fails verification makes the agent
reconnect with a fresh nonce. On uploads it treats 401 as a rejected token,
a 403 whose body carries error: project_inactive or error: ip_denied as
an authentication failure (on the connect probe and the command poll too),
413 as a payload that is too large, 415 and 422 as content the brain
refused (a multi-frame batch is split so valid siblings still deliver), and
every other non-200 status as transient, honouring Retry-After when
present.
Status codes
Section titled “Status codes”| Status | Route | Meaning |
|---|---|---|
200 |
both | Frames accounted for, or commands delivered (possibly none). |
401 |
both | Missing, malformed, unknown or revoked agent token, including a revoke that commits during a wait. A revoke detected between claims is 401 only when nothing was claimed yet; otherwise the already-claimed frames are returned with 200. |
403 |
both | The agent's project is archived; the body carries error: project_inactive. The token is still valid and the agent is admitted again once the project is active. The shipped agent treats it as an authentication failure: it logs the code and the message and backs off on its authentication schedule (10 s to 10 min) instead of its connection schedule, on every route including the connect probe. |
403 |
both | The resolved source address is outside Z4J_AGENT_IP_ALLOWLIST; the body carries error: ip_denied and details.surface: agent, the same body as the API surface. Checked once per request before the bearer is read (the connect probe included), as the WebSocket gateway checks it before reading a hello, so from outside the list a bad token, a valid token and a token of an archived project all get this same answer. The token stays valid for the next request from an admitted address; the address goes to the auth.ip_denied audit row (with no agent id, since nothing was authenticated), never to the caller. The shipped agent treats it like project_inactive: the code and the message are logged and it backs off on its authentication schedule. Any other 403 stays transient. |
409 |
both | The session under this nonce was invalidated by a signature failure; reconnect with a fresh nonce and the connect probe. |
413 |
events | Request body above the brain's body-size limit, rejected before the handler. |
422 |
both | Body or query validation: frame count, frame size, wait or max_frames out of range. |
429 |
both | Per-IP agent bucket exhausted. |
503 |
both | Session registry full and no idle session of this agent to evict; carries Retry-After: 1. |