Before you start
- The
turnsfeature on in the space, and turns recorded instoredorpointermode, withgoldfidelity (the SDK, not OpenTelemetry) and the pins therecordingdocument requires (promptsandmodelby default). Ahash_only,bronze,partialturn or one missing a required pin is kept, but not replayable: the viewer andGET /v1/turns/{turn_id}say why inreplay_blockers. - A key with the
replayscope, on a source of the CI’s own; the routes also take a person with theintegrationorsecurityrole. - Tools recorded with
@Niadra.tool(Python) orNiadra.tool(TypeScript): that is how the runner intercepts them.
1. Pick the turns
A scenario keeps up to 50 turns and up to 50 assertions. The turns are copied out of their tier when the scenario is created and stayscenario_days after its last run (180 days by default, from 7 to 365); a turn that left storage answers replay_expired.
Three ways to create one:
tool_called for each tool that answered ok, hard_respected when the turn reported applied constraints, no_denial_with_results when a result showed items, claims_traced and claims_match_state when there were claims, effect_once when there were effects, handoff_when when the turn created a handoff and budget with twice the recorded calls. From a report, the description’s words (a denial, a price or a deadline, a promise, something done twice, a human) suggest the matching assertion, and the description is read and never stored. A suggested assertion comes with suggested: true; adjust, drop or add through PATCH /v1/scenarios/{scenario_id}, which adds one to the version.
2. The assertions
Each assertion is{id, kind, args, turn_id?}: with turn_id, it applies to that turn; without, to every turn. The runner evaluates each one on the replayed turn’s record and the text it emitted, and the outcome is pass, fail or not_checked, which is never a failure:
3. Run it in CI
The SDK’s runner asks for each case, checks the pins, fetches the blobs by digest, runs the agent with the tools answering from the record and reports the run; Niadra returns the verdict.--agent names a function that makes a fresh agent per execution, which takes the case’s input and returns the text it emitted; --build names the running build, the same pins the production turns pin.
What happens in a run
- The pin check. For every pin outside
vary, the running build and the recorded one must be equal when the recording requires the pin, and when both sides carry it. A difference answers 422pin_mismatch, withpins: [{name, recorded, running}], and no case is served. To test a new prompt on purpose, name the pin:--vary prompts. - The values. In
storedmode each blob comes with its kept, masked content; inpointermode the runner reads eachpointerwith your credentials, through the SDK’s content resolver, and checks the digest. A blob that cannot be read or does not match is an infrastructure error, counted apart and never an assertion failure. - The tools. A call whose normalized arguments match a recorded call’s
args_hashanswers with the recorded result, the calls of one tool taken in order. One that matches none runs live only when the tool was markeddry_run=True(dryRun: true); otherwise it answers empty and counts as divergent. A framework tool not wrapped withtool()never runs live in a replay. - What the agent reads and writes. The pack of the time is the blob of the record’s
packread; working memory starts empty and its writes stay with the runner; a coordination check answers from the record. The replayed turn’s record stays with the runner, markedsynthetic, and is never sent as a turn. - The modes.
hermetic_turn(the default) runs the turn alone, with the recorded pack and every tool answering from the record;hermetic_conversationruns the scenario’s turns in order, until the first strong divergence;era_memoryruns with the pack of the time and the current state of the objects the tools read. - The input. Each case carries the customer’s message that opened the turn, masked, and the conversation’s history up to it (at most the last 50 public messages of the previous 7 days).
--runsruns each turn N times, numbered from 1; aparaphrasefunction rewrites the input of some executions, so an intermittent result is confirmed with other words.
4. Read the verdict
Per assertion, over the executions that completed and in which it applied (not_checked left out): passed and failed, the baseline (the same assertion in the scenario’s latest earlier run whose verdict was pass or flaky; without one, the recording, which passed as many times as the run checked), drop (the baseline’s pass rate minus the run’s) and p_value (the one-sided Fisher exact test that the run fails more than the baseline). regression is failed >= 2, drop >= 0.2 and p_value < 0.05; flaky is a failure short of a regression.
Per scenario: pin_mismatch when any execution stopped at the pins; infrastructure_error when none completed; regression when any assertion regressed; flaky when any wavered; pass otherwise. needs_paraphrase says an assertion both passed and failed across the same executions: run again with paraphrases before trusting it. The run’s verdict is the worst of its scenarios’, in the order pass, flaky, infrastructure_error, pin_mismatch, regression, and the command exits with 1 only on a regression: flaky results are listed and never block.
When an execution fails an assertion that reads what the tools showed, Niadra checks the recorded turn’s observations against the declared types and opens a data issue for the data’s owner (null_field, out_of_vocabulary, stale_source), with up to 50 references and no value: a regression that belongs to the data, not to the agent, reaches whoever can fix it.
5. Bisect with the change log
GET /v1/changes lists what changed between consecutive builds of each agent, newest first: prompt (one prompt’s version, by name), model, code (the assembler or a tool’s schema), data (the corpus digest), config (a section of the space’s configuration the turn read, by its digest: the claim contract, the types, the bindings) and compiler (the context compiler’s version), with before and after and when the new build appeared. A regression between two runs points at the changes between their builds.
Privacy
The case carries the conversation’s history masked and, inpointer mode, no recorded value: the runner reads them inside your company. A scenario’s turns are personal data with the purpose quality, stay scenario_days after the last run and are erased with the rest of the person’s turns. A run’s results hold verdicts and counts; error and detail never carry personal data.
Next steps
Turn records
what a turn records and why it is replayable or not.
Metadata only
replay with the values in your bucket.
Ask for a replay case
the reference of
POST /v1/replay/cases.Outcomes and attribution
the tool counterfactual, which runs through the same runner.

