Turning it on
Outcome measurement is the space’smeasurement feature, off by default and turned on in the features document by the security role. The definitions live in the measurement document, of the analysis role: the revenue definition (gross or net of discount, coupons, shipping, taxes, cancellations, exchanges, split orders, currency), the outcomes (an object type, its success states, the revenue field, its lines with the item, the quantity, the state and the exposure token stamp, the handoff field, and final_after), the attribution (methods, click and view windows, the reconciliation tolerance), the experiments, the minimum of people for unmet demand and the cost per turn with memory off your company measured. The routes need the analysis role.
From what the agent showed to the order
An agent shows a list of cards; the customer buys one of them, in the store’s own checkout. To tie the order line to the exact card the agent showed, the card carries a short exposure token, and the store’s app copies it, as opaque text, into a property of the cart line. The order arrives with the token, and the memory attributes the line to the exposure and the position, deterministically, without the app knowing anything else.exposure_token(exposure_id, pos) in Python, exposureToken() in TypeScript), and the interface puts it on the card. It carries only the exposure id and the position: nothing about the customer, the item or the price. The verifier, the first 4 characters of the SHA-256 of the token’s head, catches a token copied wrong; it is not a signature, and the memory accepts a token only for an exposure its space recorded, at a position that exposure had. A line whose token does not check is attributed by the other methods, never by this one.
The methods
Each outcome link (
GET /v1/measure/outcomes) carries the method, the band, the line’s value whole or none at all (in the currency’s minor units, only under the signed revenue definition), the state of the object or its line, when the outcome first counted and its finality: provisional until the exchange window ends (final_after) or the object reaches a final state of the lifecycle, whichever comes first, or expired_without_outcome when the type’s deadline passed first. An identity link carries the method’s confidence in this space: how often it names the exposure the line’s own token names, on the lines where both exist. It stays beside the value, never multiplied by it.
GET /v1/measure/attribution adds up by day, agent, method and band, in UTC days, deterministic and probable apart. An outcome that comes through a handoff appears as the receiving side’s handoff.outcome action, and the commitments and outcomes of coordination enter the memory as actions on their objects.
Reconciling with your BI
The number that counts is yours.POST /v1/measure/reconcile/uploads reserves the upload of your BI’s CSV for a period (order, line optional, value in major units as the signed definition nets it, currency), read once and deleted; POST /v1/measure/reconcile compares line by line the attributed lines with the file’s, on the lines both sides know, and returns the deviation (|ours - BI| / BI), the lines that differ most and, apart, the ones only one side knows, which never enter the deviation. The tolerance lives in the measurement document.
Experiments
The space can measure the effect of an element or of an agent by arms, in themeasurement document: an element experiment drops from one arm the constraints block, the inferred soft signals, the inferred and outcome-derived sizes, the state block or a pack section (facts, episodes, actions or objects); an agent one compares agents; an interleave one interleaves two rankings of a tool. Arms are drawn like the memory’s control group, by the HMAC of the customer’s oldest handle, so the same person stays in the same arm across every channel and vendor. Critical sources never go to a control arm, and the memory’s control group gets no block at all.
GET /v1/measure/experiments compares each arm with control on each metric (conversion and value per person), from what people did after they were assigned: incremental, as only a control group says. The difference is adjusted by CUPED, with the days before each person’s assignment as the covariate (cuped_days), when the history varies, and the report says how much of the variance the adjustment removed. The test is the mixture sequential test (mSPRT, with the document’s mixing variance): the p-value and the interval hold at any look, so the report may be read every day without inflating the error; decided says when the difference is real at the test’s alpha. The comparisons stay empty until each arm holds enough people for the normal approximation. POST /v1/measure/power says how many days of the space’s traffic a relative difference needs to be detected, looked at once at the end or every day, from the space’s history or the numbers you give.
Interleaving
When two rankings of a tool are interleaved (a team draft, seeded by the turn),team on each item of the exposure says which one contributed it, and engagement and cart go to the list that contributed the item. GET /v1/measure/interleaving counts, by day and over the period, the impressions engagement credited to a (the ranking in production) or b (the candidate) and the ties, with the two-sided sign test over the decided impressions.
The tool counterfactual
Before running an experiment on an element of the constraints block, a smaller question: does the element change what the tool returns at all? A search that ignores a filter, or a filter that removes nothing, makes every later measurement of that element a measurement of nothing. The tool counterfactual answers from recorded turns: the Python SDK’sniadra counterfactual command runs each recorded call of the tool live, inside your company, three times at the same moment (twice with the element, to measure the tool’s own noise, and once without it), and reports positions and overlaps only, never items, arguments or results. A tool that writes state runs as its dry run, and without a dry run it is never called.
overlap@k, depth-weighted with the same exposure weight as the signals (1 / log2(pos + 1)); k is the recorded list’s visible_k, or 10. The report (POST /v1/measure/counterfactual-runs, read at GET /v1/measure/counterfactual-runs/{run_id}) carries the mean overlap, the noise floor, the effect (noise_floor - overlap), the sign test, where the items the person engaged with went (kept, lost, gained, the mean shift) and the skipped cases. It also states its own limits: not_quality (it says whether the list changed, not whether it got better), model_reaction_not_measured, trivial_for_hard (a hard constraint changes a list by construction), few_cases (fewer than 30 completed), noisy_tool (a noise floor below 0.9), cases_skipped and dry_run. Niadra keeps the report, never the turn and call ids.
Unmet demand
GET /v1/measure/unmet-demand lists what people asked the tools for and did not get: by week and by combination of what was asked (on the type registry’s fields, values normalized), the calls that came back empty and the ones that came back with one or two results, and how many distinct people asked. A combination fewer people than the space’s minimum asked (10 by default) is left out: never a person.
Cost with and without memory
With turn records on, context use brings incost, per source and agent, the calls, tokens and money per turn, against the cost per turn with memory off your company measured and declared in memory_off_costs of the measurement document; difference_usd above zero costs more. Niadra adds up what the records say; the baseline is yours.
Privacy
Outcome measurement reads what the cell already has, and the result is about the agent, the tool and the order, never about the subject: the reports carry counts, aggregate values and object ids, and unmet demand only above the minimum of people. Outcome links stay 13 months. Erasing a customer removes their rows by the same lineage, and a report already computed is not redone.Next steps
Signals and constraints
the exposure the exposure token is born from.
Object types and state
the lifecycle that gives an order its final state.
Context use
the measurement that says whether the agent used what it received.
Attributed outcomes
the reference of
GET /v1/measure/outcomes.
