ModelRig

Concepts

The words ModelRig uses, and the loop they describe — See, Grade, Improve. Every ⓘ tooltip in the console points at a term here; this page is where they all live.

see where this fits ▸ /setup

The loop

See

Observe every model call your product already makes — its prompt, output, cost and latency — in one console.

Grade

Judge quality: attach a grade — a 0–1 score plus an optional comment, from your code, a human, or an AI — and the console shows what's working and what isn't.

Improve

Bake off a cheaper-or-better model on your own replayed traffic, and when the evidence holds, swap it — instant, versioned, reversible, from the CLI and committed to git; the swap is a pull request, never a surprise.

Framing words

Words that describe the loop itself rather than one table column.

lane
Three separate axes once all read as “capture” on this console — keep them apart. (1) Content custody is your org’s artifact_posture (managed / metadata / off, in Account settings): it decides whether the bytes of prompts and outputs are stored at all. (2) Optimization is a route’s capture flag: it decides whether runs are kept as replay samples for a bake-off — “Optimization on” keeps samples, off is a pure router that keeps none. (3) The verification lane is a call’s tags.lane: “raw · unverified” marks a BYOK passthrough that skipped the validate/repair pipeline, versus a verified route run. All three are orthogonal — a call can be verified, on an Optimization-on route, under managed custody, each set independently.
incumbent
The model a route serves today — the thing a bake-off measures every challenger against.
challenger
A variant run against the incumbent on your replayed traffic, to see whether it is cheaper or better on the same inputs.
ladder
The evidence ladder a route climbs before optimization is safe: rung 0 (implicit signals only), rung 1 (enough of your attached grades), rung 2+ (judges, a later phase). See the rung entry below.
regret
How far a routing choice can fall short of the top-scoring option available; the optimizer only proposes a swap inside a regret bound you set. Distinct from switching regret — the risk that changing models at all makes things worse — which is why every swap stays your decision.

Glossary

Every term the ⓘ tooltips define, grouped. The definitions here are the same ones the tooltips show.

Keys & credentials

What your app holds, and whose rows its calls become.

API key more ▸
A rig_sk_ secret naming your organization — the only ModelRig credential your app holds.
scope more ▸
One thing a key is allowed to do. A key can only ever do what it was granted.
credential scope
Which key a claim holds under: our managed account, your own (BYOK) key, or any — they differ.
project
A sub-scope inside your organization. A key bound to one stamps it on every row it ships.
ingest more ▸
The endpoint your exporter posts telemetry to. Your key decides whose rows they become.
rig
Your project's ModelRig setup as a whole; later, a published endpoint others can call.
rig manifest
Your project's settings file: name, cost dimensions, sharing preferences.
last used
When this key last authenticated. Updated every few minutes, not on every request.

Traffic & calls

The rows the See pages read — a run, its steps, and every call underneath.

route more ▸
A task your app asks ModelRig to run — a small YAML file naming prompt, schema, and allowed models.
candidates more ▸
The only models allowed to serve this route — nothing else can be billed.
version
Bumps when the route's YAML changes; telemetry records which version served each call.
inference
One model call your app made through ModelRig — one row here per attempt.
served tier
The provider capacity class that actually ran it — flex is discounted but sheddable; compare with requested.
requested tier
The capacity class the route asked for; the served tier shows what you actually got.
failure class more ▸
Typed reason a call failed — retries are budgeted per class, so one flaky network can't eat your content retries.
attempts
Retry counts by failure class for this call.
latency
Wall-clock time for the call, including provider queueing.
run
One execution of a pipeline your app instrumented — an InferWealth report, a batch job.
step
One stage of a run. A gateway-routed step has a ground-truth row; others still produce artifacts.
artifact
A work product a step produced — prompt, raw response, parsed output, evidence — stored as metadata + a hash.
provenance
How a step is grounded: ground-truth (a real gateway inference) or attested (a foreign trace span).
grade
a judgment on a call, run, or artifact: a 0–1 score plus an optional written comment, from your code, a human, or an AI model. The pass/fail verdict is derived from the score (against the route's threshold), never stored separately.Not: not the derived pass/fail — that's the verdict, computed from the score; and not a bake-off's score of a candidate against a golden (a separate system).Interacts: attach one with rig.grade(subject, { score, kind }) — subject is an inference, a run, or an artifact; ten grades on a route reach Rung 1 on /coverage, and /insights clusters on a route's low-scoring grades.
content hash
sha256 of the artifact's bytes. Proves integrity; the bytes themselves are not stored this mission.
taint
Everything that derives from this artifact — its downstream closure over lineage + supersession. If this one is bad, these are the ones to re-check (condemn or clear).
route mirror
The copy of a route's resolved bundle the hosted gateway serves. Git is the source of truth; this is a cache, stamped with when it last synced.
bundle sha
A content hash of the resolved bundle. When it changes, the route changed — the console flags a mirror whose bundle has fallen behind your repo.

Quality & conformance

Whether an output is actually usable — not just valid JSON.

conformance
Output actually matched the required schema — not just valid JSON.
no-failure share
the fraction of a Group-by slice's calls with failure_class === null, over routed and raw calls alike.Not: not schema-conformance: raw calls have no schema to conform to, yet still count here as 'no failure'.
run success
The run passed its run-outcome@v1 grade — the pipeline finished with an accepted outcome, not just no error.
value accuracy
Whether the values are right, not just schema-shaped — valid JSON can still be wrong.
grounded rate
Share of a model's probed outputs that grounded their answer in retrieved evidence — probed, not self-reported.
native rung
Whether the provider enforces your JSON Schema natively (server-side) — the probed native-schema pass rate. Unrelated to the optimization 'rung' ladder.
repair rate
How often an invalid output needed the repair pass before conforming — repair cost is shown, never hidden.
ci95
The 95% confidence interval — where the true rate plausibly lies given this sample size.
samples
How many calls this statistic is computed from — small n means wide intervals.
fingerprint
An anonymous shape-signature of your schema — never its contents.

Cost

What a call really costs, and how spend is grouped, capped, and attributed to a customer.

dimension
A label you attach to calls so costs can be grouped by it.
declared vs observed
Declared = named in rig.yaml with a label; observed = any tag we've seen in your data — both work.
required
Warns when calls omit this tag, so cost reports don't silently under-count.
Reserved
Well-known tag keys (subject, feature) always offered as cost dimensions — even before a call has carried one.
subject
the tag that attributes a call to one end-customer: tags: { subject: customerId }. On first sight it becomes a tenant on /tenants.Not: not a plain cost dimension like client or feature — those group spend but do not create a tenant statement.
tenant
One of your end-customers — derived from its subject tag (or a project-scoped key), the unit a statement is per.
tenant statement
A per-end-customer view derived on read: attempted · verified · billable, with the true cost measured. Never stored.
verified outcome
A call whose output passed deterministic conformance — the only thing billed. Judge and human grades are context, never billed.
unattributed
The always-visible row catching usage with no subject tag and no project-scoped key — spend not yet tied to a tenant.
cached tokens
Input the provider had cached — billed at a discount.
cache hit rate more ▸
Share of this route's input tokens the provider served from cache — its own reported usage, replays excluded.
effective input rate more ▸
What input really costs per token once cache reads and write premiums are counted — not the list price.
effective cost of conformance more ▸
What it really costs to get 1,000 usable outputs — repairs and retries included.
category percentile
This route's effective cost ranked against your routes with the same schema shape — 0 = cheapest.
unpriced
The model has no pricing entry: cost shows an honest $0 and it is never ranked as cheap.
envelope more ▸
A spending cap for a unit of work — spend hard-stops at the cap with a typed error.
budget_exhausted
The typed error a call gets when its envelope's cap is hit — nothing further spends.
utilization
Spend so far against the envelope's cap.
verified savings
(baseline cost − actual cost) × conformant outputs since the swap — measured from telemetry, never projected.

Optimization & proof

How a cheaper-or-better model is proven on your own traffic before you swap. Here “capture” is the route's Optimization/replay flag (keep runs as replay samples) — a different axis from content custody (are bytes stored) and from a call's raw-vs-verified lane.

bake-off more ▸
Replaying your own captured traffic through model variants to compare cost and conformance.
variant more ▸
A named alternative serving setup for a route — different model, prompt, or repair policy.
arm
One contestant in a bake-off — the incumbent you run today, or a challenger variant being measured against it.
sample input
One replayed captured input; the incumbent and each challenger are run on the same input so you can read their outputs side by side.
sample
A kept incumbent-vs-challenger output pair for one input, stored only when a bake-off ran with `--keep-outputs`; bounded and expiring.
capture more ▸
Recorded inputs+outputs for replay, written to the local SQLite `captures` table — the telemetry exporter has no code path to it. Turning capture on is Optimization on for that route: the bounded review samples a bake-off keeps do reach your organization here, where you can read, export or delete them.
replay
Re-running captured inputs through a variant offline — real calls, zero production exposure.
proposal
A variant that beat the incumbent's gate. A proposal only — swaps are your decision, in your code.
evidence
The measured bake-off attached to a proposal — conformance and effective cost on your own traffic.
trigger
The watched event that raised this proposal — a price change or a new model on a provider you use.
watch
The daily cycle that diffs model prices and new arrivals against your routes, raising proposals.
expires
Proposals age out after 30 days — decisions ride current data, never stale suggestions.
verdict
the pass/fail a grade resolves to: pass = score ≥ the subject's route threshold (default 0.5).Not: not the grade itself (the grade is the 0–1 score) and not a stored field — it is derived at read time.Interacts: coverage, /insights, and the console's pass badge all derive it from the one stored grade score, so they never disagree.
episode
One run of a job (your run_id) — a grade can target the whole episode instead of one call.
optimization coverage
How much evidence each route has for safe optimization: implicit signals (rung 0) plus your grades (rung 1).
rung
The evidence ladder: 0 = implicit signals only, 1 = enough attached grades, 2+ = judges (later phase).
registry
What we've verified about each model — declared claims, probed measurements, observed traffic.
declared
What the provider's own documentation claims the model can do.
probed
What our published, reproducible tests actually measured — dated, with confidence intervals.
observed
What live routed traffic has shown — aggregates only.
discrepancy
A declared claim that didn't hold up when probed — shown, not hidden; that's the point.

Custody & screening

What ModelRig holds (content custody — the bytes-at-rest axis, distinct from a route's Optimization/replay flag and a call's raw-vs-verified lane), and how untrusted content is handled before dispatch.

content custody
The org-level artifact_posture your account is on — the single, server-enforced control over what ModelRig captures, set in account settings (the console or the set_org_settings MCP tool). Three values: managed (signup default — full inputs and outputs stored, PII/PHI scrubbed), metadata (counts and hashes only), off (capture nothing — the server drops even metadata at ingest).Not: Not a deployment environment variable you set. The customer opt-out is this account posture, not a MODELRIG_ARTIFACTS env var — one API key turns everything on and you dial it back per-org from account settings.Interacts: It is the master gate above every route: a content-enabled route stores nothing while the org is on metadata or off. A zero_retention route keeps nothing even under managed.
zero retention
Opt in — require: zero_retention on a route, or zeroRetention on a single run — to route only to zero-retention–designated endpoints under our managed account (a BYOK key is scoped separately). Fail-closed: no designated candidate, no dispatch.
screening more ▸
Isolating untrusted content as data and flagging known injection shapes before dispatch. It does not prevent injection — no detector can.
spotlighting more ▸
Rendering untrusted content (search results, third-party variables) into a JSON-encoded user-slot block, never the trusted system prompt.
quarantine more ▸
A flagged turn is dispatched with tool calls barred (tool_choice:none, or dropped and counted) — the untrusted turn can't actuate a tool through ModelRig.
canary more ▸
A unique token seeded into the system slot; if it appears in the model's output, the receipt counts a canary leak (a prompt-echo signal).
screened share more ▸
Share of a route's production samples that ran through a screen: policy. Counts only — never a matched value.

Billing

The money vocabulary the /billing page shows.

prepaid
You fund a balance up front; usage draws it down. When it's empty, requests stop — the balance is its own spending cap.
balance
Money you've funded, minus what usage has drawn. Recomputable from every statement entry — never a guessed counter.
available balance
The part of your balance you can spend right now. Frozen funds are not part of it.
frozen balance
Funds held aside during a payment dispute — not spendable, and not ours. A lost dispute keeps them frozen; a won one releases them.
top-up
Adding funds to your balance on Stripe's hosted checkout. Your balance credits when the payment clears.
statement
Every ledger entry, newest first — top-ups, draws, fees, refunds. The running balance after each one is shown.
provider cost
What the model provider charged for a call, at list price — drawn from your balance on managed keys.
metered fee
Our flat 2% of metered usage at provider list prices — the only thing we keep. Stripe's processing cost passes through separately.
fee credit
A credit back to your balance — a promo or a correction. Rare; shown for a complete record.
refund
Unused balance returned to your card. Always refundable — just ask; consumed usage is not.
adjustment
A manual correction to your balance — e.g. dunning after a failed off-session charge. Signed, and always explained in the note.
card on file
A saved payment method for the monthly metered fee, entered on Stripe's page. We never see the card number.
dispute
A chargeback your bank raised on a payment. It freezes the disputed funds while we respond with your usage record.

Foundations

Prerequisites a first-timer may want a one-line reminder of.

YAML more ▸
A human-friendly config format — indentation defines structure.
JSON Schema more ▸
A contract describing what valid output looks like.
environment variable more ▸
A named value your shell passes to programs — how keys and settings reach ModelRig without living in code.
CLI more ▸
The command line — you type a command in a terminal and the program runs.