ModelRig

Setup — point your coding agent at it

Point your coding agent at it — it reads the docs, finds any LLM calls you already have, and turns them into routes. That path comes first on every step below: copy the prompt, paste it into Claude Code, Codex, or Cursor from inside your project, and your agent does the step. Prefer your own hands? The exact snippet sits right under it, complete as shown.

8 steps, each checked off by real data this console reads — not by clicking a box. Prerequisites carry a ⓘ — YAMLi, JSON Schemai, environment variablesi, and the CLIi each get one plain sentence and a link out.

modelrig 0.9.0 is on npm (Apache-2.0, Node >= 20) — the commands below install and run. It builds a native module (better-sqlite3), so give it a toolchain. Prepaid balances are the remaining piece: nothing here is chargeable yet. The probe suite, registry and leaderboard stay public and reproducible.

Launch terms

BYOK is free during launch — there is no platform fee on usage routed on your own provider keys while launch terms are in effect. Managed keys are unchanged (provider cost at list + 2%, from a prepaid balance).

Launch terms end when we announce pricing. You get 60 days’ notice, and organizations created during launch keep launch terms for 6 months after the announcement.

Fair-use capNo card · your capsCard on file
Routed BYOK requests / month1,000,00010,000,000
Telemetry retention90 days1 year
Concurrent bake-offs210
Published receipts / month50unlimited
Managed keys—yes (prepaid)

Adding a card lifts your fair-use caps and unlocks managed keys — it never adds a charge during launch. Orgs without a card stay free and capped. Caps are soft: exceeding one shows a notice and this prompt; nothing is blocked mid-run.

Add a card to lift caps →

Step 0 — mint your ModelRig key

One rig_sk_ key is the only ModelRig credential your app ever holds — it names your org, so your telemetry lands in yours and nowhere else. Create one below, copy the MODELRIG_API_KEY=… line it shows once, and set it beside your provider keys. Your coding agent never sees it — it reads the name from the environment you set.

Sign in to create a key. Keys are issued to an organization, so this page needs to know which one you belong to.

No key needed for the open lane: with only your provider keys set, routes, probes and telemetry run entirely on your machine — this console and synced telemetry are what the key unlocks. You can add it later; nothing below requires it.

Instrument your project

The fastest migration is the one your agent performs — that path comes first. Prefer your own hands? The three steps, spelled out.

Paste this into Claude Code, Codex, or Cursor from inside your project. Your agent reads the docs, finds your LLM call sites, and turns them into routes:

prompt — paste into your coding agent
Migrate this project's LLM calls to ModelRig (https://modelrig.dev).
Work the details from the docs yourself; STOP and present to your human at each ⏸ checkpoint.

OPERATING RULES (read first; they override anything you remember):
R1. Ground in live surfaces, never memory: which models are probed, which knobs a
    provider honors, what a claim means — ask https://modelrig.dev/llms.txt, the
    raw-lane support matrix (modelrig.dev/route-bundles.html), rawKnobSupport(),
    or the MCP oracle. What the routing engine does and does not do at run time
    (fall-through triggers, the repair rung, what is stored, what is NOT built) is
    https://modelrig.dev/routing-reliability.html — read its "Not yet implemented"
    section before proposing any behaviour that depends on a route-declared judge or
    raw-lane structure. Declared capability and probed behaviour differ — that gap
    is ModelRig's entire thesis.
R2. Key custody: never create accounts, sign in, or read/transmit/store provider
    key VALUES. Existence checks by NAME only; humans place keys.
R3. Behaviour-identical, or a NAMED change presented at a checkpoint. Never a
    silent change — not sampling, not caching, not grounding, not enforcement.
R4. A ModelRig knob that is absent is byte-identical to before; a declared knob a
    provider cannot honor FAILS CLOSED with a classified error. If you hit one,
    that is the product working — consult the matrix, don't work around it.
R5. If the human's answers conflict with each other or with the code you read,
    STOP and surface the conflict. If they ask for something ModelRig does not
    support, say so directly and propose the closest supported alternative —
    never improvise.
R6. The rig has a lifecycle: create it ONCE (module singleton for a server;
    per-invocation for a one-shot job) and await rig.close() before exit, or
    telemetry is lost and timers hang the process.
R7. Run the project's own tests before proposing any merge; show the diff.
R8. Every pipeline execution is a run. Wrap each execution in run.start /
    run.scope, name one step() per model-call family, and artifact.save the work
    products; tag subject and feature (inside a run context the SDK stamps run_id
    AND step for you — a value you pass yourself always wins). subject is your
    end-customer's opaque id — it powers the per-customer tenant statements on
    /tenants; a tag named client/tenant is only a cost dimension. Acceptance: the run
    appears on /runs with its steps and artifacts — a migration is not done until
    one real execution shows a run with steps there.

0. PERSISTENT GROUNDING (do this before anything else):
   a. Add to this repo's agent instructions (CLAUDE.md, AGENTS.md, .cursorrules —
      whichever exists; create CLAUDE.md if none):
        "Before writing ModelRig code, fetch https://modelrig.dev/llms.txt —
         probed facts change; do not rely on memorized models or parameters."
   b. If your tool supports MCP servers, register the ModelRig oracle and ASK it
      rather than recalling (query_registry, get_leaderboard, explain_pricing,
      get_call_notes):
        { "mcpServers": { "modelrig": { "command": "npx", "args": ["modelrig-oracle"] } } }
   c. Keys: if MODELRIG_API_KEY is not set and the human wants hosted
      telemetry/console sync, STOP and ask them to create an org and issue a
      rig_sk_ key at https://app.modelrig.ai (about two minutes, human-only).
      Without it, proceed in local-only mode; that is fully supported.

1. DISCOVERY-LITE (ask, don't infer, these three; the code answers the rest):
   - What is this project, and is ModelRig replacing existing LLM calls or
     starting fresh?
   - Which providers does it (or should it) use, and which provider keys exist?
   - Hosted telemetry (app.modelrig.ai) or local-only?
   Then read https://modelrig.dev/routing-reliability.html (how a call actually
   behaves and what is not built), https://modelrig.dev/migration-playbook.html
   (this prompt follows its T0-T2 autonomy ladder and checkpoints) and
   https://modelrig.dev/quickstart.html.

2. START AT THE BOTTOM OF THE LADDER, not at routes. T0: point an existing
   OpenAI-compatible client at the ModelRig raw lane by swapping its base URL —
   transport only, nothing about the request or model changes, telemetry lights
   up. T1: run `modelrig observe` to see runs without migrating anything. T2
   (routes, the steps below) is a per-call-site graduation you earn AFTER T0/T1
   prove the plumbing — not the opening move.

3. ⏸ CHECKPOINT — fit and lane. Inventory the LLM call sites (generation calls
   incl. JSON-extractor followups; embeddings and web-search tool calls are out
   of scope). If the codebase already has a router layer, runtime-assembled
   prompts, or env-resolved model choice, do NOT externalize it into routes —
   migrate ONE naive call site as a reference route and integrate the router at
   the rig.runRaw seam instead. Present this as a MIGRATION RECOMMENDATION and
   get explicit approval before writing any code:
     Use case: <one sentence>            Lane: <T0 / T1 / T2 routes / runRaw seam>
     Call sites: <count, by provider>    Keys present: <names only>
     Knobs to carry (each with why): <sampling / cache / grounding / tier / reasoning …>
     Named behaviour changes (if any): <e.g. provider-strict JSON -> own extractor>
     Open questions / assumptions: <what you inferred that the human should confirm>
     Ready to proceed?

4. ⏸ CHECKPOINT — caching, grounding AND SAMPLING inventory, BEFORE writing any
   route or seam call: search for provider caching (cachedContent, cache_control,
   prompt_cache_key, cached-token usage fields), provider-native search/grounding,
   and explicit sampling params (temperature, top_p, max_tokens/max_output_tokens).
   Caching must be carried through (the cache handle on rig.run / rig.runRaw —
   modelrig.dev/caching-lifecycle.html); grounding gaps must be REPORTED as named
   behaviour changes; sampling must be PRESERVED by declaring policy.sampling
   { temperature, top_p, max_output_tokens } on the route (or the same fields on
   runRaw) — a classifier at temperature 0.1 is a deliberate choice, so carry it
   over exactly (absent, adapters use their own defaults; Gemini runs at 1.0).
   Never migrate past any of the three silently.

5. Write routes (T2; install with the project's own package manager — pnpm/npm/
   yarn; Node >= 20): modelrig/routes/<task>.yaml with the prompt in a FILE
   (prompt.system: ./prompts/<task>.md, never inline), schema, candidates,
   policy. Replace call sites with rig.run("<route>", { input, tags }) — tags is
   REQUIRED and every tag key must be declared as a dimension in rig.yaml.
   result.output arrives parsed AND schema-validated — delete your manual
   JSON.parse / fence-stripping. Routes resolve from ./modelrig/routes relative
   to the process CWD: in a monorepo put modelrig/ at the app root and run from
   there, or set MODELRIG_ROUTES_DIR. A runRaw-only seam needs no routes at all —
   construct with routesDir: null. Keep prompt and schema semantics unchanged; if
   the original was JSON-mode + app-side validation, adding schema enforcement is
   an UPGRADE, not identity — present that choice at the prove-it checkpoint, or
   set policy.json: json_mode for closest identity. To load-check the bundle
   WITHOUT any provider key, run `modelrig validate` — it structurally loads
   rig.yaml + every route and reports per-route OK / config error.

6. ⏸ CHECKPOINT — candidates. Any model an adapter can reach may be a candidate —
   probes gate CLAIMS, not serving. Pin what production resolves to today as
   candidate #1, always. Probed models (llms.txt or the public registry) carry
   evidence; an UNPROBED candidate is flagged as such and costs exactly this: no
   json: native claim (the emulation ladder serves structure), no leaderboard/
   bake-off priors, and possibly conservative envelope pricing until the next
   sweep. Never silently substitute a probed sibling for the model production
   runs; offer a probe request or a probed sibling as candidate #2. Present the
   list for the human to ratify.

7. ⏸ CHECKPOINT — prove it: run the project's tests, show the diff and every
   changed file, and after first traffic compare the route's cache-hit rate and
   cost against the pre-migration baseline in the console (this metric needs a
   REAL run — a single probe cannot produce it; say so if it must wait). A zero
   hit rate on a route that cached before is a regression; say so.

8. ⏸ CHECKPOINT — see the run (R8). Open https://app.modelrig.ai/runs: your
   execution is listed, expands to its steps in order, and each step carries its
   inference and saved artifacts; /projects groups the same run by its tags. An
   empty /runs after a real execution means it was not wrapped in run.start /
   run.scope — fix that before calling the migration done. `modelrig status`
   prints "runs recorded: N (last: …)" as the local mirror of the same fact.

9. ATTACH A GRADE, THEN PROVE A CHEAPER MODEL. Capture is not the whole
   loop — a grade is what makes /coverage readiness and /insights findings
   meaningful. A grade is your judgment — from your code, a human, or an AI — a
   0–1 score with an optional comment. If the pipeline already has a quality
   signal (a validator, a downstream accept/reject, a human QA gate), emit it as
   a grade — score 1 on a clean pass, 0 on a rejection:
     rig.grade({ kind: "inference", id: result.meta.inferenceId }, { score: 1, kind: "human" }); // one call
     rig.grade({ kind: "run", id: run_id }, { score: 0, kind: "human", comment: "…" });          // whole run
   The write is synchronous and local, mirrored by the exporter, and never blocks
   or throws into your path. A route with ten attached grades reaches Rung 1 on
   /coverage; a route with no schema to conform to (a prose output) gains proposal
   confidence only this way. Then prove a cheaper model on YOUR own traffic, on the
   lane that fits: a DECLARED route uses `modelrig bakeoff --route <route>
   --replay-last N`; a RAW/run-based pipeline (a rig.runRaw seam, no routes)
   promotes a real output to a golden (rig.artifacts.artifact.promoteToEval(handle,
   { task }), or the console's one-click "make golden") and then runs `modelrig
   bakeoff --from-eval-cases <task> --variants default,<challenger>`. A bake-off
   measures; it never switches what serves a route — that is still a human-approved
   PR.

10. CONTENT CUSTODY IS ON BY DEFAULT — and Optimization on is the recommended
   default. Set `capture: true` on your routes so ModelRig keeps that route's
   telemetry rows and the bounded review samples your bake-offs replay — row-scoped
   in your organization, visible and deletable in the console — repaid with community
   priors and evidence on your own routes at the same flat 2%. Your organization's
   content custody posture is managed by default, so a route that captures content
   lands its prompts, model outputs and evidence in your managed store, scrubbed for
   the PII/PHI you classify, yours to see and export in the console. Opting out is one
   setting, not a default you have to escape: from your account settings (the console or
   the `set_org_settings` MCP tool) set posture `off` to capture nothing, or posture
   `metadata` to keep hashes only, or require `zero_retention` on a route to keep
   nothing there.

The console shows the truth; actions live in your code.

It won't create accounts, sign in, or touch your API keys — those stay in environment variables you set yourself (Step 0). It swaps call sites to rig.run() without changing prompt or schema semantics, runs your tests, and reports every file it changed for you to review.

How ModelRig fits your stack

You keep writing normal code. ModelRig sits between your app and the model providers: it serves each call from the route's candidates, records what happened, and shows you the truth here. Actions stay in your code — nothing swaps a model on its own.

What we keep — your choice, per route

Pure router (capture off)
Nothing retained — the honest answer for a ZDR-strict or data-residency-bound workload. Real and supported, not a downgrade of the routing.
Optimization on (recommended)
Keeps per-inference telemetry and bounded review samples for the bake-offs you create — in your org, row-scoped by the database, visible, exportable and deletable here.

Optimization on is the recommended default: your traffic is what makes your routing smarter and your bill smaller, so giving us the data is the normal, encouraged path — and every telemetry row it keeps, you can see, export and delete right here. Pure router means your traffic never becomes community knowledge: your rows are never folded into the shared model-performance cells, whatever any individual rig's contribution flag says. It does not mean we stop recording your own telemetry — that is what your usage is billed from, and it stays in your organization, visible and deletable here. Optimization on adds your conformance stats to the shared cells. Both lanes cost the same flat 2% — contributing is how the corpus gets better for you, not a discount you trade privacy for.

How calls are keyed & billed

Your keys (BYOK)
Your provider keys stay in your own environment — the exporter has no code path to them. The first 1,000,000 requests each month are free, then a flat 2% of list price.
Managed keys
ModelRig fronts the provider's inference cost, so managed calls are provider cost + 2% from the first request — no key for you to hold.

The same flat 2% either way — contributing is how the corpus gets better for you, not a discount you trade privacy for. Nothing is chargeable yet; prepaid balances land with billing. See your costs →

ArtifactsoffNo runs recorded yet. The run → step → artifact chain records automatically once the app that runs your pipeline has a ModelRig API key — on by default under a control plane (set posture off in account settings to opt out).
First run visibleNot yet. Wrap a pipeline execution in run.start() / run.scope() (one step per model-call family) so it appears here — this is the last step of setup.open /runs ▸

Checklist — 0/8 detected from ModelRig's own traffic — this is the public demo

  1. →1. Create your ModelRig key

    One `rig_sk_` key is the only ModelRig credential your app ever holds — it names your organization, so telemetry lands in yours and nobody else's. It is shown once; rotate any time, and the old one stops working the moment the new one exists.

    Create your key ▸
    let an agent do this step (the fast path) ▸
    Migrate this project's LLM calls to ModelRig (https://modelrig.dev).
    Work the details from the docs yourself; STOP and present to your human at each ⏸ checkpoint.
    
    OPERATING RULES (read first; they override anything you remember):
    R1. Ground in live surfaces, never memory: which models are probed, which knobs a
        provider honors, what a claim means — ask https://modelrig.dev/llms.txt, the
        raw-lane support matrix (modelrig.dev/route-bundles.html), rawKnobSupport(),
        or the MCP oracle. What the routing engine does and does not do at run time
        (fall-through triggers, the repair rung, what is stored, what is NOT built) is
        https://modelrig.dev/routing-reliability.html — read its "Not yet implemented"
        section before proposing any behaviour that depends on a route-declared judge or
        raw-lane structure. Declared capability and probed behaviour differ — that gap
        is ModelRig's entire thesis.
    R2. Key custody: never create accounts, sign in, or read/transmit/store provider
        key VALUES. Existence checks by NAME only; humans place keys.
    R3. Behaviour-identical, or a NAMED change presented at a checkpoint. Never a
        silent change — not sampling, not caching, not grounding, not enforcement.
    R4. A ModelRig knob that is absent is byte-identical to before; a declared knob a
        provider cannot honor FAILS CLOSED with a classified error. If you hit one,
        that is the product working — consult the matrix, don't work around it.
    R5. If the human's answers conflict with each other or with the code you read,
        STOP and surface the conflict. If they ask for something ModelRig does not
        support, say so directly and propose the closest supported alternative —
        never improvise.
    R6. The rig has a lifecycle: create it ONCE (module singleton for a server;
        per-invocation for a one-shot job) and await rig.close() before exit, or
        telemetry is lost and timers hang the process.
    R7. Run the project's own tests before proposing any merge; show the diff.
    R8. Every pipeline execution is a run. Wrap each execution in run.start /
        run.scope, name one step() per model-call family, and artifact.save the work
        products; tag subject and feature (inside a run context the SDK stamps run_id
        AND step for you — a value you pass yourself always wins). subject is your
        end-customer's opaque id — it powers the per-customer tenant statements on
        /tenants; a tag named client/tenant is only a cost dimension. Acceptance: the run
        appears on /runs with its steps and artifacts — a migration is not done until
        one real execution shows a run with steps there.
    
    0. PERSISTENT GROUNDING (do this before anything else):
       a. Add to this repo's agent instructions (CLAUDE.md, AGENTS.md, .cursorrules —
          whichever exists; create CLAUDE.md if none):
            "Before writing ModelRig code, fetch https://modelrig.dev/llms.txt —
             probed facts change; do not rely on memorized models or parameters."
       b. If your tool supports MCP servers, register the ModelRig oracle and ASK it
          rather than recalling (query_registry, get_leaderboard, explain_pricing,
          get_call_notes):
            { "mcpServers": { "modelrig": { "command": "npx", "args": ["modelrig-oracle"] } } }
       c. Keys: if MODELRIG_API_KEY is not set and the human wants hosted
          telemetry/console sync, STOP and ask them to create an org and issue a
          rig_sk_ key at https://app.modelrig.ai (about two minutes, human-only).
          Without it, proceed in local-only mode; that is fully supported.
    
    1. DISCOVERY-LITE (ask, don't infer, these three; the code answers the rest):
       - What is this project, and is ModelRig replacing existing LLM calls or
         starting fresh?
       - Which providers does it (or should it) use, and which provider keys exist?
       - Hosted telemetry (app.modelrig.ai) or local-only?
       Then read https://modelrig.dev/routing-reliability.html (how a call actually
       behaves and what is not built), https://modelrig.dev/migration-playbook.html
       (this prompt follows its T0-T2 autonomy ladder and checkpoints) and
       https://modelrig.dev/quickstart.html.
    
    2. START AT THE BOTTOM OF THE LADDER, not at routes. T0: point an existing
       OpenAI-compatible client at the ModelRig raw lane by swapping its base URL —
       transport only, nothing about the request or model changes, telemetry lights
       up. T1: run `modelrig observe` to see runs without migrating anything. T2
       (routes, the steps below) is a per-call-site graduation you earn AFTER T0/T1
       prove the plumbing — not the opening move.
    
    3. ⏸ CHECKPOINT — fit and lane. Inventory the LLM call sites (generation calls
       incl. JSON-extractor followups; embeddings and web-search tool calls are out
       of scope). If the codebase already has a router layer, runtime-assembled
       prompts, or env-resolved model choice, do NOT externalize it into routes —
       migrate ONE naive call site as a reference route and integrate the router at
       the rig.runRaw seam instead. Present this as a MIGRATION RECOMMENDATION and
       get explicit approval before writing any code:
         Use case: <one sentence>            Lane: <T0 / T1 / T2 routes / runRaw seam>
         Call sites: <count, by provider>    Keys present: <names only>
         Knobs to carry (each with why): <sampling / cache / grounding / tier / reasoning …>
         Named behaviour changes (if any): <e.g. provider-strict JSON -> own extractor>
         Open questions / assumptions: <what you inferred that the human should confirm>
         Ready to proceed?
    
    4. ⏸ CHECKPOINT — caching, grounding AND SAMPLING inventory, BEFORE writing any
       route or seam call: search for provider caching (cachedContent, cache_control,
       prompt_cache_key, cached-token usage fields), provider-native search/grounding,
       and explicit sampling params (temperature, top_p, max_tokens/max_output_tokens).
       Caching must be carried through (the cache handle on rig.run / rig.runRaw —
       modelrig.dev/caching-lifecycle.html); grounding gaps must be REPORTED as named
       behaviour changes; sampling must be PRESERVED by declaring policy.sampling
       { temperature, top_p, max_output_tokens } on the route (or the same fields on
       runRaw) — a classifier at temperature 0.1 is a deliberate choice, so carry it
       over exactly (absent, adapters use their own defaults; Gemini runs at 1.0).
       Never migrate past any of the three silently.
    
    5. Write routes (T2; install with the project's own package manager — pnpm/npm/
       yarn; Node >= 20): modelrig/routes/<task>.yaml with the prompt in a FILE
       (prompt.system: ./prompts/<task>.md, never inline), schema, candidates,
       policy. Replace call sites with rig.run("<route>", { input, tags }) — tags is
       REQUIRED and every tag key must be declared as a dimension in rig.yaml.
       result.output arrives parsed AND schema-validated — delete your manual
       JSON.parse / fence-stripping. Routes resolve from ./modelrig/routes relative
       to the process CWD: in a monorepo put modelrig/ at the app root and run from
       there, or set MODELRIG_ROUTES_DIR. A runRaw-only seam needs no routes at all —
       construct with routesDir: null. Keep prompt and schema semantics unchanged; if
       the original was JSON-mode + app-side validation, adding schema enforcement is
       an UPGRADE, not identity — present that choice at the prove-it checkpoint, or
       set policy.json: json_mode for closest identity. To load-check the bundle
       WITHOUT any provider key, run `modelrig validate` — it structurally loads
       rig.yaml + every route and reports per-route OK / config error.
    
    6. ⏸ CHECKPOINT — candidates. Any model an adapter can reach may be a candidate —
       probes gate CLAIMS, not serving. Pin what production resolves to today as
       candidate #1, always. Probed models (llms.txt or the public registry) carry
       evidence; an UNPROBED candidate is flagged as such and costs exactly this: no
       json: native claim (the emulation ladder serves structure), no leaderboard/
       bake-off priors, and possibly conservative envelope pricing until the next
       sweep. Never silently substitute a probed sibling for the model production
       runs; offer a probe request or a probed sibling as candidate #2. Present the
       list for the human to ratify.
    
    7. ⏸ CHECKPOINT — prove it: run the project's tests, show the diff and every
       changed file, and after first traffic compare the route's cache-hit rate and
       cost against the pre-migration baseline in the console (this metric needs a
       REAL run — a single probe cannot produce it; say so if it must wait). A zero
       hit rate on a route that cached before is a regression; say so.
    
    8. ⏸ CHECKPOINT — see the run (R8). Open https://app.modelrig.ai/runs: your
       execution is listed, expands to its steps in order, and each step carries its
       inference and saved artifacts; /projects groups the same run by its tags. An
       empty /runs after a real execution means it was not wrapped in run.start /
       run.scope — fix that before calling the migration done. `modelrig status`
       prints "runs recorded: N (last: …)" as the local mirror of the same fact.
    
    9. ATTACH A GRADE, THEN PROVE A CHEAPER MODEL. Capture is not the whole
       loop — a grade is what makes /coverage readiness and /insights findings
       meaningful. A grade is your judgment — from your code, a human, or an AI — a
       0–1 score with an optional comment. If the pipeline already has a quality
       signal (a validator, a downstream accept/reject, a human QA gate), emit it as
       a grade — score 1 on a clean pass, 0 on a rejection:
         rig.grade({ kind: "inference", id: result.meta.inferenceId }, { score: 1, kind: "human" }); // one call
         rig.grade({ kind: "run", id: run_id }, { score: 0, kind: "human", comment: "…" });          // whole run
       The write is synchronous and local, mirrored by the exporter, and never blocks
       or throws into your path. A route with ten attached grades reaches Rung 1 on
       /coverage; a route with no schema to conform to (a prose output) gains proposal
       confidence only this way. Then prove a cheaper model on YOUR own traffic, on the
       lane that fits: a DECLARED route uses `modelrig bakeoff --route <route>
       --replay-last N`; a RAW/run-based pipeline (a rig.runRaw seam, no routes)
       promotes a real output to a golden (rig.artifacts.artifact.promoteToEval(handle,
       { task }), or the console's one-click "make golden") and then runs `modelrig
       bakeoff --from-eval-cases <task> --variants default,<challenger>`. A bake-off
       measures; it never switches what serves a route — that is still a human-approved
       PR.
    
    10. CONTENT CUSTODY IS ON BY DEFAULT — and Optimization on is the recommended
       default. Set `capture: true` on your routes so ModelRig keeps that route's
       telemetry rows and the bounded review samples your bake-offs replay — row-scoped
       in your organization, visible and deletable in the console — repaid with community
       priors and evidence on your own routes at the same flat 2%. Your organization's
       content custody posture is managed by default, so a route that captures content
       lands its prompts, model outputs and evidence in your managed store, scrubbed for
       the PII/PHI you classify, yours to see and export in the console. Opting out is one
       setting, not a default you have to escape: from your account settings (the console or
       the `set_org_settings` MCP tool) set posture `off` to capture nothing, or posture
       `metadata` to keep hashes only, or require `zero_retention` on a route to keep
       nothing there.
    
    Focus on this part only: create a ModelRig API key and put it in the environment. Stop there and show me what changed.
    or create one from a terminal ▸
    # Create a key from a terminal (the console's Keys page does the same call).
    # Scopes: ingest, query, grade, decision, oracle-act — grant the narrowest set that works.
    curl -X POST "$MODELRIG_SERVER_URL/v1/keys" \
      -H "authorization: Bearer $YOUR_SESSION_TOKEN" \
      -H "content-type: application/json" \
      -d '{"name":"production exporter","scopes":["ingest"]}'
    # The response carries the secret exactly once. Store it; it is hashed at rest.

    The console shows the truth; actions live in your code.

    How you’ll know it worked: the new key appears at the top of the API keys list, and this step ticks green here.

  2. ○2. Install & connect

    Needs an API key, so telemetry has somewhere to land first. Do “Create your ModelRig key” ▸

  3. ○3. Define your first route

    Needs finish "Create your ModelRig key" first first. Do “Create your ModelRig key” ▸

  4. ○4. Run your first call

    Needs a route to call first. Do “Define your first route” ▸

  5. ○5. Prove a cheaper model

    Needs at least one recorded run to replay first. Do “Run your first call” ▸

  6. ○6. Attach a grade

    Needs finish "Create your ModelRig key" first first. Do “Create your ModelRig key” ▸

  7. ○7. Cap a run's spend

    Needs finish "Create your ModelRig key" first first. Do “Create your ModelRig key” ▸

  8. ○8. Tag your costs

    Needs finish "Create your ModelRig key" first first. Do “Create your ModelRig key” ▸

Stuck on a step and out of ideas? open an issue — a human reads them ▸

The loop, in three moves

  1. See. Every call writes a telemetry row locally first — the run never blocks on the cloud — then an async exporter mirrors it, plus the run → step → artifact graph, into this console. When someone asks “why did the pipeline do that?”, you pull up the run and show them.
  2. Prove. Replay your own recent traffic through route variants — a cheaper model, a different prompt — and get the effective cost of conformance for each, with confidence intervals, on Bake-offs. The registry feeds routing from the side: a probed fact beats a declared claim.
  3. Save. A winning variant is a proposal for you to apply in your route YAML — nothing swaps a model automatically. Your routes were always YAML in your git, so leaving is a git rm.

The console shows the truth; the actions live in your code. Full architecture: how it fits ▸

More capabilities

Beyond the checklist, ModelRig also gives you per-customer tenant statements, shareable published receipts, a metadata-only analytics query API, provenance-first assembly & trust, and pairwise, regret-bounded bake-offs. The whole surface, one line each: how it fits ▸