---
name: openrouter-decisions
description: Find where an app or agent should use a decision model (a prompt-and-parse LLM call, a keyword or similarity heuristic, a human review queue, a misfiring router, filter, or classifier) and implement it through OpenRouter's Decisions API, which returns probabilities instead of generated text. Covers routing, classification, moderation, guardrails, grounding checks, scoring, ranking, dedupe, escalation and approval gates, bounded extraction, and thresholds. Use when decision models, Jev, TypeSafe, or System One come up. Not for chat completions, text generation, or general OpenRouter setup.
---

# OpenRouter Decisions

A decision model reads application state and answers typed questions with probabilities, never text. The model judges, code computes. Code owns the workflow and every deterministic step, and the model fills in the judgments code cannot make. It is not a chat model, a coding agent, or a replacement for the LLM that drives one.

Every decision model on OpenRouter speaks the same Decisions API and the same three primitives, evaluated independently and in parallel within one request.

| Need | Primitive | Returns |
| --- | --- | --- |
| One option from a defined set | `choice` | `choice`, `probabilities` per option, `confidence` |
| Whether a condition holds | `noul` | `noul`, the probability of yes |
| Degree along ordered levels | `score` | `score` (probability-weighted position), `probabilities` per level, `confidence` |

Read [references/decisions-api.md](references/decisions-api.md) when writing a request or reading a response. Read [references/models.md](references/models.md) when picking between the decision models the live catalog returns. Read [references/decision-model-limits.md](references/decision-model-limits.md) whenever a question involves quantities, dates, negation, indirection, or untrusted text, since those are the things a decision model does not do.

## Steps

1. **Find the decision points.** Read the code path and list every place that turns unstructured input into a bounded outcome. The usual forms: a chat-completion call parsed into a label, boolean, or number, a keyword, regex, token-overlap, or embedding-similarity heuristic standing in for a judgment, a retrieval step whose similarity threshold decides whether the top result is relevant (retrieval stays in code, relevance is the judgment), an if-chain over free text, a human review queue that mostly confirms the obvious, and a step that picks or reranks candidates. A spot fits when the answer is bounded, needs judgment rather than computation, and a probability would let code act, defer, or escalate. It does not fit when the step must produce text, values, or code, or when the answer is already in a field code can read. Done when each candidate has a one-line description of the judgment and the action code takes on the answer.

2. **Split judgment from computation.** Arithmetic, counting, date ordering, string and pattern matching, lookups, and threshold comparisons stay in code, and a value code already holds is never re-asked of the model. When code-side facts settle the action for an input (a path rule, a hard policy, an amount over a cap, an empty input), code returns that action and skips the model for that input. When code parses a value out of free text before comparing it, write the parser against the formats in real inputs and treat a parse miss as its own outcome. Generation goes to a generative model, and a decision model can then pick among its candidates. Done when every remaining item is a judgment that fits one primitive.

3. **Pick the primitive per judgment.** Mutually exclusive alternatives are a `choice`, which is relative and settles which option. An independently testable condition is a `noul`, which is absolute and can be low for every label, so labels that can co-occur are one `noul` each. A degree along one dimension is a `score` whose levels each describe a situation that stands alone. One unordered label from a written rubric is one `choice` whose criteria restate the labels, not several `noul`s recombined in code. Ordered levels (a 1 to 5 rating, low to critical) are a `score` whose criteria restate the levels in order, even if code stores the result as a label. Split into `noul`s only when code combines them with facts of its own or reuses them elsewhere. Comparable per-item `score`s ranked in code give graded ranking. Done when each judgment has a primitive and a reason.

4. **Build the state.** Send only what the questions read. Source text, identities, relationships, policies, and current facts, each in a named JSON field, filtered and retrieved in code first. Anything only code uses stays out (flags, counts, computed booleans, and any field that feeds a code-side rule such as a plan tier that caps a score), and code applies that rule to the returned answer. Unrelated material costs accuracy. Done when every field in `state` is read by at least one question and nothing a question needs is missing.

5. **Write the questions.** `instructions` states the judgment and carries the task's own inclusion and exclusion rules (what counts, what looks similar but does not). `criteria` defines the possible answers and reads as an extension of the instructions, asked one way with no double negatives. Ask about the fact, not the text. "Is this change breaking" holds where "Does the description mention breaking changes" and "Does the report state that" both fail, so phrase each question as a property of the thing judged ("Is this review promotional") and let the model infer from the state. Point at a nested field with a backticked path such as `` `ticket.messages[0].text` ``. A `choice` lists every viable option plus a no-match option such as `none` or `not_stated` whenever nothing may fit. When the options are runtime candidates, key the state by option name and name the option in its criterion, never an array index. Question keys are for code and are not sent to the model, so each question carries its full meaning. Independent questions over the same state go in one request. A second request is needed only when an answer decides what state to fetch or which options to offer next. Done when a careful person could answer every question given only `state`, `instructions`, and `criteria`.

6. **Pick a model and call the Decisions API.** The set of decision models changes, so never pick from memory. List what OpenRouter serves right now with `scripts/models.ts` (or `GET /api/v1/models?output_modalities=decisions`) and shortlist by the criteria in [references/models.md](references/models.md): the context length holds your real state with room to spare, the price fits the call volume, and the provider footprint fits your availability needs. The catalog cannot tell you how a model answers your questions, so when more than one candidate remains, run the step 8 probe set through every candidate with `decide.ts --compare` and pick on the observed answers, latency, and cost. Pin the chosen `canonical_slug` (or the versioned `id`) in one config value and never an alias, since an alias can shift probabilities under you. Swapping models later is that config change plus a rerun of step 8. Keep the API key server-side. For request types, validation, and the HTTP and SDK calls, import or copy `parseRequest` and `decide` from [scripts/lib.ts](scripts/lib.ts) instead of rewriting them. Done when a request returns one typed answer per question key and the response `model` string is logged with the answer.

7. **Gate in code.** Thresholds and business rules live in code as named constants next to the raw `noul` and `probabilities`. Before probing, a `noul` gate is `>= 0.5` and a `choice` is its `choice` field, and every stricter number waits for step 8. Compare in the direction the question asks, so `is_breaking.noul >= T` means breaking and "safe" is the `else`. A single gate fits only when the code path has two outcomes. When it already has a fallback (a review queue, a default, a bigger model), keep it as the outcome for probabilities near the gate and set that band's width in step 8. Use `confidence` only when such a fallback is cheaper than a wrong answer. It measures how concentrated the distribution is, not whether the answer is correct or the action is authorized. A `noul` near 0.5 means yes and no are similarly likely, not medium intensity. A threshold does not carry from a `noul` to a `choice` or from one model to another, and `P(yes)` from one question plus `P(not yes)` from another need not sum to 1. Done when every threshold is a named constant with the consequence of each mistake written next to it.

8. **Probe before you trust.** Run the real questions over representative inputs with the bundled script and read the raw probabilities. Cover the clear cases, an ambiguous case, a no-match case, an empty or off-topic input, a negated statement, and adversarial text that argues for its own classification. Set thresholds from those numbers, not from cookbook defaults, and repeat when the model changes. Done when each edge case produces the routing you want and every threshold traces to an observed number.

```bash
cd <skill-path>/scripts && npm install
npx tsx models.ts request.json                     # live catalog with context, price, providers, and fit for this request
npx tsx decide.ts request.json --model <model-id>  # raw HTTP to one model
npx tsx decide.ts request.json --sdk               # through @openrouter/sdk
npx tsx decide.ts request.json --compare           # same request to every pinned model in the catalog
```

`request.json` holds `{ "state", "questions" }` and optionally `"model"`. `--model` overrides the request's model, then `DECISION_MODEL` from the environment fills in when neither is set. The scripts print the answers, resolved model version, latency, and cost, and `--compare` prints one entry per model with the error inline when a model fails.
