Jev Decision Layer: Save Frontier Tokens on Closed Decisions
Send closed agent choices to TypeSafe Jev at $0.042 per million input tokens, with output free. Qualixar adds the local gate and receipt. The host still acts.
The expensive part of an agent loop is often the closed choice in the middle. Which of these three tools should run? Does this ticket belong to billing or engineering? Which supplied file should be read first? A frontier model can answer. It also charges you for the reasoning tokens and for the paragraph wrapped around a value the rest of the program only needed as a selection.
Qualixar Jev Decision Layer 1.0.8 exists so that selection leaves the frontier turn. It sends the bounded question to hosted TypeSafe Jev, checks the typed answer with a local gate, and writes a receipt. The frontier model stays on the patch, the draft, or the plan. The host still executes. A recommendation is not a shell grant, a file write, a browser click, or a publish button.

Move the closed choice off the frontier turn. Choice, gate, receipt. The host still acts.
Where the tokens and the money go
TypeSafe published the comparison this layer is built around. In its 15 September 2026 announcement, Jev is reported at 193.6× faster and 444.6× cheaper than the LLM reference on TypeSafe's workflow comparisons. The same post puts the price and the latency next to ordinary chat models:
| Chat-model range TypeSafe published | Jev, as TypeSafe published it | |
|---|---|---|
| Input | $0.20 to $10 per million tokens | $0.042 per million tokens |
| Output | about 5× the input price | free |
| End-to-end response | 3 to 329 seconds | 70–500 ms |
| Workflow result | the LLM reference on those workflows | 193.6× faster, 444.6× cheaper |
The workflow multiples come from TypeSafe's evals. They are TypeSafe's workflow-level figures, not numbers you can derive from the price table above, and this article does not reproduce the eval token volumes behind them. TypeSafe says they are on the higher end of the gains it expects, the reference answer is the average of GPT-6 Astra and Fable 5.1, and the LLM side used TypeSafe's System One wrapper. The list price is in the model documentation: Jev 1.13 (jev-1.13.0) at $0.042 per million input tokens, output free. When this article was checked, jev-latest and jev-preview both pointed at that version. The documented context budget is 64k tokens for the state plus every question, and 32k for the state plus the longest question. Input is text. TypeSafe also says the published rate limits, 250,000 tokens per second and 1,200 requests per minute, can change while demand is high.
That table is the practical view. A harness that asks its chat model to narrate every closed choice pays frontier input and frontier output for a selection. A harness that asks Jev pays the input price on the state and the questions, pays nothing for the typed output, and gets a probability over options you already defined. Questions that share one state belong in one call. TypeSafe's primitives documentation says its parallel-questions cookbook shows 13 questions in a single call at 11.5× lower cost and 9.6× lower latency than 13 separate calls.
Optional local Laya-MLX, on a supported Apple-Silicon Mac, is the other way to keep a restricted decision off a hosted bill. It is separately installed and attested. Jev-only mode never silently switches to it.
The Qualixar layer is the part a harness can reuse: 20 MCP tools, 38 recipe specifications, a local gate, a receipt, and adapters for Codex, Claude Code, Hermes, Antigravity, and VS Code. You supply the options. Jev returns the distribution. The gate tells the host to verify or to ignore. The frontier model is reserved for the work that still needs language.

TypeSafe's own chart, from the System One announcement. Jev is the pink point on the left. The line is their frontier: nothing in that plot is both cheaper and more accurate. Source: typesafe.ai and the live evals.

The same announcement shows why the bill drops. The work is split into small questions the code can compose. A frontier model is not asked to narrate the whole incident in one turn.
What a typed answer actually contains
TypeSafe's question primitives come in three shapes. Each one is a judgment a person could make quickly if the relevant state were already on the desk.
| Question | Ask it when | What comes back |
|---|---|---|
| Choice | The answer is one of a known set, with no order between the options | The selected option, a probability for every option, and a confidence summary |
| Score | The answer sits on a scale whose levels you define | A position on that scale, the level legend, a probability for each level, and confidence |
| Noul | The useful signal is the probability that a yes/no statement is true | A value from 0 to 1. There is no separate confidence field |
A Choice fits a ticket route, a document type, or a tool name. A Score fits severity, frustration, or review attention, provided you write what each level means. A Noul fits "does this message request a refund?" It does not fit "how strong is this candidate?" A Noul of 0.5 means yes and no were equally likely under the question you asked. It does not mean the candidate is medium. If you wanted a level, ask for a Score and define the levels.
TypeSafe's docs use a duplicate-charge ticket as a teaching state: a customer says order A-104 was charged twice and asks for a refund, the order shows two captured charges of $49, and the policy says duplicate charges are eligible. Two Noul questions can point at exact fields, one for whether the message requests a refund and one for whether the policy supports it given the charges. That example is TypeSafe's, and it shows the shape. The model is asked to judge named parts of a state, not to invent a new policy.
Choice and Score confidence is derived from the reported probability distribution. TypeSafe documents that derivation on its confidence page. A peaked distribution produces higher confidence. Confidence is a summary of that distribution. It is not a separate accuracy audit, and the Qualixar recipe gate does not treat it as one. The gate applies two configured keys, min_confidence and min_selected_probability (see plugins/qualixar-jev-decision-layer/runtime/jev_auto/recipe_gate.py), as conservative policy. Neither bar has been calibrated on this repository's tasks.
TypeSafe also says schema matching is guaranteed, so a Jev answer cannot fall outside the options or levels you supplied. That is a type-safety claim, and TypeSafe is explicit that its zero type-error figure is not an empirical count of factual correctness. A typed answer can still be the wrong option. The surrounding system has to decide what to do with a wrong, uncertain, or malformed result. That is the job of the local gate.
One more TypeSafe pattern matters for cost. Its primitives documentation says the parallel-questions cookbook shows 13 questions in one call at 11.5× lower cost and 9.6× lower latency than 13 separate calls, with the same answers. Questions that share a state are evaluated together. Adding a question you might ignore is cheap relative to opening another frontier-model turn. The Qualixar layer uses that shape. It does not republish those cookbook timings as its own benchmark.
What the layer adds around the call
A direct Jev call already returns a typed answer. The layer is the harness around that call.
Version 1.0.8 ships 20 MCP tools (registered in plugins/qualixar-jev-decision-layer/__init__.py plus the typed route), 38 recipe specifications (in recipes/, one fixture each in fixtures/), 114 synthetic offline fixtures, and adapters for Codex, Claude Code, Hermes, Antigravity, and VS Code. The 20 original workflow contracts correspond to 20 of the 38 recipes. They are not 20 extra recipes on top of the catalog. Two of the recipes are aimed past the person writing the plugin: qualixar.work-item-priority compares one work item with a rubric and priorities you supply, and qualixar.content-repurpose selects among content formats you supply.
The local broker checks the workspace scope you reviewed, screens recognizable secrets, enforces daily request limits, and sends the bounded state and questions for that call. Jev-only mode does not silently fall back to Laya. Hybrid mode can send selected restricted decisions to an attested local Laya route. Laya is optional, separately installed, and limited to Apple-Silicon Macs that pass attestation. Hosted TypeSafe and OpenRouter are separate providers, with separate keys and separate consent.

You supply the state and the allowed answers. The gate in 1.0.8 caps a would-be pass at verify.
A local run of the offline self-test against the published 1.0.8 source tree, on 29 September 2026, reported the result below. The run used the same file tree as tag v1.0.8:
plugins/qualixar-jev-decision-layer/scripts/jev selftest
It reported all_passed: true, 114 cases, 114 passed, 38 recipes. The tool's own disclaimer is the right reading: this is an offline synthetic gate contract check. It is not provider accuracy, cost, or latency evidence. Each recipe carries three synthetic cases. One is clear-cut and must clear the gate. One is ambiguous and must not. One hides an instruction inside material that is supposed to be data, and that case must not clear the gate either. A fixture cannot be marked passed by recording whatever the gate happened to do.
The gate, and why a pass still says verify
A live jev_recipe_try returns a host_action and a policy receipt ID. The response is marked EXPERIMENTAL_ADVISORY. The receipt records the recipe status, the provider receipt ID, the provider and model, and the local gate result. None of that runs the recommendation.
host_action |
Meaning in 1.0.8 | What you should do |
|---|---|---|
act |
The answer cleared the configured gate | No shipped recipe can keep act. Every recipe is SPECIFICATION_NOT_MODEL_EVALUATED, so a would-be act is capped to verify |
verify |
The answer is advice that still needs an independent look | This is the strongest result a shipped recipe can return |
ignore |
The model selected unknown |
Do not treat it as a recommendation. Decide the step the way you would have without it |
A missing, malformed, or out-of-range threshold fails closed to verify. So does a value outside its domain, a label the model ranked below another, and an act that would carry no recommendation. An explicit unknown can return ignore.
That cap is deliberate. The recipes have specifications and synthetic fixtures. They have not been evaluated against a labeled set of provider answers. Until that evaluation exists, the layer will not tell a host that a live answer is ready to execute. The host, or the person using the host, still reads the result.

The broker stays on the computer. Hosted Jev sees only the reviewed request. The receipt comes back as advice.
The security boundary matches that picture. Recognizable credentials are blocked in every hosted mode. That screen is best-effort, not a data-loss-prevention product. Jev maximum can include client or confidential workspace text only after an explicit review, and only if you have the authority to share it. Ordinary contact details and workspace paths are accepted only inside an enrolled internal or restricted scope. Private state for the supported macOS runtime uses Keychain. The same-user process is outside the plugin's tamper-resistance claim: this layer does not protect you from hostile code already running as you.
Four places the same mechanism shows up
The mechanism does not change when the person changes. What changes is the state they supply and the options they are willing to accept.
A developer choosing a tool or a skill
You have two plausible skills, or three tools, and the next step should load one of them. Put the task and the candidate names in the state. Ask a Choice. The answer names one candidate and shows the probability of the others. If the distribution is flat, or the model returns unknown, the gate will not dress that up as a command. You load the skill, or you do not.
This is also the right shape for a long tool result. Before the host spends a frontier turn reading a huge log, a bounded question can ask whether the log contains anything that answers the current goal. The original log stays authoritative. The answer only tells the host whether the read is worth doing.
A manager ordering one work item
qualixar.work-item-priority takes a work item, a rubric, and priorities you already wrote. It does not discover your roadmap. It compares the item you supplied with the rules you supplied and returns an advisory ranking. A manager can disagree with it in one look, because the options were theirs. The useful part is the receipt: which rubric went in, what came back, and that the host did not move the ticket by itself.
A creator choosing a format
qualixar.content-repurpose selects among formats you provide, such as a short video, a newsletter, or a carousel, using a brief you provide. It does not write the piece. A creator who already knows the three packages they can actually ship can use the score to break a tie, then edit. If the brief is too thin for the formats to be distinct, an unknown or a flat distribution is a better result than a confident guess.
A reviewer deciding where to look first
A proposed patch can be scored for review attention. The Score levels have to say what "look here first" means: a public API change, a migration, a test-only edit. The layer can suggest the first area. It does not review the patch, and it does not run the tests. Pair it with a completion gate, such as the one in Bounded Loops & Graphs, when the question is whether the work is finished. Jev answers a bounded choice. A completion gate answers whether the required property holds. They fail in different ways, and they should stay separate.
In all four, the person supplies the evidence and the allowed answers. The model returns a distribution. The gate classifies the result. The host, or the person, performs the action.
Five adapters, with unequal proof

One local broker. Five adapters. Advisory only.
The shared runtime lives in plugins/qualixar-jev-decision-layer/. The adapters stay separate because the hosts do not share a hook contract. Codex registers a PostToolUse matcher over documented read-only tools. Claude Code can take an advisory-only PreToolUse hint. Antigravity does not receive a PreToolUse registration from this package, because that contract can change host permission. Hermes has its own hook and tool entry points. VS Code has no hook surface, so the adapter writes a workspace MCP entry.
The host guide separates evidence from boundary:
- Codex. Installed MCP tools and a live synthetic TypeSafe route were verified on macOS. That does not establish Linux, and it does not establish automatic-hook coverage.
- Claude Code. The plugin installs from the marketplace, validation passes, commands and skills load, the MCP launcher completes an
initializehandshake, and the hook launcher stays silent on an unenrolled workspace. A native tool turn inside a live session is not yet verified. In the desktop app's Code tab, the plugin-provided server does not load. - Hermes. The 1.0.8 source registers its declared tools, including self-test, verify, and rerank, plus its hook. A native model and tool turn still needs verification after the installed copy is refreshed.
- Antigravity. The package includes a PreInvocation advisory hook and skills, and it does not request PreToolUse authority. A native turn is still unverified.
- VS Code. Tests cover writing and merging
.vscode/mcp.jsonin the documentedserversshape. No live Copilot agent-mode turn has been run. There is no extension. Registration is the integration.
An adapter in the tree means the package ships a shim. It does not mean a live turn has been run. native_status stays NOT_RUN until someone runs one.
Platform scope is separate from the adapter list. 1.0.8 is supported and verified on macOS. Linux code remains experimental and unverified: generic Linux CI does not establish the Secret Service, broker, host, and provider path together. Windows hosted runtime entry points fail closed. The Windows private-state contract did not pass CI, and that lane is out of this release. Local Laya stays on Apple Silicon.
Paste a prompt. Codex or Claude Code does the rest.
The repository is the product. Install it, then give the host one of the prompts below. You do not write MCP JSON. The host calls the recipe. You read the choice, the probabilities, host_action, and the receipt.
These three prompts use the synthetic cases shipped with the recipes. The fixture files also contain hand-authored mock answers so the gate can be tested offline. Those mocks are not a live Jev or Laya prediction. If you send the same text to a provider, the distribution can differ, and in 1.0.8 a would-be pass still comes back as verify.
A developer, choosing a tool
Use
qualixar.tool-selection. Goal: Identify where a cache key is read. Tool descriptions: search, exact symbols; read, inspect candidates; test, execute the test suite; review, inspect existing evidence. Return the choice, the probabilities,host_action, and the policy receipt id. Do not run a tool.
The recipe's own sample is that goal and those four categories. Selection does not grant permission and does not build a shell command.
A manager, ordering one item
Use
qualixar.work-item-priority. Work item: Fix the cache expiry defect that blocks the committed release. Stated priorities: first address work that blocks a committed release, then non-blocking improvements. Known constraints: this defect blocks the committed release, and no capacity estimate was supplied. Decision horizon: this week. Return the score, the level probabilities,host_action, and the receipt id. Do not move the work item.
A creator, choosing a format
Use
qualixar.content-repurpose. Source: an original diagram shows how evidence enters a completion gate, gets checked, and produces a pass-or-review result. Audience: junior developers learning a three-step process. Goal: explain the sequence visually. Constraints: five panels, English text, no video production this week. Rights: the author created the diagram and confirmed permission to adapt it. Return the format, the probabilities,host_action, and the receipt id. Do not write the piece.
The allowed formats are carousel, newsletter, short post, short video, or unknown. The shipped nominal fixture expects carousel at RECOMMEND, observed in the self-test as capped to verify. Unclear rights are supposed to come back as unknown.
See both routes without spending anything
Every recipe can be replayed with no TypeSafe call and no Laya call. From a host that has the tools, paste:
Run
jev_recipe_selftest. Do not call TypeSafe. Do not call Laya. Tell me how many cases passed.
On 29 September 2026 that offline replay, against the published 1.0.8 tree, reported 114 passed out of 114, across 38 recipes. It checks the local gate. It does not spend provider money, and it is not an accuracy score.
The same zero-spend view is in the workbench. From a checkout:
plugins/qualixar-jev-decision-layer/scripts/jev workbench --workspace /absolute/path/to/your/project
It opens on loopback, starts offline, and lets a manager or creator pick a recipe, read the question, and run a synthetic fixture. Hosted Jev and local Laya are the two live routes behind those same recipes. The offline replay shows the question, the gate, and the receipt before either route is switched on. A later live Jev answer uses the TypeSafe key and the $0.042 per million input price, with output free. A later Laya answer stays on a supported Apple-Silicon Mac after a separate install and attestation, and Jev-only mode never silently switches to it.
Install on Codex or Claude Code
To install without cloning, add the marketplace and the plugin. The blocks below are documented install steps from the getting-started guide, shown so you can copy them; the only executed run in this article is the self-test further down. The marketplace name is qualixar-jev-layer while the repository is qualixar/jev-decision-layer. On Codex Desktop the documented commands are:
codex plugin marketplace add qualixar/jev-decision-layer
codex plugin add qualixar-jev-decision-layer@qualixar-jev-layer
On Claude Code:
claude plugin marketplace add qualixar/jev-decision-layer
claude plugin install qualixar-jev-decision-layer@qualixar
If a plugin is already installed, finish the active task and fully quit the host before upgrading. A running Codex task can still call a hook from the previous cache after the files move. The documented refresh is codex plugin marketplace upgrade qualixar-jev-layer, then a new task. Check the installed version in the Plugins screen or with codex plugin list. Do not paste a provider key into chat to fix a missing local file.
Setup is a reviewed local wizard, not an .env file. In a new task, ask the host to set up Qualixar Jev Decision Layer for this workspace. The wizard shows the folder, the provider, the text scope, the expiry, and the daily limits before you confirm. Modes are Jev public, Jev reviewed internal, Jev maximum, Jev plus Laya hybrid, and Laya-only. Automatic prompt guidance starts off. Explicit tools start on after the reviewed setup. Installing the plugin does not trust its hooks. If the host asks, inspect the current hook definition before you allow it.
After setup, paste one of the three prompts above in a new task. Open Codex or Claude Code, point it at this repository's plugin, and the host makes the call. jev doctor is offline. Before enrollment, ACTION_REQUIRED with NOT_ENROLLED and exit status 2 are expected. A marketplace install does not add a global jev command, so that CLI check uses a local checkout (git clone https://github.com/qualixar/jev-decision-layer, then the plugins/qualixar-jev-decision-layer/scripts/jev entry in that tree).
What you should still do yourself
Use Jev for the closed choice, and leave the frontier turn for the patch, the draft, or the plan. Write the options yourself. Write what each Score level means. Put every question that shares a state into one call, because a second call is a second bill. Read the distribution, not only the winning label. Treat verify as advice you still check. Treat ignore as no advice, and decide the step without a recommendation.
Keep tool execution, tests, and publishing on the host. If the task is "is this finished?", use a gate that checks the artifact. If the task is "which of these should we try?", use the typed decision and keep the receipt. The receipt shows what was asked, what came back, and that the host did not spend a frontier generation on that selection.
The price and the workflow multiples above are TypeSafe's published figures. The getting-started page is the current source for the Mac commands. Start with the offline self-test. Then send one synthetic choice, and read the receipt before you send a real one.
Varun Pratap Bhardwaj builds AI Reliability Engineering tools at Qualixar. ORCID 0009-0002-8726-4289