~/blog/jev-router-claude-code

Jev as a Claude Code prompt router: routing by task size

A Jev classifier in a Claude Code prompt hook sizes each message and routes it to a helper agent. Thresholds, measured cost, and failure modes.

Krzysztof Słomka15 min read

Every prompt I send to Claude Code runs on the same model, whether it is a rename or a question about why checkout double-charges under load. The session is the same, and the token bill scales with context rather than with how hard the question is. That default is wrong. Most of my prompts are small, and the expensive model is wasted on them. The hard ones deserve that model, but they get less attention than they should when they are buried in a stream of trivial work.

This article describes a small experiment: a classifier in front of Claude Code that reads each message, decides how big the job is, and routes it to the smallest helper that can handle it. The classifier is Jev, from TypeSafe. Jev does not write anything. It answers typed questions with calibrated probabilities, which is the property a router needs.

I'll walk through why a classifier fits and a generator does not. Then I'll cover how the hook is wired into Claude Code, where the confidence threshold sits and why, and what a routing decision costs in money and latency. I'll finish with where it fails and what I have not measured yet.

Jev answers questions, it does not write#

Jev is a decision model. You send it a piece of text, called state, and a map of questions. It returns one answer per question. There are three question types. A noul returns a bare probability that the answer is yes. A choice picks one option from a set and returns the full probability map. A score places the input on an ordered scale.

The router uses one choice question with four classes:

{
  "size": {
    "type": "choice",
    "instructions": "A developer sent this message to a coding agent. What is the smallest class of model that can do the job well? Judge only the work the message asks for. Ignore politeness, length and tone.",
    "criteria": {
      "tiny": "A lookup, a rename, a single fact, a one-line answer, a trivial mechanical edit.",
      "everyday": "An ordinary piece of writing or a contained change: a normal email, a post, a short document, one function, one small file.",
      "large": "A multi-step build, research across sources, a full report, changes spanning several files.",
      "hardest": "Strategy, architecture, ambiguous debugging, or any judgment where being wrong is expensive to undo."
    }
  }
}

The instruction tells Jev to judge the work, not the tone. A polite two-line request for a refactor across four files still lands in large. The criteria carry the boundary cases, because Jev follows the words it is given and anything I want it to decide has to be written down in them.

Each response carries a confidence value derived from how the probability is spread across the options. It describes the model's certainty about its own answer. The router is built around that reading.

The same pattern works for any triage over a pile of text. TypeSafe's documented limitations for jev-1.13 suggest the following practices:

  • Ask several questions per call. The state is tokenised once, so extra questions cost little more than one.
  • Keep arithmetic, counting and date logic in code. The docs say Jev is not a calculator and reads dates as text. Compute the value yourself and ask Jev only about the result.
  • Filter the state before sending it. Unrelated context is a distractor, and accuracy drops as it grows.
  • Write the boundary cases into the criteria. Jev reads literally and does not infer the intent behind a question. Double negatives and multi-hop reasoning are answered less reliably, so name the field the question depends on.
  • Do not use it to generate. It is not trained for text generation, and the docs say it will be slow and weak at it. Use it to pick an option, then let a generative model write.
  • Treat the input as untrusted. Jev does not treat the state as hostile by default, so anything a stranger wrote is data, and the criteria are the only defence against injected instructions.
  • Test choice lists in both orders. The docs note a tendency to favour earlier options, so a close call can flip with the ordering.

Where the hook sits in the prompt path#

A prompt card passes through the UserPromptSubmit hook into a classifier, which fans out to four agent blocks

The hook is a gate, not a router. It only reads the prompt and hands a size to the classifier, which fans out to four helpers. The lead model stays in the path for the messages the thresholds send back.

The hook runs on every submitted prompt. It reads the prompt text and the working directory from the hook input, asks Jev for a size, and returns the answer to Claude Code as additionalContext. The lead model then sees an instruction to hand the work to a subagent named route-<size>. Four subagents exist: route-tiny, route-everyday, route-large and route-hardest. They are ordinary Claude Code agents with the full tool set, and each one pins its own model and effort in its frontmatter:

SizeHelperModelEffort
tinyroute-tinyhaikulow
everydayroute-everydaysonnetmedium
largeroute-largeopusmedium
hardestroute-hardestopushigh

Four rungs labelled tiny, everyday, large and hardest, with a prompt routed to the everyday rung

A prompt lands on one rung. The other three stay dimmed, and the rung it lands on is the only thing the helper choice depends on.

Each helper ends every reply with a fixed closing line that names itself and its model, for example done by route-tiny (haiku). The line shows which tier produced the answer, and the hook asks the lead model to keep it intact when it passes the result on.

For the wider coordination layer that sits on top of Claude Code agents, the Ruflo post covers what a swarm adds and what it costs.

Routing is opt-in per project. The hook is installed once, globally, but it does nothing unless a flag file exists for the current directory. A repository never carries a file because of the router, and turning it off is one command.

Thresholds: two overrides that keep work in the lead model#

The hook does not trust the classifier blindly. Two rules override a delegation decision, and both send the message back to the lead model:

DECISION="delegate"
if (( ${#PROMPT} < FRAGMENT_CHARS )); then        # FRAGMENT_CHARS=80
  DECISION="inline"
elif awk -v c="$CONF" -v m="$MIN_CONFIDENCE" 'BEGIN{exit !(c < m)}'; then   # MIN_CONFIDENCE=0.60
  DECISION="inline"
fi

Short messages are excluded because they usually only make sense inside the conversation. "yes, do that one" is a size-less instruction. Jev sees the text, not the history, so it cannot know what "that one" refers to. The second rule handles uncertainty: below a confidence of 0.60, the router declines to decide. The threshold is a starting point I picked, not a calibrated value. Jev's own guidance is to act automatically above roughly 0.8 and route the rest to a human or a slower path. A router that sends everything below 0.8 back to the lead model would delegate almost nothing, so I set the cut lower and accepted more misroutes in the middle band.

The hook fails open. If the key is missing, the request times out after eight seconds, or the response has no answer, the script exits without output and the message proceeds as if the router did not exist. A router that can block a prompt gets switched off the first time it misfires.

The router had never fired#

When I checked the log, it was empty, and the hook's own status command said so: no messages routed, not even a failed call. An empty log with no failure rows means the script exited before it reached the API. The reason is one field name. The hook reads the prompt from .prompt_text, but Claude Code's UserPromptSubmit input carries the text in prompt. The variable came out empty, and the guard [[ -n "$PROMPT" ]] || exit 0 quietly ended every run. Nothing was blocked and nothing was routed. That is the fail-open behaviour working as designed, and it is also why the failure was invisible.

A message passes a sealed checkpoint unchanged while a grey log line beside it stays empty

The message went through untouched, which is the design. The log line beside it stayed empty, which was the only sign anything was wrong.

The fix is a single line, reading .prompt // empty in place of .prompt_text // empty. The rest of this article describes the logic as designed. The numbers below come from calling the endpoint directly, because the installed hook had not yet produced a single real routing decision.

What a routing decision costs#

I did not have production data to report, and the installed router had produced none. Instead I sent six synthetic prompts through the same endpoint with the same question:

Prompt (shortened)Jev sizeConfidenceLatencyCost (USD)
Rename the variable userId to accountId in src/auth.tstiny0.950.46 s0.0000194
What is the default port of PostgreSQL?tiny1.000.33 s0.0000192
Write a short follow-up email to a client who has not paid an invoiceeveryday1.000.41 s0.0000200
Refactor the payment module so the retry policy is shared across four fileslarge0.990.28 s0.0000200
Checkout intermittently double-charges under load, find the root causehardest0.970.29 s0.0000202
Add a dark-mode toggle to the navbar componenteveryday0.930.27 s0.0000192

Six prompts are a smoke test, not a benchmark. Even so, the shape is clear. Each decision costs about two thousandths of a cent, so ten thousand routed messages come to roughly twenty cents. At that price cost is not the limiting factor. Latency is. Every message waits roughly a quarter to half a second before the lead model starts, and the trivial messages pay that wait too.

The sizes look plausible on a read-through. A dark-mode toggle is a contained change to one component, which makes it everyday rather than large, even though it touches styling, state and tests.

Where it fails#

The router sees one message at a time, and that produces four failure modes I can name.

  • Follow-ups lose their context. "now do the same for the other two services" is forty characters and has no size of its own. Jev judges it from the text alone, so it can land in tiny while the work depends on everything above it. The hook tells the lead model to keep follow-ups inline, but that is an instruction in prose, and the model has to choose to follow it. Nothing enforces it.
  • Subagents start cold. A helper gets the prompt and whatever the lead model put in the delegation, not the conversation. That suits self-contained work and does not suit anything that depends on files the lead model read earlier. Writing that shared context down once, in CLAUDE.md, is covered in the Obsidian post.
  • The text leaves the machine. Every routed prompt goes to OpenRouter and from there to TypeSafe. The hook says so when it is switched on. For client work or anything under NDA, the router stays off, and the per-project flag is what makes that safe to forget about.
  • Literal reading. Jev follows the words. A message such as "figure out why it is slow" is short and vague, and a model that reads "figure out" as a small request could size it as everyday. The criteria help only if they name that case.

Gate on confidence, log everything, fail open#

Four habits follow from the failure modes above.

Send low-confidence decisions to the lead model without drama. The threshold is a dial between misroutes and wasted delegation, and it should be set from logged data rather than from taste.

Spell the boundary cases out in the criteria. They are the only lever the router has. "Changes spanning several files" needs to be explicit, because a rename across two files and a refactor across four look alike in a short message.

Log every decision with its confidence and cost. The hook writes one JSON line per call, and status summarises sizes and spend. The log is what tells you whether the router works, so an empty log after a week is itself a finding.

Keep the classifier out of the way of the user's message. It sits in the hot path, so it must never be able to block the prompt. The same applies to privacy: prompt text is the only thing that leaves the machine here, but a prompt can carry anything the user typed, so the flag belongs only on projects where that is acceptable.

What I have not measured#

  • Routing accuracy. I have judged six sizes by eye. I have not labelled a set of real prompts and scored Jev against it, and I will not claim an accuracy figure until I have.
  • Whether the delegated answers are better or worse. Routing saves money only if the helper handles the work as well as the lead model would have. I have not compared outputs side by side.
  • Real latency under load. The six timings were sequential, from one connection, on one day.
  • Whether the follow-up problem costs anything in practice. I suspect it does, but I have not counted how often a delegated follow-up needed correction.

Until the log has a few hundred real rows, treat every number in this article as an illustration of the mechanism, not as a result.

Set it up with one prompt#

If you want the same router on your own machine, paste the prompt below into Claude Code, from any directory. It builds the pieces this article describes, uses the correct prompt field, and stops before anything leaves the machine. Nothing is routed until you enable a project yourself.

Set up a Jev prompt router for Claude Code on this machine. Do the steps in order. Stop and ask me
wherever a step says so.

1. Check the key. In a subshell, run
   security find-generic-password -a openrouter -s openrouter-api-key -w >/dev/null
   If it fails, stop and tell me to store my OpenRouter key in the macOS Keychain (account
   "openrouter", service "openrouter-api-key"). Never print the key or write it to a file.

2. Create ~/.claude/hooks/jev-router.sh and make it executable. It is a UserPromptSubmit hook:
   - Read the stdin JSON. Take "cwd" and the prompt from the "prompt" field. Not "prompt_text".
   - Exit 0 silently unless a flag file exists at ~/.claude/jev-router/enabled/<cwd with every
     non-alphanumeric character replaced by ->.
   - POST to https://openrouter.ai/api/alpha/decisions with model "typesafe/jev-1.13" and one
     "choice" question named "size" with criteria tiny, everyday, large, hardest. Write the
     criteria so that a lookup or rename is tiny, one ordinary piece of writing or one contained
     change is everyday, a multi-file or multi-step job is large, and architecture, ambiguous
     debugging or costly-to-reverse judgment is hardest.
   - Use an 8 second curl timeout. Fail open on any error: no output, exit 0.
   - Keep the lead model inline when the prompt is under 80 characters or confidence is below 0.60.
   - Otherwise print {"hookSpecificOutput":{"hookEventName":"UserPromptSubmit","additionalContext":"..."}}
     telling the lead model to hand the work to subagent route-<size> and to keep that agent's
     closing line intact.
   - Append one JSON line per call to ~/.claude/jev-router/log.jsonl: ts, project, size,
     confidence, cost, decision, and error when there is one.
   - Support subcommands: on <dir>, off <dir>, status (print enabled projects, sizes and total cost).
     "on" must print a warning that prompt text will be sent to OpenRouter and TypeSafe.

3. Create four agents in ~/.claude/agents/:
   - route-tiny: model haiku, effort low
   - route-everyday: model sonnet, effort medium
   - route-large: model opus, effort medium
   - route-hardest: model opus, effort high
   Each one answers directly, says when the task is bigger than its class, and ends every reply
   with exactly this line, on its own: done by route-<size> (<model>).

4. Register the hook in ~/.claude/settings.json under hooks.UserPromptSubmit, command set to the
   script path, timeout 10. Show me the diff and wait for my yes before writing it.

5. Do not create any enabled flag. Routing stays off everywhere. I will run
   ~/.claude/hooks/jev-router.sh on <dir> myself for a project I choose.

6. Verify without sending anything real: run the script with a sample hook input on stdin for a
   directory I have enabled, but only after I approve the sample prompt, because it goes to
   OpenRouter. Confirm one log row is written.

7. Report what you created, what you did not verify, and the exact paths changed.

Do not install anything else, do not change other settings, and do not push or commit anything.

Keep it off for client work and anything under NDA. The first real call sends the prompt text to OpenRouter and from there to TypeSafe.

One rule#

Let the classifier choose the tier, never the data boundary. Whether a prompt may leave the machine is a per-project decision made by a person, and the flag file exists so that the decision is explicit, written down and reversible. Saving money on routing does not help if the cheap path also sends client text to a vendor.

$ related posts