jevcode

v0.2 · MIT · no index, no vectors

A coding agent that decides with a model that cannot write a word.

Jev answers typed questions — is this true, which of these, where on this scale — with calibrated probabilities: 256 in one request, about 200 ms, a hundredth of a cent. So jevcode asks about everything before it moves, and a small fast model does the typing under instructions it never chose.

No key to look around: jevcode --demo runs the whole interface off a canned table, in a throwaway copy of a sample project.

~/cart
$ jevcode "make Cart.total accept a discount argument,
          taken off before tax"

step 1 search  p=0.50 conf=0.54
        searched 'Cart' → 2 files
step 2 read    p=0.90 conf=0.98
        opened cart.py (21 lines)
step 3 edit    p=0.80 conf=0.86
        writing cart.py Cart.total (lines 14-15) — 6 candidates
        dropped 4: same as candidate A
        chose A (p=0.72, works=0.93, out-of-scope=0.10)
        python3 -m unittest discover -s tests -q passed

done — task carried out and checked
73 decisions in 6 requests (5.2s, 41k tokens in)
6 writer calls (12.6s) · 7.6s total
100%of tasks solved, tests deciding
34smedian, start to green tests
31decisions per task, on average
9tasks, each judged by its own tests

The trade

Reverse the price of a decision and the agent changes shape.

An ordinary coding agent spends its budget on deliberation: every step replays a growing transcript through a large model. Steps are slow, expensive and precious, which is why such an agent commits to the first plausible move and finds out about the alternatives by walking into them.

ordinary agentjevcode
a decisionseconds, cents, a full context replay ~200 ms, ~$0.0001, no transcript
decisions per steponedozens, in one request
looking aheadtake the step and find out score the whole tree first
picking among draftskeep the first one write six, judge six, keep the best
checking a commanda deny list of regexes three questions about this exact command
context growthevery file read stays in the prompt forever state is assembled per question

That last row matters more than it looks. Nothing accumulates in a prompt here: the state handed to Jev is built fresh for each question, so a long session never slowly poisons itself with everything it has ever read.

One step

Every question the step might need, asked in a single request.

The decision itself and the arguments for each action it might choose go out together. Most of the answers are thrown away — whichever the chosen action does not need. They are speculative on purpose: a hundred extra questions cost about as much as one.

{
  "action":         choice({read, search, edit, create, run, finish}),
  "file":           choice(up to 255 files, described),
  "region":         choice(the functions and classes of the open file),
  "command":        choice(what the project declares: make, npm, pytest),
  "query":          choice(literal strings taken from the task),
  "done":           noul("carried out AND confirmed by a command?"),
  "needs_human":    noul("a decision only the owner can make?"),
  "worth_read", "worth_edit": noul("would this produce anything new?"),
}

Candidates, not a candidate

An edit asks the writer for six drafts at once. Drafts that do not parse, came back empty, or match the current code are dropped in Python before anything is judged — facts first, opinion second. If the tests reject the winner, the runner-up is already written and already judged.

A gate in front of every command

Would this destroy work, is it unrelated to the task, does it reach outside the repository — asked about the actual command, every time. A deny list only catches the shapes somebody thought of; this reads find . -delete the way a person does.

A broken toolchain is not a broken patch

A missing test runner fails exactly like a wrong change. The output is scored on a four-level rubric, and an environment fault keeps the edit instead of throwing away a change that was fine.

Three moves ahead, one request

A beam search over future actions: each branch carries its own assumed history, each question addresses its branch by path, so one request answers "what next" for the whole frontier. Width three, depth three, about a second.

Measured, not claimed

Numbers on this page are generated from the results file.

Nine tasks, one attempt each. Both jevcode rows were re-run on 2026-09-21 against commit 94b0162; the opencode rows are the 2026-09-20 runs, which our own code cannot move. Each agent is shown with both writers it was measured on; the head-to-head pair shares one. Test files and Makefiles are checksummed, so a suite made green by rewriting its own tests does not count, and a run that edits the original task library instead of its own copy is not scored at all.

Every contestant, measure by measure
Leaderboard, by tasks solved
Seconds per task
Solved, by task
Cost per task
tasksolveddecisionsrequestswall
cartyes14112.9s
csvparseyes14167.6s
durationyes14134.2s
jsdedupeyes14118.8s
jsonflagyes14120.8s
pagesizeyes826147.5s
retryyes30278.7s
slugifyyes86882.4s
ttlcacheyes14115.2s
agentsolvedmedian timecost per task
jevcode · Jev + MiniMax-M2.79/9 (100%)34.2s0.10 RUB
jevcode · Jev + mercury-29/9 (100%)4.7s0.62 RUB
opencode · MiniMax-M2.78/9 (89%)24.7s—
opencode · mercury-26/9 (67%)72.5s—
49% → 78%

What the judge adds. The writer produces N drafts for 80 HumanEval tasks, the tests decide which work, and Jev picks without ever seeing them: llama-3.2-3b goes from 49% one-shot to 78%, against a ceiling of 85% — 1148 questions in 80 requests. With a writer already at 94% there is nothing left to win, and it wins nothing.

13 / 15

Finding the right file with no index at all. Fifteen questions about this repository, each with a known answer, every run reading the tree from scratch: 30 questions in 15 requests, 12 seconds. Nothing to embed, nothing to keep fresh.

The head-to-head against another terminal agent lands here after the next run. The harness is already in the repository — bench/compare.py, a plain list of command templates — because putting another agent's name in a table before it has run would be worse than an empty section.

In the terminal

A prompt you can live in, not a dashboard.

jevcode with no arguments opens a session in the current directory. Everything scrolls, everything can be copied out, and the only thing that repaints is the status line.

a jevcode session in the terminal

Recorded with jevcode --demo, which answers from a table so the interface can be shown without a key. The same recording moving, and the raw cast beside it if you would rather replay it yourself.

  • @path puts a file in front of the agent, matched on any tail of its path.
  • !command runs a shell command yourself without leaving.
  • /undo moves the files, not just the transcript — so it works in a directory that is not a git repository at all.
  • Sessions survive the terminal closing: jevcode -c carries on, /sessions lists them.
  • Before an edit or a command you get the diff and a question, with "yes, and stop asking" as the second option.
  • AGENTS.md is read if the project has one — the same file the other terminal agents look for.
command
/new /clearstart over
/sessions /resumelist and switch
/undo /redomove the files back and forward
/diffeverything this session changed
/costdecisions, requests, seconds, money
/models /modelwho decides, who writes, swap the writer
/detailsshow or hide each decision as it is made
/permissionask, allow or deny — how much it asks
/initwrite an AGENTS.md

Install

Two keys: one decides, one writes.

Put it on the machine

pipx install git+https://github.com/AutoPasha/jevcode
# or: uvx --from git+https://github.com/AutoPasha/jevcode jevcode "..."

Point it at a model that decides and one that types

export TYPESAFE_API_KEY=...          # console.typesafe.ai/keys
export JEVCODE_WRITER_KEY=...        # any OpenAI-compatible endpoint
export JEVCODE_WRITER_URL=https://api.openai.com/v1/chat/completions
export JEVCODE_WRITER_MODEL=gpt-4o-mini

The writer should be small and fast. It is asked for six drafts at a time and judged on all six, so throughput is worth more here than pedigree. Any gateway speaking the same protocol works in place of TypeSafe.

Work

jevcode                      # a session here
jevcode -c                   # carry on where you left off
jevcode run "rename the --verbose flag to --loud everywhere"
jevcode where "the retry backoff"     # find code, two requests
jevcode plan "add a discount to Cart.total"  # three moves ahead
jevcode stats --days 7       # what it has cost you

Honestly

What it is not good at.

Jev is a System One model and its rough edges are documented by the people who trained it. It reads literally, it does not count, it is not a calculator, and accuracy drops as irrelevant detail piles into the state. Everything numeric in this agent is done in Python for that reason.