AI · Automation · Engineering

Can an LLM Judge Make Coding Agent Iteration Cheaper?

By Lazar MilicevicOctober 7, 20269 min read
Lines of source code on a monitor, representing LLM-judged coding agent iteration and evaluation

The expensive part of improving a coding agent is not writing the change. It is finding out whether the change helped. A prompt tweak takes a minute to write and hours of compute to evaluate honestly, and by the time you have your answer you have already queued three more tweaks behind it. A new MIT and Sakana AI framework called SIFT attacks that loop directly, and I want to work through what it gets right, and what I would still refuse to trust it with.

The problem: evaluation is the bottleneck, not ideation

Self-improving coding agents work by modifying the things that guide them: prompts, tools, and the code of the harness itself. Proposing a patch is cheap. Judging it is not. The SIFT paper names the poor signal-to-cost trade-off of existing evaluation methods as the main bottleneck: full benchmark runs are reliable but costly, and small task subsets are cheaper but noisier (arXiv:2609.19526).

If you have ever run evals on an agent, you know both horns of this. Run the full suite and you learn the truth, slowly, at real cost. Run ten tasks and you get a number that moves by several points when you change nothing at all. The paper cites prior self-improvement runs costing on the order of $22,000 to evaluate on SWE-Bench and consuming thousands of CPU hours (citing Zhang et al., 2025b and Wang et al., 2025). The extract I read of that sentence is fragmentary, so treat the exact framing with care, but the order of magnitude matches what anyone who has scaled agent evals will recognize.

That cost shapes behavior. When each evaluation is expensive, you try fewer ideas. When you try fewer ideas, a search process has less to work with. The whole point of self-improvement is breadth of exploration, and an expensive judge starves it.

What SIFT actually does

SIFT stands for Recursive Self-Improvement via Fast Tree Search. The core move is to put a separate language model in front of the expensive benchmark and let it compare candidate agents before anyone pays for a real run (VentureBeat).

The mechanics, as the paper describes them:

  • Each new candidate is compared against incumbent agent harnesses by the judge.
  • The win-loss record is aggregated with a regularized Bradley-Terry model, which assigns every node in the search tree a global strength score (arXiv HTML).
  • Expensive downstream task evaluations are reserved for the most promising nodes only. The search is disaggregated, so judge scores guide exploration without waiting on slow evaluation runs (arXiv abstract).

I like the Bradley-Terry choice. Pairwise comparisons are a much more natural task for an LLM than absolute scoring. Asking a model "is this agent a 7 out of 10" gets you mush. Asking "which of these two is better" gets you something you can aggregate statistically, and the regularization keeps a node with two lucky wins from looking like a champion.

The disaggregation is the other important design decision. In most agent search loops, exploration blocks on evaluation. Here the cheap judge runs ahead and the expensive benchmark runs behind, only on survivors. It is the same pattern I use in content pipelines: a cheap filter in front, an expensive verifier at the end.

The numbers, and how I read them

The per-step cost breakdown in the paper is the most useful part for anyone budgeting this kind of system:

Step Approximate cost
Patch proposal about 12 cents
Pairwise judge call about 4.4 cents (up to 10 per candidate, roughly 44 cents)
Evaluating an agent on 50 Polyglot tasks about $6 and 2.6 CPU hours

Look at the ratio. Judging a candidate with up to ten pairwise calls costs roughly 44 cents. Evaluating on 50 tasks costs about $6. So the judge screens a candidate for roughly an order of magnitude less than a real evaluation, and only the best candidates earn the $6 run.

On results, the researchers report that one SIFT run reached 35.1% accuracy on Polyglot in under five hours of wall-clock time, using 42 CPU hours and about $150 in API credits (VentureBeat). The paper text also reports about 35% on Polyglot with o3-mini (under 50 CPU hours, 5 hours wall-clock) and 32% with Qwen3-30B (under 250 CPU hours, 7 hours wall-clock) (arXiv HTML).

For context, the Darwin Gödel Machine from Sakana AI, the baseline lineage here, improved from 14.2% to 30.7% on Polyglot and from 20.0% to 50.0% on SWE-bench (Sakana AI).

Now the caveats, because I would not put these numbers in a slide deck without them. The 35.1% and the "about 35%" may be different runs or roundings. I could not confirm whether it is a best run or an average, or what the variance looks like across seeds. A single run reaching 35.1% tells me the method can work. It does not tell me how often it does. I also saw a secondary source quoting a much lower dollar figure for a full search, which the primary sources do not confirm, so I am sticking with the researchers' $150. And this is a v1 preprint (submitted 17 Sep 2026), not a peer-reviewed result.

The real question: does the judge reward things that only look better?

This is the part I care about most, because it is where cheap evaluation usually goes wrong. A third-party review of the paper states the load-bearing premise plainly: the judge's pairwise preferences must correlate with true downstream agent quality, and if the judge signal is no better than random, the claimed speedups disappear (Pith Review).

I agree with that framing, and I would add a nastier version of it. A judge does not need to be random to hurt you. It only needs to be systematically biased toward something other than quality. LLM judges are known to like certain surface features. Longer diffs, more confident explanations, more "thorough-looking" scaffolding. A self-improving agent is an optimizer, and optimizers find whatever the judge likes.

The paper is candid that this applies to the evaluation harness too. Candidate patches may exploit harness weaknesses, relax constraints, or optimize for evaluation artifacts. The authors mitigate this by sandboxing execution, restricting writable files, and rejecting patches that modify benchmark files (arXiv HTML). VentureBeat reports the same thing from the practitioner side: the researchers had to block patches that loosened the evaluation harness, and downstream benchmarks are still needed to confirm that a change actually improved the agent.

Those mitigations are sensible, and they are the minimum. But here is what I could not find in the sources: a quantified measure of how often the judge's ranking disagreed with downstream results. That number is the one that tells you whether the judge is a filter or a flatterer. Until it is published, "the judge works" is a hypothesis supported by a headline result, not a measured property.

How I would test it before trusting it

If I were adopting this pattern for my own agent work, I would not start by wiring the judge into a search loop. I would first measure the judge as an instrument. Here is the procedure.

  1. Build a calibration set. Take a few dozen candidate agent variants where you already know the true downstream score from full evaluations. Include some deliberately bad ones that look good: verbose, heavily commented, large diffs that do nothing.
  2. Run the judge pairwise on all of them, with position swapped on every pair to cancel order bias.
  3. Fit the Bradley-Terry scores and compute rank correlation against the true downstream scores. This single number is your trust level.
  4. Inspect the disagreements, not just the correlation. Read the top ten cases where the judge ranked something high and downstream ranked it low. That tells you what the judge is actually rewarding.
  5. Add adversarial controls. Include a candidate that only edits comments, one that only lengthens the prompt, one that tries to touch the evaluation files. If any of them climb the ranking, the judge is gameable.
# Position-swapped pairwise judging, aggregated before ranking
def judge_pair(a, b):
    r1 = judge(a, b)      # returns 1 if first wins
    r2 = 1 - judge(b, a)  # same question, order flipped
    return (r1 + r2) / 2  # 0.5 means the judge could not decide

Treat that 0.5 as information. If a large fraction of your pairs come back undecided, your judge is not seeing a signal and the tree search will be steering on noise.

Then, once the judge passes, keep a held-out downstream set that the search never touches. The expensive benchmark is not just a final check, it is the thing that keeps the whole system honest. If the cheap judge guides the search and the same tasks also select the winner, you have built a closed loop that can drift.

Where this fits in production

I build autonomous systems that run unattended, and the self-learning content loop behind BizFlowAI ContentStudio (measure, learn, target, generate, publish) taught me the same lesson this paper is circling. A cheap proxy lets you iterate fast, and a real-world signal keeps the proxy honest. In that system the real signal is search performance. For a coding agent it is the downstream benchmark. The architecture is identical: a fast, imperfect judge for breadth, a slow, trustworthy verifier for decisions.

What I would not do is let the proxy make irreversible decisions. In an unattended system, a judge that ranks candidates is fine. A judge that promotes a change to production with no downstream check is how you wake up to a regression that nobody saw coming.

What I'd do

  • Use the pattern, verify the judge first. Cheap pairwise screening in front of expensive evaluation is a sound design. Measure rank correlation on a calibration set before you let it steer anything.
  • Never let the agent write to its own evaluation. Sandbox execution, lock the writable paths, and reject any patch touching benchmark or harness files. The paper does this, and so should you.
  • Keep a held-out set the search cannot see. The judge guides exploration, the held-out benchmark decides what ships.
  • Add adversarial controls to your judge tests. If a comment-only edit can climb the ranking, you have your answer.
  • Wait for the disagreement data. The headline numbers are promising, but I want the judge-versus-downstream disagreement rate and multi-seed variance before I call this a solved problem.

I am genuinely optimistic about this direction. Cutting the cost of a trustworthy signal is what makes self-improving agents practical outside well-funded labs. Just remember that the cheaper the signal, the more carefully you have to check what it is a signal of.

If you are building agent evaluation or self-improvement loops and want to compare notes, you can reach me at lazar-milicevic.com/#contact, or keep reading on the blog for more hands-on write-ups from production systems.

Frequently asked questions

What is SIFT and how does it make coding agent evaluation cheaper?

SIFT (Recursive Self-Improvement via Fast Tree Search) is an MIT and Sakana AI framework that places an LLM judge in front of the expensive benchmark. The judge compares each new candidate agent against incumbent harnesses, and the expensive downstream evaluations run only on the most promising nodes. Because judging a candidate costs roughly 44 cents versus about $6 for a 50-task run, most weak candidates never reach the costly evaluation. In my read, the saving comes from filtering cheaply first and verifying expensively last.

How does SIFT use a Bradley-Terry model to rank coding agents?

SIFT collects pairwise win-loss results from the LLM judge and aggregates them with a regularized Bradley-Terry model. That model assigns every node in the search tree a global strength score. Pairwise comparison suits LLMs far better than absolute scoring, since asking which of two agents is better yields signal you can aggregate statistically. The regularization stops a node with a couple of lucky wins from looking like a champion.

How much does it cost to evaluate a self-improving coding agent?

Full evaluation is expensive: the paper cites prior self-improvement runs costing on the order of $22,000 on SWE-Bench and thousands of CPU hours. In SIFT's own breakdown, a patch proposal costs about 12 cents, a pairwise judge call about 4.4 cents (up to 10 per candidate, roughly 44 cents), and evaluating an agent on 50 Polyglot tasks about $6 and 2.6 CPU hours. One reported SIFT run reached 35.1% on Polyglot in under five hours using 42 CPU hours and about $150 in API credits.

Can you trust an LLM judge to decide which coding agent is better?

Only to the extent that its pairwise preferences correlate with true downstream agent quality. If the judge signal is no better than random, SIFT's claimed speedups disappear, so that correlation is the load-bearing premise. The main risk is the judge rewarding changes that merely look better rather than perform better. I would use it as a cheap filter and still keep real benchmark runs as the final verifier.

What results did SIFT report on the Polyglot benchmark?

The researchers report one SIFT run reaching 35.1% accuracy on Polyglot in under five hours of wall-clock time. The paper text also reports about 35% with o3-mini (under 50 CPU hours, 5 hours wall-clock) and 32% with Qwen3-30B (under 250 CPU hours, 7 hours wall-clock). For context, the earlier Darwin Godel Machine improved from 14.2% to 30.7% on Polyglot. These come from a v1 preprint, and it is unclear whether the figures are best runs or averages or what the variance across seeds looks like.

Lazar Milicevic

Lazar Milićević

Senior Technical Engineer. I build AI automation, GenAI/LLM systems and cloud architecture — autonomous systems that run while you sleep. Founder of BizFlowAI.

Building something hard with AI or automation? I am open to talk.

Get in touch

← All posts