>
eval_set: refund_qa · 240 examples
metric: rubric · pass-rate
OPTIMIZING PROMPT...
. . ✶ . . . ✶ ✶ ✶ . . ✶ ✶ ✶ ✶ . . ✶ ✶ ✶ . . . ✶ . .
v3 -
v5 -
v7 -
Latest Writing
Thoughts on software development, AI, and building systems
Escalating Evals: 72% Fewer LLM Judge Calls
The cheapest LLM-as-judge is not doing the judge at all! While the Jev output is not on-par with the LLM-as-judge, using it with escalation gives us great cost saving with small amount of difference in verdicts
TDD in the age of agents
The original TDD, test modules, one test per unit of code, is dead now that LLMs write most of the code. But TDD as a discipline didn't die; it moved up an abstraction layer
Agent-as-Judge: Harder Evals, More Computation
What happens when same eval is changed from passive grading to an agent that can investigate the whole run (and previous runs) and base the score on that. For my tests, it was massive improvement on the overall scoring, reducing the common mistakes LLM-as-judges did.