Jev vs OpenAI Decisions: 2,094 Real Eval Checks
Does it matter what System One model you should be using? On same tasks, how well do Jev and OpenAI's Decision API agree, and which one turns out to be stronger
Does it matter what System One model you should be using? On same tasks, how well do Jev and OpenAI's Decision API agree, and which one turns out to be stronger
The cheapest LLM-as-judge is not doing the judge at all! While the Jev output is not on-par with the LLM-as-judge, using it with escalation gives us great cost saving with small amount of difference in verdicts
The original TDD, test modules, one test per unit of code, is dead now that LLMs write most of the code. But TDD as a discipline didn't die; it moved up an abstraction layer
What happens when same eval is changed from passive grading to an agent that can investigate the whole run (and previous runs) and base the score on that. For my tests, it was massive improvement on the overall scoring, reducing the common mistakes LLM-as-judges did.
System One models like Jev can be used in evals to give confidence signal for LLM-as-judge results, which can be easily used to debug eval setup problems.
What needs to be true, to automate work to the agents, while knowing the output is what we wanted.
My exploration of harness engineering, and the primitives that allow the agents to do work on loop even when designing harnesses
I wanted to see how far I could push autonomous code generation. Not by writing better prompts, but by building a system where agents could implement, verify, and fix their own work without me watching. A DOCX editor built from behavioral specs and pixel diffs was the testbed.
How come we have advanved agent tracing products with minimal setup, but prompt optimization itself needing rewriting your product in new frameworks in order to work. My attempt on showing that it does not need to be the case
How to create AI systems that adapt and improve their performance over time by learning from past interactions. Going over the MemGPT and Letta AI frameworks approach on bulding agents with memory systems that evolve through conversations.