Home
  • Blog
[ New ] Building a Harness: How I Got Agents to Verify Their Own Code

Hi, I'm Sami!

AI ENGINEER & BUILDER
GitHub
INPUT PORT: 01

>

eval_set: refund_qa · 240 examples

metric: rubric · pass-rate

AGENT ENGINE SANDBOX

OPTIMIZING PROMPT...

  .   .   ✶   .   .
    . ✶ ✶ ✶ .
  . ✶ ✶ ✶ ✶ .
    . ✶ ✶ ✶ .
  .   .   ✶   .   .

CPU: 4 · TOOLS: 5 MODE: apo-loop
OUTPUT EVAL

v3 -

v5 -

v7 -

Latest Writing

Thoughts on software development, AI, and building systems

View all
Building a Harness: How I Got Agents to Verify Their Own Code
Apr 19, 2026

Building a Harness: How I Got Agents to Verify Their Own Code

I wanted to see how far I could push autonomous code generation. Not by writing better prompts, but by building a system where agents could implement, verify, and fix their own work without me watching. A DOCX editor built from behavioral specs and pixel diffs was the testbed.

Harness Engineering AI Agents Code Generation
Prompt Optimization Shouldn't Require Rewriting Your App
Feb 25, 2026

Prompt Optimization Shouldn't Require Rewriting Your App

How come we have advanved agent tracing products with minimal setup, but prompt optimization itself needing rewriting your product in new frameworks in order to work. My attempt on showing that it does not need to be the case

Prompt Optimization Langfuse
Building Memory Systems that Learn Across Conversations
Jan 8, 2026

Building Memory Systems that Learn Across Conversations

How to create AI systems that adapt and improve their performance over time by learning from past interactions. Going over the MemGPT and Letta AI frameworks approach on bulding agents with memory systems that evolve through conversations.

MemGPT Agents
View all posts