01
We gave three coding agents the same broken production task
Same repository, same acceptance criteria, same time box. The difference was not where we expected it.
STATUSPUBLISHEDOPEN LOG →
EXPERIMENTS / TEST LOG
Real tools, real workflows and measured outcomes from the edge of practical AI.
01 / TEST LOG
Same repository, same acceptance criteria, same time box. The difference was not where we expected it.
Browser Use · 24 attempts · 79% completion
Same tool schema, same evaluation set, different model sizes and serving setups. We measured reliability before speed.