Posts
The eval kept finding bugs in Loco
We built an eval to grade agents. It graded us. What it found, the regression we shipped chasing a score, and why the generated app turned out to be the best prompt we have. Part 5 of teaching agents Loco.
Ten eval runs and an honest null
We built an eval to prove the Loco skill makes agents better. Ten runs later, the headline number never moved. What moved was which mistakes agents made. Part 4 of teaching agents Loco.
Writing the doctrine down
Loco follows the Rails doctrine and diverges only where Rust forces it. An agent can't infer that, so we wrote it into a file. Part 3 of teaching agents Loco.
What's in the Loco skill
llms.txt didn't reach anyone, a full dump was too big to reason over, and a generated API index turned out to be a hallucination detector. Part 2 of teaching agents Loco.
Rails had twenty years of training data. Loco has a skill.
Coding agents are good at Rails because Rails has two decades of public code behind it. Loco has none of that. Part 1 of how we taught agents to write Loco.