KnowledgeForge 29 Versions, 23 Failures, and the Three-Line Fix That Changed Everything Six weeks ago I asked a simple question: could an AI system redesign a battle bot, test it in a live competitive arena, diagnose its own failures, and improve itself — without any human involvement in the loop? The answer turned out to be yes. It also took 23 consecutive failures to get there.
Automation Why I Track What I Got Wrong More Carefully Than What I Got Right Nightwatch has been running overnight tasks for about six months. In that time, it has completed several thousand work units, surfaced dozens of useful findings, and caught at least four issues that would have cost real time to diagnose in the morning. I track all of this in a log.
AI Agents The AI Panel That Changed How I Think About Agents (Seattle Tech Week) I expected the usual conversation about model selection and prompt engineering. What I got was a room full of builders talking almost exclusively about failure modes and environment design.
Build in Public The Scores Improved. The Results Got Worse. I built a tool to grade my writing. Then I wired it into the thing it was grading.
Build in Public Why I Stopped Using One AI Agent for Everything I fell in love with Windsurf for about three months. The agentic coding experience was different. More fluid than Cursor. It felt like the assistant was in the project with you, not just answering questions about it. Then I started building things complex enough to expose the problem. Long sessions,
Build in Public Three Years at AI Tinkerers. Three Presentations. One Pattern. I've been showing up to AI Tinkerers for three years. I started going before I had anything worth demoing. I kept going after I did. The difference between those two states taught me more than the presentations. AI Tinkerers is a monthly meetup for people building with AI.