Journal · AI
Lessons from building with AI agents
AI coding agents are fast, but speed isn't the same as correct. What I've learned about small scopes, independent checks and who owns the decisions.
I build with AI coding agents most days now. They write a lot of the code, and I wouldn’t go back. But I’ve also watched an agent report “all done, everything passes” about work that didn’t run.
That gap between what an agent says and what is true is where most of my lessons come from.
Fast is not the same as finished
An agent can produce a working-looking feature in minutes. That speed is real and it changes how I plan. It also makes it easy to stop looking.
The code compiles, the summary sounds confident, the diff looks tidy. None of that tells you it’s right. An agent grading its own work is like a student marking their own exam: not dishonest, just not a reliable second opinion.
So I treat the agent’s account of what it did as a claim to be checked, not a result.
The check has to be independent
For me, work counts when something other than the agent agrees with it. That means the test suite, the type checker, a clean production build, and sometimes a separate review pass from a different session or a different model that hasn’t seen the original reasoning.
The point is independence. If the agent writes the code and also writes the tests that check it, a bad assumption can pass both. I’d rather have a test that existed before the change, or one I’ve read myself, because those can fail for reasons the agent didn’t anticipate.
This sounds like extra work. It is. It’s also cheaper than finding out in production that a reminder email goes to the wrong person.
Keep the scope small
The best results come from narrow tasks with a clear finish line. “Add a budget field to the event form, with a test for the empty case” goes well. “Improve the event experience” produces a lot of confident changes in places I didn’t ask for.
Small scopes make review quick. I can read a diff in a few minutes and know what changed. When an agent touches twenty files, I either trust it blindly or spend an afternoon, and neither is a good deal.
I also run several things in parallel when they don’t overlap. That works only because each piece is small enough to verify on its own.
Humans own taste and decisions
Agents are good at producing options and bad at knowing which one fits. They don’t know that I want the capture box to feel nearly empty, or that a certain wording sounds too eager for this product. Those calls come from the person who has to live with the result.
I decide what gets built, what it should feel like, and what gets cut. The agent helps me get there faster. When I let that slide, the product starts to feel generic, like it was assembled from the average of everything.
I also read the code that handles anything sensitive, such as email sending or sharing permissions. Delegating the typing is fine. Delegating the responsibility isn’t.
A short version
If I had to boil it down: give the agent a small job, make it prove the job is done with a check it didn’t write, and keep the judgment calls for yourself.
None of this is clever. It’s closer to ordinary engineering discipline with a faster pair of hands added. I keep relearning it, usually right after the one time I skipped it.