Development experiments
Agent Development Lab
I use AI agents to build, review, and debug. These experiments test whether what they learn on one problem helps on the next.
Every attempt leaves notes behind: what we assumed, what we ruled out, what finally explained the bug. Some of that is worth keeping. Some of it just happened to work once, and turning it into a rule would make the next attempt worse.
Writing lessons down
Each proposed lesson stays linked to the attempt it came from, so we can go back and see what actually happened.
The real test is an unfamiliar problem. Does the lesson help an agent ask a better question, tell two possible causes apart, or notice that its explanation is wrong?
Testing them
So we run comparisons: a different problem, with and without the guidance, and a way to check what the agent actually got right.
We also examine how the agents work together. In one comparison, a group of specialists finished faster and found an important problem a single reviewer missed. They found the same number of verified problems overall and used more resources. There was value in the result, but it did not meet the rule we had set for expanding the approach.
Even the comparison needed work. An earlier attempt had given the agents different surrounding context. We had to correct that before we could make sense of the difference between their results.
A closer look at that comparison
The specialist group used about 32% more processed tokens, the units of text handled by the models, for the same three verified findings. One of its findings was an important problem the single reviewer had missed. The cost rule had been set before the comparison; the result did not meet it, so we held the proposal.
This was one comparison on one target. Processed tokens were a way to compare resource use, not a measured bill. A later experiment intended to test lessons on unfamiliar problems is still unfinished. It has not established that the proposed guidance improves future work.
Why it matters
This only matters if it helps the person using what we build get from what they want to do to having it done, faster.
That includes updates. Someone may be in the middle of their work when the software changes. Keeping their work safe, checking what they can actually use, and being able to roll back a bad release are part of the job.
The methods are still changing. The question for each new step is whether it helps us understand the problem or just adds process.
Based on development records from August and September 2026. Specific results and their limits are included in the comparison above.