- The essence of an LLM loop isn’t “autonomy”; it’s “verifiable iteration.”
- It’s not that the agent is smart and will eventually solve it if you let it run.
- You need a mechanical way to judge success/failure at every iteration.
- In a good loop, the exit condition matters more than the prompt.
- You need a clear stop signal like “tests pass,” “compile succeeds,” “metric improves,” or “zero exit code.”
- Without an exit condition, a loop is closer to gambling than engineering.
- Separate what the agent decides from what the code must enforce.
- LLM: ambiguous judgment, code generation, diagnosis.
- Deterministic code: execution, verification, logging, retry caps, rollback, diff checks.
- An agent loop is ultimately a software architecture problem.
- Loops where an agent evaluates itself are risky.
- It can keep rationalizing within the same flawed logic.
- It’s safer to separate the task agent and the evaluator agent.
- The biggest risk isn’t cost; it’s humans losing understanding.
- The faster an agent produces code, the less people know “why it ended up this way.”
- When a production failure happens later, debugging becomes very difficult.
- Planning loop → Execution loop → feed the results back into planning. Separating deciding what to build from actually building it was reportedly the single biggest efficiency gain.
- Humans don't review all the code; they only approve chokepoints like product direction, scope changes, and infra/data model decisions.
- The principle is: when agent output is bad, don't fix it by hand — fix the harness/rules and re-run. Improve the machine, not the output.
- Wherever possible, don't leave verification to the LLM; enforce it with deterministic checks like lint, type systems, static analysis, and executable specs.
- In the planning stage, multiple agents repeatedly critique the spec, applying "shift-left validation" to catch as many problems as possible before implementation.
- Implementation runs in parallel according to a ticket dependency graph → separate agent review/E2E → integration tests → deploy → prod smoke tests, all automated. Failures and fixes flow back into the planning loop.
Loops, graphs & harnesses – getting quality out of a software factory
Deep-dive on harnessing in agentic coding workflows. I'll share how one code factory can be built, with real examples, design principles and ways to not do it.
https://www.ivokund.com/loops-graphs-harnesses-getting-quality-out-of-a-software-factory/

Matt Van Horn on Twitter / X
https://t.co/DM0CAuyprS— Matt Van Horn (@mvanhorn) June 8, 2026
https://x.com/mvanhorn/status/2063865685558903149
Closing the Software Loop
How coding agents are changing software development, from feature requests to autonomous deployment with minimal human intervention.
https://www.benedict.dev/closing-the-software-loop

Seonglae Cho