One-Shot Prompting Is a Myth. Engineering Is a Loop

Most disappointed AI projects share an assumption that is rarely said out loud: that building software with AI is a one-shot act. One-shot prompting: write the prompt, get the thing, done. It is an easy belief to hold, because that is exactly what the demo looks like. You watch a paragraph of English turn into a working feature, and the obvious conclusion is that the machine now writes the software. 

The practitioners who are actually accountable for what ships have converged on the opposite conclusion. AI engineering is iterative, incremental, and planning-heavy. The work did not disappear when the typing got cheap — it moved to the two ends. Deep planning went to the front, hard verification went to the back, and the middle, the part everyone watches, is the one part that genuinely got easy. If you only look at the cheap middle, it looks like magic. If you are on the hook for the result, it is a disciplined loop. 

That is the split in the industry: some people think AI writes the software; others understand that AI generates candidates, and a disciplined loop turns those candidates into software. 

Figure 1 — When the middle of the work got cheap, effort didn't vanish; it migrated to planning and verification.
Figure 1 — When the middle of the work got cheap, effort didn’t vanish; it migrated to planning and verification. 

Not hype, and not too late 

Part of the confusion is that this feels simultaneously old and brand new — and both readings are correct. 

The ideas are years old and proven. Retrieval-augmented generation dates to 2020. Tool-use and reasoning loops — the ReAct pattern — landed in 2022–2023. Agents and the “harness” idea followed in 2023–2024. The Model Context Protocol was published in late 2024. None of this is a 2026 invention; it has been documented across more than two dozen practitioner books over the last two years. 

What is genuinely new is the ability to configure it yourself, in real tools, and trust it. Claude Code went from preview to general availability across 2025. GitHub’s Copilot agent mode began rolling out to stable in April 2025. Agent Skills and the SKILL.md convention arrived in October 2025, and through 2026 the pattern went mainstream — a skills marketplace, industry plugins, cloud-hosted managed agents, multi-agent workflows. 

The unlock underneath all of it is subtle. Models from late 2025 onward follow written guidelines and guardrails reliably enough that a governed, file-based agentic setup became production-viable. Before that, even a good operating model couldn’t be trusted at higher autonomy — the model would drift past the rules. So the honest framing is neither hype nor lateness: the operating model isn’t new theory. It’s newly buildable. That’s the “why now.” 

Figure 2 — The ideas are proven and years old; the ability to build the operating model yourself arrived only in the last ~18 months.
Figure 2 — The ideas are proven and years old; the ability to build the operating model yourself arrived only in the last ~18 months. 

The loop one-shot prompting never sees 

The real shape of the work is a cycle: plan deeply → generate fast → verify hard → own the result → repeat.

“One-shot prompting” is simply this loop with four of its five steps deleted. It sees only the generate box, because that is the only part the demo shows. 

Figure 3 — The engineering loop. One-shot sees only the cheap GENERATE step; the value lives in the steps around it.
Figure 3 — The engineering loop. One-shot sees only the cheap GENERATE step; the value lives in the steps around it.

Three findings explain why deleting the rest fails at scale. 

The first is the 70% problem, named by Addy Osmani: AI reliably nails the first ~70% of a task — the scaffolding, the boilerplate, the common patterns. The remaining 30% is judgment. Edge cases, architectural fit, security, verification. That 30% is inherently iterative; there is no prompt that skips it, because the difficulty was never in the typing.

The second is that the bottleneck moved. Code can be produced at roughly 1,600 lines a day and safely reviewed at only about 400 lines an hour — four hours of careful review for every day of output. The scarce step is no longer writing — it is verifying. Adoption is not the question: 84% of developers now use or plan to use AI tools, yet only about a third trust the output. 

.Figure 4 — Code is produced faster than it can be safely verified. What skips verification becomes trust debt. (Rates from Vibe Engineering, 2026.)
Figure 4 — Code is produced faster than it can be safely verified. What skips verification becomes trust debt. (Rates from Vibe Engineering, 2026.) 

Because there is a cost to getting this wrong that doesn’t show up on any dashboard: trust debt — the accumulating liability of shipping AI-generated code that nobody on the team actually understands or owns.One-shot prompting is the fastest way to rack it up. The fix is a cultural one: replace dump-and-review with verify-then-merge. 

Speed in the wrong direction isn’t productivity. It’s debt accumulated faster. 

The discipline has a shape 

The thing that replaces one-shot is not “more careful prompting.” It has a name — spec-driven development, or vibe engineering — and a concrete shape. Because the model is non-deterministic and forgets everything between sessions, a short, versioned spec becomes the stable anchor the whole loop rotates around. The moves are the same in every serious take on it, and the point is to run them small and often. 

Plan. Write the destination as behavior, not architecture — roughly 500 words. Grill the agent with ten or more questions before a line is written. The highest-value sections are the two people always skip: Out of Scope and Assumptions. The spec isn’t overhead — it’s what prevents the rework later.

Generate. The agent builds from the spec in vertical slices, riskiest unknown first — a tracer bullet through the whole stack rather than a horizontal layer. This is the only step that feels like one-shot prompting, and that’s fine: the spec constrains it going in, and verification catches it coming out. 

Verify. Read the code — that checks the judgment. Check acceptance criteria — that checks the behavior. Lean on tests, evals, and CI as gates. Passing tests is not “done”; it is the exit criterion that unlocks the next loop. 

Own. Findings become new issues and feed back into generate. Update the spec first, then re-prompt — never patch around a spec that’s gone stale. You come out holding an accurate mental model of how the thing behaves and how it fails. The test for whether you’ve actually done this: would you go on-call for it? 

The incremental point is the one that’s easy to lose. Each pass ships one thin, verified slice, then loops again. Not one giant one-shot — many small verified loops. The speed comes from the loop being tight, not from skipping its steps.

The same shift, read from six chairs 

The shift is real for everyone, but it reads differently depending on where you sit. None of the readings below is wrong — each is an accurate view of the same change from a particular chair. Setting them side by side is simply what makes a shared decision possible.

Figure 5 — The same shift seen from six roles: the common view, what's easy to underestimate, and where the value matures.
Figure 5 — The same shift seen from six roles: the common view, what’s easy to underestimate, and where the value matures. 

For the CEO or board, the natural frame is a productivity and cost lever — and the lever is real. What’s easy to underestimate is the 20/80 pattern: the demo is perhaps 20% of the work, while identity, data, governance, evaluation, and adoption carry the rest. The lever compounds when it is treated as iterative change rather than a one-off purchase — with trust debt, measured in rework and incidents, watched as a real risk metric. 

For the CIO, it reads as a security, risk, and integration question. What’s easy to underestimate is that the model is rarely the hard part — workflow redesign and adoption are — and that governance is more durable as a continuous practice than a one-time gate. Identity, data, and evaluation increasingly sit in the architecture, not alongside it. 

The CTO is usually close to the mark: this is iterative engineering. What’s easy to underestimate is the culture and process the discipline asks for, and the value of establishing observability before scaling autonomy. Verify-then-merge, sandboxing, and evaluation tend to come first, and scaling agents after. 

For the Product Owner or PM, it can look like the same backlog delivered faster. What’s easy to underestimate is that the two hardest parts of a spec — Out of Scope and Assumptions — are also the most consequential. The value matures when the spec becomes a craft in its own right: intent defined precisely, and read adversarially for how it might be misinterpreted.

For the developer, it is a faster way to produce code. What’s easy to underestimate is that verification effort grows as output grows. The productive shift is a change of question — from how should this work? to how could this fail? — the move from writing code to owning it, sometimes called the trust engineer. 

For the architect, the assumption is that AI writes code while the design still holds. What’s easy to underestimate is that locally correct can still be structurally wrong — a perfect puzzle piece for the wrong puzzle. Here the role’s value tends to rise: tacit taste gets codified into enforceable rules — a constitution, invariants, an agent control model the loop can respect. 

Read together, these are not six disagreements. They are six angles on one shift — which is exactly why naming it for each chair is worth the effort before committing to a direction.

AI writes the code. You verify that it deserves to exist.

Where this leaves us 

The teams pulling ahead in 2026 are not the ones prompting fastest. They are the ones planning deepest and verifying hardest — the ones who understood that when the middle of the work got cheap, the value didn’t vanish, it migrated to the ends. One-shot prompting was never the method. It was just the part of the method that looked good on stage. 

Pick the argument for the chair you’re in. For the CEO, the gains are real — and trust debt is the risk to watch. For the CIO, governance is continuous. For the CTO, discipline comes before autonomy. For the PO, Assumptions and Out of Scope are load-bearing. For the developer, the only question that matters is whether you’d go on-call for it. For the architect, locally right can still be structurally wrong. The same truth, seen from six chairs.

Running the loop is a team skill, not a tool setting. Our AI engineering enablement programs coach delivery teams through it on their own backlog — planning, verification, and ownership included.

One-shot is a myth. Engineering is a loop. 

Sources

Books: Vibe Engineering (Lelek & Skowroński, Manning 2026) — the loop, the 70% problem, trust debt, the 1,600/400 rates, the trust engineer. Spec-Driven Development (Bezael Pérez, 2026) — the spec as primary artifact, vertical slices, tracer bullets, QA as exit criterion. Shipping Enterprise AI / the FDE playbook (Bommena, 2026) — the 20/80 reality, “identity, data, evaluation are architecture.” 

Axon Frontier Radar series: #1 From chatbots to problem solvers (state of AI agents, 2026); #2 Why AI productivity gets lost between benchmarks and the balance sheet; #3 How agentic AI is turning tokens into a business metric.

Frequently Asked Questions

Is building software with AI a one-shot act? 

No. The practitioners accountable for what ships have converged on the opposite conclusion: AI engineering is iterative, incremental, and planning-heavy. The work moved to the two ends — deep planning at the front, hard verification at the back — while only the middle, the typing, genuinely got easy. 

What is the 70% problem in AI-assisted coding? 

Named by Addy Osmani, it describes how AI reliably nails the first ~70% of a task — the scaffolding, the boilerplate, the common patterns — while the remaining 30% is judgment: edge cases, architectural fit, security, and verification. That 30% is inherently iterative; there is no prompt that skips it.

What is trust debt? 

Trust debt is the accumulating liability of shipping AI-generated code that nobody on the team actually understands or owns. One-shot prompting is the fastest way to rack it up; the fix is cultural — replace dump-and-review with verify-then-merge. 

What does the AI engineering loop look like in practice? 

Plan deeply → generate fast → verify hard → own the result → repeat. Each pass ships one thin, verified slice, then loops again. The speed comes from the loop being tight, not from skipping its steps.