The Prompt Is the New Spec: Why Vague Prompts Become Expensive Bugs
🎯 Who this is for: Developers who prompt AI agents every day, and the leads and analysts who used to write the requirements those developers worked from.
Series: Part 2 of 13 — Enterprise Vibe Coding | Read time: 6 minutes
A bank's operations team asks for something simple: "Flag any transfer over $10,000 for review."
A developer pastes that sentence into an AI agent. Ten minutes later there is clean code, a passing test and a tidy pull request. It goes live on Thursday.
On Friday, compliance calls. The rule was meant to catch combined transfers from one customer in a single day, because splitting a payment into smaller pieces is the oldest trick in the book. It was meant to use the converted amount for foreign currency. And "for review" meant a specific queue with a specific owner, not an email to a shared inbox.
The AI did nothing wrong. It built exactly what it was asked. The prompt was the spec, and the spec was one sentence long.
📝 Your Prompt Already Is a Spec
For decades, the requirement and the code were written by different people at different times. An analyst wrote "flag transfers over $10,000," and a developer, who had sat in the meetings and knew the business, filled in the gaps with judgement.
AI agents remove that middle step. Whatever you type goes almost straight to code. The agent has no memory of the meeting and no instinct for your regulators. When the prompt is silent, the model fills the silence with the most statistically likely answer, which is rarely your company's answer.
The Bob team at IBM puts this sharply on its blog: "The code is the cheap part." Generating it costs almost nothing now. Deciding precisely what it should do is where the value, and the risk, has moved.
🧩 The Five-Part Enterprise Prompt
You do not need a 40-page document. You need five things the agent cannot guess. Here is the transfer example done properly:
| Part | Question it answers | Transfer example |
|---|---|---|
| Goal | What outcome, in business terms? | Detect possible structuring of payments |
| Context | Where does this live and what exists already? | TransferService, existing ReviewQueue class |
| Rules | What must always be true? | Sum per customer per calendar day, in USD equivalent |
| Edge cases | What will break a naive version? | Reversals, currency conversion, midnight boundary |
| Done means | How will we know it works? | Unit tests for each edge case, item lands in AML queue |
Five lines like that turn a one-sentence wish into something a reviewer can check. They also make the AI's job easier: fewer guesses, fewer rewrites, smaller diffs.
💡 Key insight: If you would not hand the prompt to a new contractor on their first day and expect correct work, do not hand it to an AI agent either. The agent is fast, but it has the same blind spots as a stranger to your business.
🛠️ How the Tools Are Turning Prompts into Specs
Vendors noticed the same problem and built answers into their products:
- Kiro (Amazon's spec-driven editor, which has replaced Amazon Q Developer) turns a prompt into three files per feature:
requirements.md,design.mdandtasks.md. Requirements use a structured "WHEN this happens, THE SYSTEM SHALL do that" format, and each step waits for your approval before moving on. - IBM Bob has a Plan mode that writes out the approach before Agent mode changes any files. It also supports "Literate Coding," where you write instructions as comments in the code and Bob returns an inline diff. The Bob team's recommended cycle is Explore, Plan, Implement, Verify, and its advice is short: "Iterate, do not one-shot."
- Claude Code and Cursor both offer plan-first workflows where the agent proposes steps you can edit, and both read project instruction files so standing rules do not have to be retyped.
- GitHub Copilot reads repository instruction files too, so team conventions travel with the code rather than living in someone's head.
The names differ. The idea is identical: make the spec a visible, reviewable thing before code exists. We go deeper on planning modes in Part 4 and on instruction files in Part 11.
⚖️ Match the Spec to the Risk
Not every prompt needs a ceremony. A good rule of thumb for enterprise teams:
- Throwaway or internal-only (a script to rename files, a one-off report): a clear one-liner is fine. Vibe away.
- Team feature (a new screen, a new API field): use the five-part prompt and read the plan.
- Money, safety or compliance (payments, billing, outage credits, patient data): write the spec, get it reviewed like any requirement, and keep it with the change record.
Kiro builds this in with a Quick Spec option that skips approval gates for well-understood work. That is the right instinct: the overhead should scale with what is at stake.
⚠️ The Honest Part
Specs do not make AI trustworthy on their own. A detailed prompt can still produce confident, wrong code, and longer prompts can lull reviewers into skimming the output. Independent research is a useful cold shower here. A 2025 randomized trial by the research group METR found experienced open-source developers took 19% longer on tasks when allowed to use AI tools, yet afterwards believed the tools had made them about 20% faster. It was one study, with early-2025 tools, but the gap between feeling and measurement is worth remembering.
The fix is not to abandon specs. It is to treat the spec as the start of review, not a substitute for it. Write it, let the AI plan against it, and still read what comes back.
Key Takeaways
- With AI agents, the prompt is the requirement. Whatever you leave out, the model fills with a guess.
- A solid enterprise prompt covers five things: goal, context, rules, edge cases and what "done" means.
- Kiro's spec files and the plan modes in Bob, Cursor and Claude Code all exist to make the spec reviewable before code is written.
- Scale the effort to the risk: one-liners for throwaway work, reviewed specs for anything touching money, safety or compliance.
References
- Kiro Documentation — Specs
- IBM Bob Blog — Getting the most out of Bob
- IBM Bob Blog — Announcing IBM Bob
- IBM Think — What is agentic coding?
- METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
Series Navigation
| Previous: | Part 1 — Vibe Coding Grew Up |
|---|---|
| Next: | Part 3 — Autocomplete → Chat → Agents |
🧰 From TheMaximoGuys toolbox: For Maximo teams, the best "context" you can give an agent is an accurate API map. Max_Interfaces is our open-source (MIT) library of 2,439+ Maximo endpoints as Postman and OpenAPI 3.0 specs. Drop it into your repo and Claude, Bob, Copilot or Cursor can work from real contracts instead of guessing field names.
Published by TheMaximoGuys | August 2026



