The CFO of a mid-sized manufacturer slides a printout across the table. It's a vendor slide. One big number in the middle: "10x faster."
"If this is true," she says, "why do we need half the dev budget we asked for?"
Every engineering leader is getting some version of this question in 2026. The honest answer is: that number might be true, for something. The job is figuring out what that something is, and whether it looks anything like your work.
📊 The Numbers on the Table
Let's use IBM Bob as our case study. Not because IBM's numbers are worse than anyone else's (they aren't), but because IBM published several different kinds of numbers, which makes it a perfect teaching example. Every figure below is IBM-reported.
| Claim | What it actually is | Scope |
|---|---|---|
| 45% average productivity gain | Self-reported survey of internal IBM users (80,000+ at GA) | All kinds of work, IBM staff only |
| 69% time savings | IBM Maximo dev team "estimated" savings; one sentence, no method published | Code generation and refactoring tasks "that normally take days" |
| 70% less time | IBM Instana survey: "average 70% reduction in time spent on selected tasks," about 10 hrs/week | Selected tasks only |
| 10x (30 days to 3) | Blue Pearl customer story: a Java 11 to 21 upgrade, 160+ hours saved | One well-defined migration project |
| 10x faster | APIS IT customer story: architecture analysis and documentation of legacy code | One public-sector documentation effort |
| 20-40% faster | IBM Z page: complex engineering work (50-80% less effort on structured workflows) | Mainframe work |
Look at that spread: 20% to 10x, from the same product. These numbers don't contradict each other. They just measure completely different things.
🔍 Four Questions to Ask Any Number
Before any productivity figure goes into a business case, run it through these four questions.
1. Who measured it? A vendor measuring its own product isn't automatically wrong, but it's not independent. The Register put it sharply: IBM is the "latest to lean on its own staff to 'prove'" AI efficacy. That applies to plenty of other vendors too.
2. How was it measured? "Developers told us in a survey" and "we timed it against a baseline" are very different evidence. Self-reported surveys measure perceived productivity. People who like a tool tend to feel faster, and feeling faster isn't the same as shipping more working software. The best-known cautionary data point is a METR study, which IBM itself has cited: developers were 19% slower with AI tools while believing they were 20% faster. One study, one setting, but it's a sharp warning about trusting surveys alone.
3. What work was in scope? "Selected tasks" is doing heavy lifting in the 70% claim. Every team has tasks where AI shines (boilerplate, test scaffolding, version upgrades, documentation) and tasks where it barely helps (untangling an ambiguous requirement, a two-day production incident). A number from the first group says little about the second.
4. What happened afterward? Speed at the keyboard is easy to measure. Review time, rework and bugs that reach production are harder, and they're where hidden costs live. Blue Pearl's story, to its credit, reports zero post-deployment defects and test coverage going from 0% to 92%, which is exactly the kind of downstream detail worth looking for.
And always check the baseline. IBM's April GA release describes Blue Pearl's job as "a typical 30-day Java upgrade in just 3 days." IBM's July release describes Blue Pearl work "originally projected to take nine months with 14 engineers" done "in just three days." Those may well be different projects, but the "before" figure is what makes any multiplier, and here two IBM releases frame it very differently.
🧭 The analyst view: When IDC assessed Bob at general availability in May 2026, it noted IBM "reports strong internal adoption and early external customer results, though the external evidence base remains limited at GA." That isn't a knock on one product. It describes almost the entire AI coding market: lots of vendor numbers, very few independent, controlled studies.
🎯 Why the Big Numbers Cluster Where They Do
Notice where the 10x numbers come from: a Java version upgrade and documentation of legacy code. Not "building a new claims portal."
That isn't a coincidence. AI coding agents are at their best on work that is:
- Well-defined (move from Java 11 to Java 21, document this JCL job),
- Repetitive (127 deprecated API calls, each fixed the same way),
- Verifiable (it compiles, the tests pass, the doc matches the code).
Modernization and legacy work tick all three boxes, which is why Part 6 called legacy code "AI's best job." Greenfield product work, where the hard part is deciding what to build, ticks almost none of them.
So when your CFO waves "10x," the right answer is: "Yes, plausibly, on our migration backlog. Not on everything."
🧪 Run Your Own Pilot
The only number you should put in a budget is one you measured yourself. Here's a lightweight pilot design that works for a bank, a utility or a manufacturer alike.
- Pick a fixed task set. Ten to twenty real, representative tasks across categories: a bug fix, a small feature, a refactor, a test-writing job, a legacy-code explanation. Same tasks for everyone.
- Capture a baseline first. How long do these kinds of tasks take today, from ticket start to merged PR? Pull it from your existing tooling before anyone gets the AI.
- Use a control group. Two comparable teams or task halves, one with the tool and one without, for four to eight weeks. Without a comparison, you're measuring enthusiasm.
- Measure outcomes, not keystrokes. Track cycle time, review time, rework (how often PRs bounce), escaped defects and cost per task (licenses plus usage credits). Survey sentiment too, but as a separate column.
- Report by task type. "38% faster on test writing, 5% on feature work, slower on incident fixes" is far more useful, and more believable, than one blended number.
Expect the result to be lower than the vendor slide. That's fine. A modest number you can defend beats a big number you can't. And you'll be ahead of most peers: IBM's Institute for Business Value reports that 79% of executives see AI productivity gains, but only 29% can confidently measure ROI.
Key Takeaways
- Read the method before the headline: most AI coding numbers are vendor-reported, self-reported, or from a single project.
- Big multipliers come from narrow work: upgrades and documentation produce 10x stories; broad averages land far lower.
- The independent evidence is thin: IDC's "evidence base remains limited" applies market-wide, not just to Bob.
- Measure downstream costs: review time, rework and escaped defects matter as much as typing speed.
- Your pilot beats their slide: fixed tasks, a baseline, a control group, results by task type.
References
- Introducing IBM Bob (IBM Newsroom, Apr 2026)
- IDC: IBM Bob Advances IBM's Position in Agentic SDLC Development (May 2026, PDF)
- IBM's AI coding 'partner' Bob hits general availability (The Register, Apr 2026)
- IBM Advances Enterprise AI Software Development with Multi-Agent Capabilities (IBM Newsroom, Jul 2026)
🧰 From TheMaximoGuys toolbox: Need a fixed task set for a Maximo pilot? Max_autoscripts is our open-source (MIT) collection of 67 Jython sample scripts, 28 templates, coding standards and Java-to-Jython conversion guides. It makes a ready-made, consistent benchmark for comparing Claude, Bob, Copilot, Cursor, or any agent on the same work.
Series Navigation
| Previous: | Part 9 — Receipts, Please |
|---|---|
| Next: | Part 11 — Teaching the AI Your House Rules |
Published by TheMaximoGuys | September 2026



