It's Tuesday morning at a regional insurer. One developer is renaming a variable across forty files. Another is asking the AI to explain a 2,000-line claims-adjudication module nobody has touched since 2011. A third is planning how to split a monolith into services.

All three are using the same AI coding tool. Should all three be talking to the same AI model?

Almost certainly not. And how your tool answers that question says a lot about who is really in control of your development pipeline.

🧠 Why One Model Can't Do Every Job

Think of AI models like vehicles in a utility company's fleet. You don't send the bucket truck to pick up printer paper, and you don't send the hatchback to repair a transmission line.

Coding work splits the same way:

  • Small, frequent tasks (completions, renames, docstrings) want a model that is fast and cheap. Waiting eight seconds for a one-line suggestion kills flow.
  • Big, gnarly tasks (cross-file refactors, architecture plans, legacy code archaeology) want a model that is deep, even if it is slower and pricier.
  • Sensitive tasks (anything touching customer data or regulated code) want a model that runs where your compliance team says it can run, which may rule out the best model on the market.

Speed, depth, cost, and data location rarely point at the same model. So the real question isn't "which model is best?" It's "who decides which model handles which job?"

🔀 Three Ways Tools Pick a Model

Today's tools fall into roughly three camps.

Camp 1: You pick. Cursor and GitHub Copilot put a model menu in front of the developer. Claude Code lets you switch between Claude models. Maximum control, maximum responsibility: every developer becomes a part-time model-selection analyst.

Camp 2: The tool picks. IBM Bob's SaaS edition routes work automatically across models including Anthropic Claude, Mistral, IBM Granite and fine-tuned models. Per IBM's own FAQ, users cannot pick or restrict which model runs. IBM's July 2026 release went further, saying Bob now "matches models to tasks, coordinates AI execution across agents." The marketing tagline sums up the philosophy: "Stop managing models. Start managing outcomes."

Camp 3: You pin one model you host. For air-gapped and sovereign setups (see Part 7), choice shrinks to whatever you can run inside the fence: typically one open-weight model served on your own GPUs (often via an inference server like vLLM), sized to your hardware. Routing, if you want it, is yours to build. IBM has said an on-premises Bob option is planned, but for now this camp is mostly DIY open-model setups.

ApproachExample toolsYou gainYou give up
Developer picksCursor, GitHub Copilot, Claude Code (within Claude)Control, transparency, pinning for auditsDecision fatigue, inconsistent choices across teams, cost surprises
Tool routes automaticallyIBM Bob (SaaS)Simplicity, optimized speed and costVisibility into which model wrote what
One self-hosted modelOpen models on your own GPUs, other on-prem setupsData sovereignty, predictabilityBest-model-for-the-job flexibility

None of these is "right." Each is a different answer to the same trade-off.

⚔️ The Double-Edged Sword

When Bob hit general availability in April 2026, RedMonk analyst Kate Holterhoff put it neatly. Automatic model selection, she said, is "a double edged sword, as developers can be suspicious of black box tools," while it also "eliminates the paralysis of choice."

Both halves are true, and enterprises feel both.

The upside is real. Most developers do not want to benchmark five models before writing a unit test. A sensible default that sends cheap work to cheap models can trim cost without anyone thinking about it. IBM's July announcement framed the problem directly: many enterprise engineers are manually choosing models while trying to balance cost against performance.

The downside is also real. Mainframe IBM Champion Uwe Graf flagged opaque model choice as a concern in practice. If you work in a bank and your model-risk team asks "which model generated this code that moved into production?", "the router decided" is not a great answer. Some regulated shops want to pin a model version for exactly that reason.

There is also a pricing wrinkle. Bob bills in Bobcoins (roughly $0.50 each), and IBM doesn't publish how many tokens a Bobcoin buys. When the tool picks the model and the conversion rate is opaque, forecasting spend becomes guesswork until you have a few months of real usage data. Plans range from $20 Pro to custom Enterprise; Cursor and Copilot sit in a similar $10 to $40 per-seat band.

And even perfect spend data only tells you half the story. As IBM's Adam McDaniel put it in an IBM Think piece: "Token consumption is a cost signal, not a value metric." A router that saves you 30% on tokens is only a win if the cheaper model's output doesn't cost you more in review and rework.

💡 The honest take: Automatic routing isn't a feature or a flaw; it's a policy choice someone else made for you. Before you roll out any multi-model tool, ask the vendor three questions: Can we see which model handled each request? Can we pin or exclude a model? Can we get spend broken down by model? If the answer is "no" three times, make sure your risk team is fine with that.

🧭 How to Choose Your Camp

A quick decision guide for leads and architects:

  1. Highly regulated, audit-heavy code (payments, trading, safety systems): lean toward pinned or self-hosted models. Predictability beats cleverness.
  2. Broad internal teams with mixed skill levels: automatic routing often wins. Fewer knobs, fewer bad choices, easier support.
  3. Senior platform teams and AI-curious developers: give them a picker. They will find the right model for weird jobs faster than any router.
  4. Most real enterprises: all of the above, by team. Analyst Mitch Ashley of Futurum even suggested "the best answer may well be a combined IBM plus Claude team" for Bob shops. Mixing tools is normal.

Whatever you choose, write it down. "Team X uses auto-routing, Team Y pins Model Z, here's why" is a two-paragraph policy that will save a very long meeting later.

🔌 Keep the Plumbing Model-Agnostic

Here's the part people miss. Models will change. Routers will change. Your vendor's default model this quarter will not be its default next year.

What shouldn't change is how your AI tools reach your enterprise systems. If your integration with the ERP, the asset system, or the ticketing platform is hard-wired to one assistant or one model, every routing change becomes a re-integration project.

That's why the Model Context Protocol (MCP) matters (Part 3 covered it). An MCP server exposes your system's capabilities as tools any MCP-capable agent can call. Swap the model, swap the assistant, and the connection still works. Treat models as replaceable engines and your integrations as the permanent road.

Key Takeaways

  • Different tasks, different models: speed, depth, cost and data location rarely line up behind one model.
  • Three camps: developer picks (Cursor, Copilot, Claude Code), tool routes (Bob SaaS), or one pinned self-hosted model (open models on your own infrastructure).
  • Auto-routing is a double-edged sword: less choice paralysis, less transparency. Ask vendors about per-model visibility, pinning and spend breakdowns.
  • Pick by team, not by company: regulated code, broad teams and power users have different needs.
  • Keep integrations model-agnostic with MCP so the next routing change costs you nothing.

References

🧰 From TheMaximoGuys toolbox: If Maximo is part of your world, Max_mcp is our Maximo MCP server (175 tools across 20 modules, available on npm and GitHub). It doesn't care which model is behind the wheel: plug it into Claude, Bob, Copilot, Cursor, or any MCP-capable agent.

Series Navigation

Previous:Part 7 — Vibe Coding Behind the Firewall
Next:Part 9 — Receipts, Please

Published by TheMaximoGuys | September 2026