Picture a mid-sized utility, six months after rolling out an AI coding assistant. The internal audit team is doing its annual review of change management on the outage-reporting system. An auditor points at a block of code and asks a simple question:

"Who wrote this?"

The developer pauses. "Well... I asked the AI, it wrote most of it, I reviewed it. I think."

That "I think" is the problem. In a bank, an insurer, a government agency or a utility, I think doesn't pass audit. Vibe coding is fun right up until someone asks for the receipts.

🧾 The Four Questions Every Auditor Asks

Auditors have asked the same four questions about every change for decades. AI doesn't change the questions. It just makes them harder to answer.

  1. Who initiated the change?
  2. What exactly changed, in the code and in any connected systems?
  3. When did it happen?
  4. Who approved it before it took effect?

With a human developer, the answers come from source control, the ticket and the pull request. With an AI agent that can edit files, run commands and call external systems, there are suddenly many more actions to account for, and some of them never touch a commit at all. An agent that queries a production API to "check something" leaves no trace in git.

So the bar for AI-assisted work is: every agent action should be attributable to a human, reviewable after the fact, and reversible where possible.

📍 Where the Receipts Live

No single tool gives you the whole picture. Receipts come in layers.

Auditor's questionWhere the receipt lives
Who started it, and when?The AI tool's activity logs (e.g., CADF-format logs in IBM Bob), forwarded to your SIEM
What did the agent actually do?Per-action logs, plus the tool's approval and rollback history
What changed in the code?Git history and pull request diffs, same as always
Who approved it?In-tool approvals for risky actions, plus the PR reviewer
How much AI is in our codebase, and what did it cost?Usage analytics (e.g., IBM Bobalytics adoption, "Bob factor" and spend)

IBM Bob is a useful example of how vendors are filling these gaps, because IBM has leaned hard into the governance story. Its Enterprise tier offers activity logs in CADF (Cloud Auditing Data Federation, an open standard for audit events). An August 2026 update added Splunk SIEM forwarding, so those events land where your security team already looks, plus admin group policies pushed through MDM or GPO. Bob V2 added rollback per tool call, and while reads are auto-approved, edits, commands, MCP calls and skills still need a human click.

Other tools offer pieces of this in their own way, through hooks, enterprise admin consoles and audit logs. When you evaluate any of them, don't ask "do you have logging?" Ask "can I get every agent action, attributed to a named user, into our SIEM?"

⚠️ Read the dashboard carefully. Bobalytics reports a "Bob factor": the percentage of committed lines Bob created. It's a fine adoption metric. It is not a quality metric. A high Bob factor could mean your team is shipping faster, or that it's accepting a lot of verbose code without trimming it. Lines of code were a bad productivity measure in 1995, and AI didn't fix that.

🚪 APIs, Not the Database

Here is the receipt most teams forget, and it's architectural rather than a feature.

When an agent needs to do something in an enterprise system, such as creating a work order, updating a customer record or adjusting inventory, there are two ways it can get there:

  • The shortcut: connect straight to the database and write rows.
  • The front door: call the system's APIs, which run the same business rules, validation and security checks a human user would hit.

The shortcut is tempting, because it's fast and the agent is perfectly capable of writing SQL. It is also a governance disaster. Direct writes skip validation, skip the application's own change history, and run under whatever powerful service account holds the database credentials. Your receipt for that change is... a row that appeared.

The consensus among practitioners in our own Maximo world is blunt: agents act through object structures, APIs and automation scripts under the application's normal security, never directly on the database. The same principle applies to your ERP, your core banking system and your claims platform. Give the agent its own named, least-privilege API identity, and the system's own audit trail becomes your receipt for free.

🔍 An Honest Caveat

Logging tells you what happened. It doesn't stop the wrong thing from happening.

IBM describes Bob's security layer in broad strokes (prompt normalization, sensitive-data scanning, secrets detection, real-time policy enforcement), but the mechanisms aren't publicly documented in detail. Practitioners have also noted that agent commands run as ordinary child processes, without an OS-level sandbox. And remember the January 2026 PromptArmor research (Part 5), where a beta build could be tricked into running malware once a command was auto-approved.

None of that is unique to one vendor. It's a reminder that receipts are your second line of defense. The first is limiting what the agent can touch in the first place.

✅ A Five-Minute Receipt Checklist

Before an AI coding agent gets anywhere near a regulated system:

  1. Every agent session is tied to a named human, not a shared account.
  2. Agent activity logs flow into your SIEM, with retention matching your audit policy.
  3. Writes, commands and external calls need approval; reads can be relaxed.
  4. Agents reach enterprise systems only through APIs, with least-privilege identities.
  5. AI-assisted pull requests get the same review as human ones. No "the AI wrote it, so it's fine" fast lane.

That last item is where the strain shows. In GitLab's 2026 survey (a figure IBM itself cites), 85% of DevSecOps professionals agreed AI has shifted the bottleneck from writing code to reviewing it. IBM's Adam McDaniel says it plainly: "Review capacity becomes the limiting factor." If your agents produce three times the pull requests and your reviewers stay the same, your approval receipts turn into rubber stamps. Budget reviewer time the way you budget licenses.

Key Takeaways

  • Auditors ask who, what, when and who approved, and AI-assisted work must answer all four.
  • Receipts live in layers: tool activity logs (such as Bob's CADF logs), SIEM, git, pull requests and usage analytics.
  • Agents should use the front door: APIs and business logic under normal security, never direct database writes.
  • Volume metrics aren't value metrics: treat the Bob factor and similar numbers as adoption signals, not proof of quality.

References

🧰 From TheMaximoGuys toolbox: Want your agents to use the front door in Maximo? Max_Interfaces is our open-source (MIT) Maximo API library: 2,439+ endpoints across OSLC and NextGen REST, as Postman and OpenAPI collections, covering Maximo 7.x, 8.x and MAS 9. Point Claude, Bob, Copilot, Cursor, or any agent at the APIs instead of the database.

Series Navigation

Previous:Part 8 — One Model Doesn't Fit All
Next:Part 10 — Reading the Productivity Numbers

Published by TheMaximoGuys | September 2026