AI-Augmented Inspection: Visual Inspection & Large Vision Models
🎯 Who this is for: Bridge program managers evaluating drone-and-AI inspection, and IT/deployment leads — especially at FedRAMP-bound agencies — who need to know what the AI actually does, what hardware it needs, and where the watsonx.ai dependency draws a hard line through the feature set.
Series: Part 4 of 5 — Maximo Civil Infrastructure on MAS 9 | Read time: 16 minutes
🤖 Where the Modern Story and the Honest Caveat Meet
Drones and computer vision are genuinely transforming bridge inspection. A drone can scan a deck or the underside of a span in an hour that would take an inspection crew a shift with a snooper truck, and an AI model can flag every crack and spall with a consistency no tired human matches at hour six. This is real, and Civil Infrastructure is built to use it.
But this is also the part where a vendor demo most often oversells, because the most impressive AI capability — the Large Vision Models — carries a dependency that determines whether a given agency can use it at all. So this part does two things at once: it explains what the AI does and how, and it draws the deployment line honestly, because the primary audience for this application is government, and government lives inside FedRAMP.
💡 Key insight: Hold two facts side by side for the rest of this part: CNN-based Visual Inspection works wherever you have GPUs. The Large Vision Models for civil infrastructure require watsonx.ai, which is FedRAMP-authorized only as a SaaS on AWS GovCloud, not on IBM Cloud. Everything else is detail. If you conflate the two — promising a federal customer the LVM experience on a cluster that cannot run watsonx.ai — you will be wrong in the worst possible setting, a signed statement of work.
🔍 How AI Augments the Inspection
Civil Infrastructure integrates with Maximo Visual Inspection (MVI) for AI-augmented inspections. The flow is straightforward:
- Drone-captured imagery of bridge decks, piers, and beams is collected in the field.
- MVI models analyze the imagery and automatically detect defects — spalling, delamination, cracking, efflorescence, section loss.
- The visual evidence links to element-level inspection records — a detected spall on the deck attaches to the deck element, a corroded region on a girder to that girder.
The critical design point is that the AI output does not float free; it binds to the element model from Part 2. A detected defect becomes evidence on a specific element, which means it can inform that element's condition-state quantities, seed a deficiency, and end up traceable through the same deficiency-to-work-order loop as a human finding. The AI does not create a parallel inspection universe — it feeds the one authoritative record.
Augmentation, not replacement
This must be said plainly because the regulatory context demands it: AI augments the qualified inspector; it does not replace them. Federal bridge inspection requires a qualified inspector, and the compliance record rests on their judgment. The AI's role is to extend coverage (scan more, more consistently), pre-screen (flag candidate defects for the inspector to confirm), and document (attach objective visual evidence). The inspector reviews, confirms or rejects, and owns the record. A program that treats an AI detection as an unreviewed inspection finding has misunderstood both the technology and the regulation.
💡 Key insight: The right mental model is "AI as a tireless first-pass inspector whose findings a qualified human adjudicates." That framing is both technically accurate and regulatorily safe. It also sets up the honest expectation for accuracy: the AI will produce false positives (flagging a stain as a crack) and occasionally false negatives, and the human-in-the-loop is not a compliance nicety — it is how those errors get caught before they reach the record.
🧠 CNN Versus Large Vision Models
MVI offers two generations of model technology, and the difference is the crux of this part.
Classic CNN models
The established MVI approach uses convolutional neural networks (CNNs) — deep-learning architectures optimized for image analysis. In practice you use transfer learning: a pre-trained backbone (ResNet, EfficientNet) fine-tuned on your labeled images, which reduces how much data you need but does not eliminate it. You still collect and label a meaningful volume of images per defect class.
The documented minimum training-data guidance makes the cost concrete:
| Model type | Minimum images | Recommended | Notes |
|---|---|---|---|
| Classification | 50 per class | 200+ per class | Balance classes evenly |
| Object detection | 100 with annotations | 500+ with annotations | Cover diverse angles and lighting |
| Anomaly detection | 100 "good" images | 500+ "good" images | Only normal examples needed |
For a defect like deck spalling, that means gathering and annotating hundreds of images — real work, and the reason a CNN-based program has a data-collection phase measured in weeks.
Large Vision Models (MAS 9.1)
The Large Vision Models introduced in MAS 9.1 are foundation models specifically trained for civil infrastructure — bridges and roads — that require minimal training data. This is a meaningful shift: instead of building a per-defect dataset from scratch, you lean on a model already trained on civil-infrastructure imagery, dramatically shortening the path to useful detection.
The trade-off is the dependency. Large Vision Models are foundation models, and in the MAS architecture foundation models run on watsonx.ai. That single fact is what makes the LVM capability powerful and constrained.
| Dimension | CNN models | Large Vision Models (9.1) |
|---|---|---|
| Training data needed | Hundreds of labeled images per class | Minimal |
| Time to first useful model | Weeks (collect + label + train) | Much faster |
| Compute for inference | GPU (T4-class) | GPU + watsonx.ai foundation-model dependency |
| Works on any GPU cluster | Yes | No — requires watsonx.ai availability |
| FedRAMP availability | Wherever GPU compute is authorized | Gated by watsonx.ai FedRAMP status |
🚧 The watsonx.ai and FedRAMP Reality
Now the honest part, and the reason this series treats deployment as a first-class concern rather than a footnote.
The Large Vision Models depend on watsonx.ai, IBM's foundation-model platform. For the government agencies that make up the core Civil Infrastructure audience, watsonx.ai's availability is not a given:
- watsonx.ai is not available on IBM Classic infrastructure, where many federal Maximo 7.6 environments have historically lived.
- watsonx.ai is not FedRAMP-authorized on IBM Cloud.
- watsonx.ai is FedRAMP-authorized on AWS GovCloud only — IBM announced the authorization on April 1, 2026 for watsonx.ai deployed exclusively as SaaS on AWS GovCloud (U.S.); the path runs through the hyperscaler, not IBM Cloud.
The practical consequence: an agency can stand up Civil Infrastructure and run CNN-based Visual Inspection wherever it has GPU compute, but the Large Vision Models will not light up until watsonx.ai is available and authorized in that agency's boundary. For a FedRAMP-bound customer, that means confirming the LVM capability can use the AWS GovCloud-authorized watsonx.ai from within its boundary, and in the interim delivering AI inspection with CNN models — which are perfectly capable, just data-hungrier.
💡 Key insight: This is the single most important thing to get right in front of a government customer. The correct posture is: "AI-augmented inspection with CNN models works today wherever we can give MVI GPUs. The civil-infrastructure Large Vision Models — the ones that need almost no training data — depend on watsonx.ai, whose FedRAMP authorization covers only the AWS GovCloud SaaS, so we confirm that path fits your boundary before we commit to it." Saying that is credible. Promising the LVM experience on a FedRAMP boundary that cannot run watsonx.ai is how a deal turns into a dispute.
🖥️ The Hardware and Capture Requirements
AI inspection is not a software toggle; it is a compute and capture commitment. Two real line items:
GPU compute
Visual Inspection requires GPU for both training and inference, exposed to the MVI pods through the OpenShift GPU Operator:
| Component | Minimum | Recommended | Notes |
|---|---|---|---|
| Training GPU | NVIDIA T4 (16 GB) | NVIDIA A100 (40/80 GB) | More VRAM = larger batches = faster training |
| Inference GPU | NVIDIA T4 (16 GB) | NVIDIA T4 or A10 | One GPU can serve multiple models |
| Edge GPU | NVIDIA Jetson Nano | NVIDIA Jetson Xavier NX | For edge/field deployment |
For an agency procuring its own compute — common on IBM Classic or an on-prem OpenShift — GPU bare metal is a budget line, not an afterthought. On a hyperscaler, it is GPU instance types (for example, T4- or A100-backed instances).
Cameras and drones
Capture quality drives detection quality:
| Use case | Camera type | Resolution | Notes |
|---|---|---|---|
| Drone bridge inspection | Drone-mounted camera | 12 MP+ / 4K video | The workhorse for deck and span imagery |
| Close-up element inspection | Industrial camera | 5 MP+ still | For detailed element documentation |
| Mobile / opportunistic | Smartphone camera | 8 MP+ | Field capture without dedicated gear |
The drone is what makes the economics work for bridges — coverage and safety (no lane closures or under-bridge access for the first pass) — but it is also an operational program (pilots, flight authorizations, data handling), not just a purchase.
🔧 Worked Example: A Drone Deck Scan on Bridge 04512
Return to Bridge 04512 and add a drone-and-AI pass to the routine inspection from Part 2.
Step 1 — Capture. A drone flies a programmed pattern over the deck and around the piers and girders, collecting 12 MP imagery and 4K video. What took a snooper truck and lane closures is a one-hour flight.
Step 2 — Analysis. The imagery is processed by MVI. If the agency is running CNN models, they were trained on the agency's labeled spalling/cracking/efflorescence images (the weeks-long data phase). If the agency has watsonx.ai available and is running the Large Vision Models, minimal agency-specific training was needed. Either way, MVI returns detected defects with locations and confidence scores.
Step 3 — Evidence binds to elements. A cluster of detected spalls on the deck attaches as evidence to the deck element; girder corrosion attaches to the girder. The AI output is now sitting on the Part 2 element records, not in a separate report.
Step 4 — The inspector adjudicates. The qualified inspector reviews the detections: confirms the deck spalling (and uses the AI's quantification to help distribute the deck's condition-state quantities), rejects a false positive where the model flagged a water stain as a crack, and confirms the girder corrosion. The inspector's confirmed findings are the record.
Step 5 — Into the loop. The confirmed spalling becomes part of the deck's condition-state distribution and a deficiency; the girder corrosion updates that element and its deficiency. From here it is the same deficiency-to-work-order loop from Part 3 — now with objective visual evidence attached to each finding.
The AI did not replace the inspection; it made the inspection faster, more consistent, better documented, and — for the LVM case — far cheaper to stand up. And every bit of it landed on the element model that the federal submission and the work-order loop already depend on.
⚠️ Edge Cases and Gotchas
False positives and confidence thresholds. AI detection produces false positives; stains, shadows, and old repairs get flagged. Set confidence thresholds deliberately and keep the human-in-the-loop — an unreviewed detection is not an inspection finding.
Model drift. A model's accuracy in production can drift as conditions change (new deck surfaces, seasonal lighting). Monitor model accuracy over time (an MVI analytics capability) and retrain when it degrades, rather than trusting a model indefinitely.
Assuming LVM availability. The recurring trap: scoping the Large Vision Models for a customer whose boundary cannot run watsonx.ai. Confirm watsonx.ai availability in the target boundary before committing to LVM, and default to CNN models where it is absent.
Edge versus central inference. Field/edge inference (Jetson-class) trades accuracy and model size for connectivity independence. Decide where inference runs based on field connectivity, not by default.
🩺 Troubleshooting AI Inspection
- If MVI will not train or serve models, it means GPU is not available or the GPU Operator is not installed, so verify GPU nodes and the OpenShift GPU Operator before anything else.
- If the Large Vision Models are unavailable, it means watsonx.ai is not present or not authorized in the boundary, so confirm the watsonx.ai/FedRAMP status and fall back to CNN models in the interim.
- If detection accuracy is poor, it means training data is too thin or unrepresentative (for CNN), or capture quality is low, so expand and balance the dataset or improve camera/drone imagery.
- If AI findings are not appearing on inspection records, it means the MVI-to-element evidence link is not configured, so verify the integration binds detections to element records.
- If false positives are reaching the record, it means the human review step is being skipped or thresholds are too low, so enforce inspector adjudication and tune confidence thresholds.
📋 Practical Notes for Rollout
- Confirm watsonx.ai availability in the target boundary first, and set customer expectations accordingly — CNN today, LVM once the AWS GovCloud-authorized watsonx.ai is confirmed reachable for FedRAMP customers.
- Budget GPU compute as a real line item (bare metal or GPU instances) plus the OpenShift GPU Operator.
- Stand up the drone program as an operational capability, not just a hardware purchase — pilots, authorizations, and data handling.
- Design the human-in-the-loop from day one: AI pre-screens, the qualified inspector adjudicates, the inspector owns the record.
- Monitor model accuracy and plan retraining so detection quality holds up in production rather than silently drifting.
Key Takeaways
- Civil Infrastructure integrates with Visual Inspection to analyze drone imagery and auto-detect deck, pier, and beam defects, binding the visual evidence to element-level inspection records.
- Classic CNN models work wherever GPU compute exists but need hundreds of labeled images per class; the MAS 9.1 Large Vision Models for civil infrastructure need minimal training data but depend on watsonx.ai.
- The watsonx.ai dependency is a real FedRAMP constraint — LVM is not automatically available to government agencies and should be planned around the AWS GovCloud-only authorization, with CNN models delivering AI inspection in the interim.
- AI augments the qualified inspector, it does not replace them — a human-in-the-loop adjudicates detections, and the inspector's judgment remains the record of authority.
- GPU compute and drone cameras are real infrastructure commitments, not free features, and belong in the budget and the rollout plan.
References
- IBM Maximo Visual Inspection overview (IBM Documentation)
- IBM watsonx.ai overview (IBM Documentation)
- NVIDIA GPU Operator for OpenShift (Red Hat)
- FHWA — Uncrewed Aircraft Systems (drones) for bridge inspection (FHWA)
- Maximo Application Suite — Manage add-ons and industry solutions (IBM Documentation)
Series Navigation
| Previous: | Part 3 — Pavement, Tunnels & the Deficiency-to-Work-Order Loop |
|---|---|
| Next: | Part 5 — Compliance, Reporting & Rollout |
About TheMaximoGuys: We help Maximo developers and teams navigate the move to MAS 9 with practical, no-hype guidance grounded in how the platform actually behaves.
Published by TheMaximoGuys | July 2026




