MAS 9 + IBM watsonx.data: Building the Maximo Open Lakehouse

🎯 Who this is for: Data engineers and architects scoping an IBM-native lakehouse, reliability and maintenance leaders who have hit the ceiling of Cognos and Maximo Health, IT and platform owners standardized on IBM who want a lakehouse decision that doesn't add a second vendor, and any Maximo practitioner who has heard "IBM has its own answer to Databricks" and wants to know exactly what that means for their WORKORDER and ASSET tables.

Series: Series Index β€” MAS 9 + IBM watsonx.data Lakehouse | 6 parts | Total read time: ~1.75 hours

πŸ“– Why This Series Exists

Every organization that upgrades from Maximo 7.6 to MAS 9 inherits a genuinely better analytics story than it had before. Operational Dashboards replace Start Centers. Cognos Analytics replaces BIRT. Maximo Health scores assets. Maximo Predict trains failure models. Maximo Monitor ingests IoT streams and detects anomalies. The AI Service, backed by watsonx, answers natural-language questions about your work orders. None of that is marketing. It is real, entitled, and worth using first.

And then, six months into running MAS 9 in production, the same questions start showing up in planning meetings. Can we join maintenance cost against SAP financials? Can more than three people build a Cognos report without a separate license? Can we train a custom failure model on our own vibration data instead of Predict's fixed catalog? Can we keep ten years of sensor history queryable instead of aging out of Monitor's retention window? The honest answer to all four, inside MAS 9 alone, is no. Our MAS-DATABRICKS series already answered that gap with Databricks, the platform most MAS 9 customers are already evaluating. This series answers the identical gap with IBM's own product: watsonx.data, an open, hybrid data lakehouse and the data pillar of the broader watsonx platform.

This series is not an anti-Databricks piece, and it is not IBM sales copy dressed up as Maximo content either. It draws the same line the Databricks series draws, from the other side of the fence: MAS 9 native analytics is real and should be your first stop. A lakehouse β€” whichever vendor's β€” earns its cost only past a specific, nameable ceiling. What differs here is the platform past that ceiling, and the case for it: if your shop is already IBM-standardized β€” Maximo, Cognos, watsonx.ai through the MAS AI Service, Red Hat OpenShift β€” watsonx.data reuses that footprint instead of introducing a second vendor relationship, a second support path, and a second runtime to govern.

πŸ’‘ Key insight: The organizing question of this series is the same one the Databricks series asks, answered from the IBM side: "which of my analytics questions does MAS 9 already answer, and which ones require data MAS 9 was never built to hold β€” and given that I need to go past MAS 9, does staying inside the IBM stack change the answer?" For an IBM-standardized shop, it usually does.

🧭 Where This Series Fits

This series is deliberately about the boundary between MAS 9's native analytics stack and an IBM-native lakehouse past it. It does not re-teach Maximo Health, Predict, or Monitor from scratch, and it does not re-litigate the Databricks path our companion series already covers in depth. There is a division of labor across TheMaximoGuys content so you are never reading the same material twice:

If you want…Read…
Deep, hands-on coverage of Maximo Health scoring, risk matrices, and degradation curvesMAS-HEALTH series
Deep, hands-on coverage of Maximo Predict's ML model catalog and prediction workflowsMAS-PREDICT series
Deep, hands-on coverage of Maximo Monitor's IoT ingestion and anomaly detectionMAS-MONITOR series
The non-IBM lakehouse path for the identical MAS 9 analytics ceiling β€” Databricks, Delta Lake, Unity CatalogMAS-DATABRICKS series
Where native MAS 9 analytics stops, and how to extend it with IBM's own watsonx.data lakehouseThis series
πŸ’‘ Key insight: If a stakeholder is mid-evaluation between watsonx.data and Databricks, hand them both series' Part 1 back to back β€” they open from the same MAS 9 ceiling table and diverge only at the platform choice. That symmetry is deliberate; it makes the two series a fair side-by-side, not a race to convince you of one answer.

πŸ“Š The Series at a Glance

PartTitleFocus AreaRead Time
1Why watsonx.data: IBM's Open Lakehouse Answer for MaximoNative ceiling recap; Apache Iceberg vs Delta Lake; one-vendor governance case; watsonx.ai already bridged via MAS AI Service16 min
2Getting Maximo Data into watsonx.dataMIF REST/JSON (OSLC, synonym-domain values), Kafka streaming, MAS 9.1 bulk export; Cloud Pak for Data as the integration fabric17 min
3The Iceberg MedallionBronze/silver/gold mapped to WORKORDER, ASSET, FAILUREREPORT, MEASUREMENT; ACID, schema evolution, time travel, Db2-Iceberg interop18 min
4Fit-for-Purpose EnginesPresto, Presto C++/Velox, Spark, Db2 Warehouse, Netezza over one Iceberg copy; when to route which EAM workload where; RU pricing16 min
5From Lakehouse to Actionwatsonx.ai custom ML, AutoAI for RAG, Milvus/OpenSearch/OpenRAG over maintenance manuals and WO text; three closed-loop write-back patterns18 min
6watsonx.data vs. Databricks vs. MAS NativeIceberg vs Delta, Presto C++ vs Photon, IKC vs Unity Catalog; decision framework; governance and PII masking (series finale)17 min

πŸ”‘ What Makes This Series Different

Three ideas run through every part. Internalize them and the series reads as one argument instead of six standalone articles.

Native-first, lakehouse-second is not a hedge β€” it is the architecture, regardless of vendor. Every part follows the same sequence: name what MAS 9 already does natively, name precisely where it stops, then show the watsonx.data pattern that picks up from there. A shop that skips the "what MAS already does" step ends up rebuilding Cognos-shaped dashboards in Presto SQL for no reason β€” the single most common wasted-spend pattern in lakehouse projects touching an EAM system, whichever lakehouse vendor you picked.

Open by design is the specific IBM argument, and this series names it precisely instead of gesturing at it. watsonx.data standardizes on Apache Iceberg β€” an open table format readable by Presto, Spark, Db2, Netezza, and even Databricks via Unity Catalog federation β€” rather than pairing a proprietary execution engine with a single vendor's table format. IBM's own published benchmarks claim better price-performance for Presto C++/Velox against Databricks' Photon engine at equal query runtime on a 100 TB TPC-DS workload. Those are IBM-internal benchmarks, not independent third-party results, and this series repeats that caveat every time the numbers come up β€” a fair comparison belongs in your own proof-of-concept, not an assumption this series makes for you.

Closed-loop is still the pattern that actually pays back, on either lakehouse. A watsonx.data build that only produces dashboards nobody acts on is exactly as much of a cost center as an idle Databricks build. The pattern every proven case study in this space follows is closed-loop: predictions and scores write back into Maximo β€” as a work order, a service request, or an updated reorder point β€” through the MAS AI Service, direct REST calls, or watsonx Orchestrate, so the lakehouse output becomes an action inside the system your technicians already work in, not a report that sits in a folder.

πŸ’‘ Key insight: If you remember only one thing from the index, make it this: name the ceiling, extend past it with the platform that fits your vendor posture, close the loop back into Maximo. Every part is an application of that triad. Skip the last step and you get a beautiful Iceberg lakehouse nobody's Maximo workflow ever touches.

πŸ—ΊοΈ What MAS 9 Already Gives You (and Where It Stops)

Before the part-by-part guide, this table grounds the whole series in what native MAS 9 analytics actually covers today β€” the identical baseline the MAS-DATABRICKS series measures against, because the native ceiling does not change with your lakehouse choice.

MAS 9 Native CapabilityWhat It Does WellWhere It Stops
Operational DashboardsRole-based, cross-app views pulling from Monitor, Health, and Manage in one screenNot a BI platform β€” no ad-hoc joins, no external data sources, no complex statistical analysis
Cognos AnalyticsDrag-and-drop report authoring against the Maximo database, scheduled distributionOnly 3 administrative authoring seats included; BIRT reports must be rebuilt, not migrated; hard to join non-Maximo data
Maximo HealthConfigurable 0–100 health/risk scores from asset, work order, and meter data β€” as of MAS 9.1, delivered as "Maximo Manage with Health"Cannot incorporate external factors β€” weather, SCADA load, supply-chain risk β€” into the score formula
Maximo PredictPre-built failure-probability, RUL, and anomaly models trained via watsonx.ai on Cloud Pak for DataFixed model catalog with limited customization; requires IoT data already flowing through Monitor; cannot ingest external data sources
Maximo MonitorReal-time IoT dashboards, built-in anomaly detection, alert-to-work-order automation via KafkaNot built for multi-year historical trend analysis or custom ML algorithms beyond its built-in anomaly functions
Maximo AI Service / AssistantNatural-language Q&A against Maximo data, FMEA content generation, similarity tracking β€” and the sanctioned bridge into watsonx.aiScoped to Maximo data only; ~10-AppPoints entitlement with a monthly usage ceiling before additional watsonx licensing is required

That right-hand column is not a criticism of MAS 9 β€” every EAM system on the market has an equivalent list. It is the honest starting point for deciding whether your organization actually needs a lakehouse at all, or whether the gap you are feeling is really an under-used native capability. Part 1 opens with exactly that diagnostic, the same way Part 1 of the Databricks series does.

βš–οΈ watsonx.data vs. Databricks at a Glance

Because this series exists specifically as the IBM-native counterpart to MAS-DATABRICKS, the index earns a direct comparison up front rather than deferring it entirely to Part 6.

Dimensionwatsonx.dataDatabricks
Primary table formatApache Iceberg (open, multi-engine)Delta Lake (open format, runtime-coupled)
Execution enginePresto (Java) + Presto C++/VeloxPhoton (proprietary vectorized)
GovernanceIBM Knowledge Catalog (now watsonx.data intelligence) + semantic layer + Data Product Hub, or Apache RangerUnity Catalog (Databricks-centric)
RAG / vectorEmbedded Milvus + OpenSearch + OpenRAG (Docling/Langflow)Own runtime + partner vector integrations
AI layerwatsonx.ai β€” already bridged via the MAS AI ServiceMosaic AI / partner models
Vendor postureSingle IBM support and licensing path (Maximo, Cognos, watsonx, OpenShift)Separate vendor relationship alongside IBM
IBM's own price/perf claim"less than 60% the cost" of Photon at equal runtime (IBM-internal, 100 TB TPC-DS)β€”

Both platforms federate to the other: watsonx.data can zero-copy federate to Databricks Unity Catalog, and Presto can query Uniform-enabled Delta tables through the Iceberg REST Catalog API. A hybrid "Delta in Databricks, Iceberg in watsonx.data" architecture is technically feasible for organizations mid-migration between the two β€” but it requires deliberately reconciling governance across both catalogs; it does not happen automatically. Part 6 expands this table into a full decision framework, including when the honest answer is "stay on MAS native a while longer."

πŸ”§ One Concrete Example: Iceberg in Practice

The series stays grounded in real syntax, not slideware, starting with the index. Here is how Db2 Warehouse β€” one of watsonx.data's fit-for-purpose engines β€” creates a native Iceberg table directly inside the watsonx.data catalog, the kind of statement Part 3 and Part 4 build on repeatedly:

CREATE DATALAKE TABLE iceberg.db2exported (
    WORKORDERID   BIGINT,
    WONUM         VARCHAR(20),
    SITEID        VARCHAR(20),
    ORGID         VARCHAR(20),
    ASSETNUM      VARCHAR(20),
    LOCATION      VARCHAR(20),
    STATUS        VARCHAR(20),
    STATUSDATE    TIMESTAMP,
    FAILURECODE   VARCHAR(20),
    PROBLEMCODE   VARCHAR(20),
    LABORCOST     DECIMAL(15,2),
    MATERIALCOST  DECIMAL(15,2),
    TOOLCOST      DECIMAL(15,2)
)
STORED AS PARQUET
STORED BY ICEBERG
LOCATION 'DB2REMOTE://iceberg-bucket//iceberg/db2exported';

That single STORED BY ICEBERG clause is the practical difference this series keeps coming back to: the table Db2 just created is immediately queryable by Presto, Spark, and Netezza without a copy, an export job, or a second pipeline β€” because Iceberg, not any one engine, owns the table's metadata contract. Note the iceberg.catalog property requirement and that direct ALTER from Db2 is restricted, reflecting Iceberg's metadata-driven schema evolution model β€” a detail Part 3 covers with the full WORKORDER-to-ASSET enrichment example.

🧩 Part-by-Part Guide

Part 1: Why watsonx.data β€” IBM's Open Lakehouse Answer for Maximo

[Read Part 1 β€” Why watsonx.data](/blog/mas-watsonx-data-01-why-open-lakehouse) Β· 16 minutes

The IBM-native case, built the same way Part 1 of the Databricks series is: name the native ceiling, then show why an open Iceberg lakehouse with fit-for-purpose engines and one-vendor governance closes it for an IBM-standardized shop.

You learn: the same "do we actually need this" diagnostic, applied before an IBM purchase instead of a Databricks one; how Apache Iceberg's open format differs architecturally from Delta Lake's runtime coupling to Photon; how the MAS AI Service already bridges Maximo into watsonx.ai, and why that matters for adoption cost; and where this series explicitly hands off to MAS-DATABRICKS for shops that are not IBM-standardized.

Part 2: Getting Maximo Data into watsonx.data

[Read Part 2 β€” Getting Maximo Data In](/blog/mas-watsonx-data-02-data-extraction) Β· 17 minutes

The extraction architecture question, answered for watsonx.data specifically: MIF REST/JSON over OSLC, the Kafka source connector, and MAS 9.1's asynchronous bulk export β€” plus why Cloud Pak for Data is the fabric that makes any of them practical at scale.

You learn: how MIF's synonym-domain internal values matter for ML feature stability; the exact Cloud Pak for Data Presto connection parameters (port 443, engine port 8443); why a standalone Manage install only needs the Db2 Warehouse operator while cross-suite watsonx integration needs full CPD; and the named bronze-layer WORKORDER columns and suggested extraction frequencies this series works from throughout.

Part 3: The Iceberg Medallion

[Read Part 3 β€” The Iceberg Medallion](/blog/mas-watsonx-data-03-iceberg-medallion) Β· 18 minutes

Bronze, silver, gold β€” mapped table by table onto real Maximo objects on Apache Iceberg, with the ACID, schema-evolution, and time-travel mechanics that make the gold layer trustworthy enough for an executive dashboard.

You learn: the full ASSET_DIM, WORKORDER_FACT, FAILURE_FACT, MEASUREMENT_FACT silver-layer conformed entities and their synonym-decoded standardized keys; how Iceberg's schema evolution absorbs a Maximo attribute change without breaking downstream Spark jobs; the bucket-level-versus-table-level registration distinction between Iceberg and Delta/Hudi tables; and how the Sync metadata feature's three modes keep external changes reconciled.

Part 4: Fit-for-Purpose Engines

[Read Part 4 β€” Fit-for-Purpose Engines](/blog/mas-watsonx-data-04-fit-for-purpose-engines) Β· 16 minutes

IBM's core design premise β€” no single query engine is optimal for every workload β€” worked through for EAM specifically: Presto for interactive analyst queries and federation, Presto C++/Velox for cost-sensitive high-performance SQL, Spark for ETL/ML/table maintenance, and Db2 Warehouse/Netezza for high-concurrency BI.

You learn: the Presto coordinator/worker/resource-manager architecture and when federation beats a full extract; what Presto C++ v0.286 and the IBM query optimizer actually change versus classic Presto; which EAM workloads (silver/gold transformations, feature engineering, reliability KPIs) route to which engine and why; and how Resource Unit (RU) metering β€” USD 1/RU, per-second, one-minute minimum β€” turns engine choice into a real cost lever.

Part 5: From Lakehouse to Action

[Read Part 5 β€” From Lakehouse to Action](/blog/mas-watsonx-data-05-ai-rag-closed-loop) Β· 18 minutes

Where the Iceberg lakehouse stops being a reporting layer and starts producing decisions: watsonx.ai custom ML over gold-layer features, AutoAI for RAG, the Milvus/OpenSearch/OpenRAG stack for maintenance-manual and work-order-text retrieval, and three concrete closed-loop write-back patterns into Maximo.

You learn: how to train and serve a custom PdM or RUL model on watsonx.ai where Maximo Predict's fixed catalog stops; how Docling converts OEM manuals and SOPs into an AI-ready RAG corpus for a troubleshooting agent; the three write-back methods β€” AI Service-mediated, direct REST (POST /os/mxwo, PUT /os/mxitem), and watsonx Orchestrate via the IBM/maximo-wxo-integration reference repo; and why "use both Predict and custom watsonx.ai models" is usually the right answer, not a replacement decision.

Part 6: watsonx.data vs. Databricks vs. MAS Native β€” Choosing the IBM-Native Path (and Governing It)

[Read Part 6 β€” Choosing and Governing the IBM-Native Path](/blog/mas-watsonx-data-06-vs-databricks-governance) Β· 17 minutes (Series Finale)

The decision framework this whole series has been building toward, paired with the governance layer that turns any Iceberg lakehouse from a data-sprawl liability into an auditable enterprise asset.

You learn: the full, caveated head-to-head β€” Iceberg vs. Delta, Presto C++/Velox vs. Photon, IBM Knowledge Catalog vs. Unity Catalog, Milvus/OpenRAG vs. Databricks' vector integrations β€” with IBM's internal benchmarks clearly labeled as such; a scored decision matrix for choosing watsonx.data, Databricks, or staying MAS-native a while longer; how IBM Knowledge Catalog's row/column policies map onto Maximo's site/org security groups; and where labor transaction data (LABTRANS) carries PII that needs IKC masking before an auditor ever asks.

🧭 Recommended Reading Paths

Data Engineer / Architect

"I own the pipeline and the platform."
Read Part 2 β†’ Part 3 β†’ Part 4. Start with extraction patterns, then the Iceberg medallion mapping, then the engine strategy β€” the three parts that determine whether the platform is buildable and cost-controlled.

Reliability / Maintenance Leader

"I own uptime, cost, and the maintenance budget."
Read Part 1 β†’ Part 5 β†’ Part 6. Start with the honest case for adding a lakehouse, then the concrete AI/RAG use cases, then the decision framework β€” skip the engine plumbing unless you need it.

IT / Platform Owner

"I own the license decision and the vendor roadmap."
Read Part 1 β†’ Part 6 β†’ Part 4. Start with the ceiling MAS 9 native analytics actually has, then the watsonx.data-versus-Databricks-versus-native decision in Part 6, then the RU cost model in Part 4.

Data Scientist / ML Engineer

"I own the models."
Read Part 3 β†’ Part 5 β†’ Part 4. Start with what the gold layer actually contains, then watsonx.ai model training and RAG, then the engine choices that determine how fast your features actually materialize.

πŸ’‘ Key Themes Across the Series

MAS 9 native analytics is genuinely good β€” and genuinely bounded, regardless of which lakehouse you add. Cognos, Health, Predict, Monitor, and the AI Service are real capabilities worth using fully before any lakehouse conversation starts. Every part in this series names their specific ceiling instead of hand-waving past them, because a watsonx.data build that duplicates entitled MAS capability overspends exactly the way a redundant Databricks build does.

Open table format is the specific IBM argument β€” and this series shows the syntax, not just the slogan. Apache Iceberg's ACID transactions, schema evolution, and time travel are demonstrated with real Db2, Presto, and Spark statements throughout Parts 2 through 4, not asserted as a feature-list bullet.

Extraction method depends on your deployment model, not preference. MIF REST/JSON is the default; Kafka suits near-real-time needs; bulk export handles large periodic pulls; direct DB extraction only exists for self-managed MAS because SaaS MAS seals its database β€” identical constraints to the Databricks series, because the constraint comes from MAS, not the lakehouse vendor.

Closed-loop is what separates a working lakehouse from an expensive reporting layer, on either platform. Every AI use case in Part 5 should end with a REST or Orchestrate call back into Maximo β€” a work order, a service request, an updated reorder point β€” not just a chart nobody acts on.

IBM's own benchmark claims are labeled as IBM-internal every time this series cites them. The Photon price/performance comparison, the "up to 50%" data-warehouse cost reduction, and the "less than 60% the cost" claim are all IBM's own published figures on IBM-chosen workloads (100 TB TPC-DS). This series treats them as directional inputs to your own proof-of-concept, never as independently verified fact.

🧰 How to Use This Series

  • Scoping a lakehouse business case? Use Part 1's diagnostic to separate "we need a lakehouse" from "we're under-using MAS Health and Predict" before either an IBM or a Databricks dollar is committed.
  • Already evaluating Databricks and want the IBM-native alternative on the table? Read this index alongside the MAS-DATABRICKS index β€” both open from the identical MAS 9 ceiling table, so the platform comparison starts from the same baseline.
  • Building the actual pipeline? Parts 2 through 4 give you the extraction method, the Iceberg medallion structure, and the engine strategy in enough table-level detail to start a Phase 1 build.
  • Justifying the spend to a maintenance VP? Part 5's closed-loop use cases connect directly to the reliability and cost metrics that VP already tracks in MAS-RELIABILITY and MAS-HEALTH reporting.
  • Preparing for a data governance or compliance review? Part 6's IBM Knowledge Catalog mapping is written to survive an auditor's question about Maximo-derived data the same way the MAS-DATABRICKS series' Unity Catalog mapping does for that platform.

References

Start the Series

Begin with [Part 1 β€” Why watsonx.data: IBM's Open Lakehouse Answer for Maximo](/blog/mas-watsonx-data-01-why-open-lakehouse), or jump to the part your role needs using the reading paths above. Evaluating the non-IBM path in parallel? Start the companion [MAS-DATABRICKS series](/blog/mas-databricks-series-index) at its own Part 1 for the identical ceiling, answered with Databricks instead.

About TheMaximoGuys: We help Maximo developers and teams navigate the move to MAS 9 with practical, no-hype guidance grounded in how the platform actually behaves.

Published by TheMaximoGuys | July 2026