Product
watsonx.data
watsonx.data is IBM's AI-ready data lakehouse platform. It enables you to manage, query, and analyze data across different environments (cloud, on-premises, hybrid) from a single point. Its goal is to eliminate data silos and enable enterprises to work with large data volumes at high performance.
Overview
watsonx.data is IBM's AI-ready data lakehouse platform. It enables you to manage, query, and analyze data across different environments (cloud, on-premises, hybrid) from a single point. Its goal is to eliminate data silos and enable enterprises to work with large data volumes at high performance.

The data layer of the IBM watsonx platform.
Built on open-source technologies (Apache Iceberg, Presto, etc.).
Offers independent scaling of storage and query engines.
Supports both structured and unstructured data.
Unifies data warehouse, data lake, and lakehouse architectures in one platform.
Benefits
Reduce unnecessary costs by scaling storage and compute separately
Fast analytics on large datasets with distributed query engines
Centralized security, compliance, and data quality policies
Run in multi-cloud and hybrid environments
Seamless integration with watsonx.ai and watsonx.governance
Features
Data lake flexibility + data warehouse performance
Support for Apache Iceberg, Hive, etc
Choose between Presto, Spark, Db2, Netezza engines
Compression, indexing, partition management
Role-based access control, encryption, audit logs
Technical Specifications
Parquet, Avro, ORC, CSV, JSON
Presto (stalo), Apache Spark, Netezza Performance Server
AWS S3, Azure Blob, Google Cloud Storage, MinIO, IBM Cloud Object Storage
Red Hat OpenShift, bare metal, virtual machines, all major clouds
REST API, JDBC/ODBC, Python SDK
Horizontal and vertical scaling, node‑based capacity increase