Product

watsonx.data

watsonx.data is IBM's AI-ready data lakehouse platform. It enables you to manage, query, and analyze data across different environments (cloud, on-premises, hybrid) from a single point. Its goal is to eliminate data silos and enable enterprises to work with large data volumes at high performance.

Overview

watsonx.data is IBM's AI-ready data lakehouse platform. It enables you to manage, query, and analyze data across different environments (cloud, on-premises, hybrid) from a single point. Its goal is to eliminate data silos and enable enterprises to work with large data volumes at high performance.
watsonx.data
  • The data layer of the IBM watsonx platform.

  • Built on open-source technologies (Apache Iceberg, Presto, etc.).

  • Offers independent scaling of storage and query engines.

  • Supports both structured and unstructured data.

  • Unifies data warehouse, data lake, and lakehouse architectures in one platform.

 

Benefits

Reduce unnecessary costs by scaling storage and compute separately

Fast analytics on large datasets with distributed query engines

Centralized security, compliance, and data quality policies

Run in multi-cloud and hybrid environments

Seamless integration with watsonx.ai and watsonx.governance

Features

Data lake flexibility + data warehouse performance

Support for Apache Iceberg, Hive, etc

Choose between Presto, Spark, Db2, Netezza engines

Compression, indexing, partition management

Role-based access control, encryption, audit logs

Technical Specifications

Parquet, Avro, ORC, CSV, JSON

Presto (stalo), Apache Spark, Netezza Performance Server

AWS S3, Azure Blob, Google Cloud Storage, MinIO, IBM Cloud Object Storage

Red Hat OpenShift, bare metal, virtual machines, all major clouds

REST API, JDBC/ODBC, Python SDK

Horizontal and vertical scaling, node‑based capacity increase