Introducing Cloudflare Basin: an open, serverless data platform, now generally available

Marc Selwan and Micah Wylde

11 minute read

During Birthday Week 2025, we announced the Cloudflare Data Platform, a suite of products that ingest, store, and query your analytical data. Today, we’re announcing that the platform is generally available, and we’re giving it a new name: Cloudflare Basin.

Basin is a serverless data analytics platform built on Apache Iceberg, the open standard for data lakes, and R2 Object Storage. The Basin family includes:

  • Basin Pipelines, formerly Cloudflare Pipelines, receives events from Workers, HTTP, or Cloudflare Logpush, transforms them with SQL, and writes them as Apache Iceberg tables or files in R2.
  • Basin Catalog, formerly R2 Data Catalog, manages Iceberg metadata and automatically maintains tables to keep them fast and cost-efficient.
  • Basin SQL, formerly R2 SQL, is our serverless, distributed SQL engine for querying Apache Iceberg tables directly on Cloudflare.

Basin brings an end-to-end analytics platform to the Developer Platform, enabling you to collect data from a variety of sources, such as apps, infrastructure, devices, and other Cloudflare services, then query it to answer analytical questions.

We set out to build a data platform last year when we saw two fundamental developments that changed how modern data applications were being built. First, Apache Iceberg emerged as the standard open table format, making data portable across nearly every major query engine. Second, we started seeing developers bring their analytics data to R2 where the lack of egress charges made it practical and cost-efficient to actually access their data from different tools, teams, regions, and cloud providers.

So, when we launched Basin in open beta, developers — including our billing and infrastructure teams at Cloudflare — immediately started adopting these services for a variety of use cases including using real-time data to optimize e-commerce sites, long-term storage and reporting of billing metrics, and ingesting and querying telemetry from Cloudflare’s infrastructure to measure and improve utilization and efficiency.

“We moved our entire company's data pipeline to Basin Pipelines, Catalog, and SQL, replacing a complex AWS S3 and Athena setup with a cleaner, serverless architecture that reliably handles all of our event data. -Dax Raad, Co-Founder, Anomaly

Our early adopters taught us a great deal about what it means to run an analytics platform on the edge. We spent the past year improving Basin around three specific areas that hone in on what makes analytics in Developer Platform unique: speed, openness, and cost efficiency.

Basin is built for speed — whether it’s about getting started or executing large queries. You can create a Basin Catalog, set up a Pipeline to ingest data, and query it with Basin SQL in seconds. This matters as we see more data applications being built from prompts to coding agents, which would otherwise have to wait and poll for resources or data. As datasets grow, Basin Catalog compacts metadata and data files to reduce I/O and generates statistics for query planning. Basin SQL uses those statistics to split queries into smaller tasks and distribute them across Workers, keeping queries fast and consistent as datasets scale.

A large driving force in the development and adoption of Basin is the continued growth we are seeing in the open Apache Iceberg ecosystem. Developers have been rallying around the radical idea that you should own and be in control of your own data — separating the storage layer from the compute layer and allowing you to use the right query engine for the job. With Basin, you can read and write your data using any Iceberg-compatible engine, including PyIceberg, DuckDB, Snowflake, and Apache Spark. That kind of data portability is only possible with free egress, which allows developers to access their data in Cloudflare from the wide variety of tools in the ecosystem, regardless of region or cloud.

“Bobsled is a data product platform that the world's most advanced data teams use to build and distribute AI-ready data to partners, vendors and customers," said Julien Grobbelaar, Head of Platform at Bobsled. "Basin allows us to build data products that can be made accessible in any region of every major data and AI platform, all at production-grade reliability and a fraction of the cost thanks to zero egress fees.” 

In addition to free egress, our serverless architecture allows us to offer further cost efficiencies to our customers with usage-based pricing. You are only billed when Basin ingests, processes, or queries your data. Developers can build out analytics for hobby projects at little to no cost, while our pricing scales economically for larger enterprise use cases. There are no hourly charges or separate infrastructure costs to worry about.

If you are ready to get started, refer to the Basin tutorial for a step-by-step guide on how to use Basin Pipelines to deliver events to an Apache Iceberg table managed by Basin Catalog, and query them with Basin SQL. Read on to learn more about Basin and where we are going next.

Why Basin?

We launched these products last year as the Cloudflare Data Platform, which has served us well for the first year of availability. For our GA launch, we decided we needed a new name that encompasses our ambitions for the platform, links together all the products, and is a bit punchier.

A basin is where rivers from many sources come together to a single point. We felt that Basin perfectly captures how the platform is used: Pipelines brings data into Basin Catalog while Basin SQL makes it instantly queryable. A fun fact: roughly 20% of Earth’s land drains into endorheic basins, much like over 20% of the web sits behind Cloudflare’s network. 

Lake Vänern, Sweden - A basin created from tectonic activity 200-300 million years ago.

Today, Basin is made up of three products: Pipelines, Catalog, and SQL, covering ingestion, storage, and querying, and will expand over time with more products managing the rest of the analytical data lifecycle.

Basin Pipelines

Before you can query your data, your events need to be ingested, structured to a schema, and written to object storage. This is the role of Basin Pipelines. It accepts events through HTTP endpoints or Workers bindings, processes them according to a SQL query, and delivers them to Basin Catalog as Apache Iceberg tables or R2 as JSON or Parquet files.

Since our beta launch, users have created tens of thousands of Pipelines for a wide variety of use cases. For example, a common pattern we see is using Pipelines to transform Cloudflare HTTP logs before storing them:

INSERT INTO http_logs_sink
SELECT
  EdgeResponseStatus,
  to_timestamp_micros(EdgeStartTimestamp) AS event_time,
  upper(ClientRequestMethod) AS method,
  sha256(ClientIP) AS hashed_ip
FROM http_logs_stream
WHERE EdgeResponseStatus >= 400;

Doing this work during ingestion can significantly reduce the storage footprint, reduce noise from dynamic data sources, and can help prevent sensitive or unnecessary values from being written.

Since the beta, we have greatly expanded the scalability of Pipelines: we now support ingesting up to 3GB/s per stream. We’ve also expanded the feature set and integration with other Cloudflare systems:

  • Cloudflare Logpush integration: you can transform Cloudflare logs with SQL and store them as compressed Parquet files or Iceberg tables, ready to query with Basin SQL or another engine.
  • Worker bindings are schema-aware. Running wrangler types generates TypeScript types from a stream's schema, catching missing fields and type mismatches before deployment.
  • Data quality errors are visible. The dashboard and GraphQL API surface dropped events and distinguish missing fields, type mismatches, parse failures, and null values.
  • The entire ingestion path can be infrastructure as code. Terraform resources cover the catalog, stream, sink, and the SQL that connects them.

Next we plan to expand Pipelines capabilities even further, including:

  • Custom partitioning when writing to Basin Catalog
  • Schema migrations, and updatable configuration and Pipelines SQL
  • Support for Iceberg V3, including the Variant type for efficient querying of semi-structured data
  • Stateful processing to support workloads such as streaming aggregations, joins, and incrementally updated materialized views

Basin Catalog

Basin Catalog was the first product we launched in the family last year. Since then, we’ve seen thousands of developers use Basin Catalog for simple use cases such as giving DuckDB a structured way to access analytics data in R2, all the way to developers building complete enterprise data sharing platforms, fully taking advantage of zero egress fees and easy-to-use APIs.

Basin Catalog is the easiest way to get started with Apache Iceberg. Just run:

npx wrangler basin catalog create CATALOG_NAME

You instantly get a fully managed Apache Iceberg REST catalog that automatically performs routine maintenance required to keep those tables performant and healthy.

When we announced the Data Platform, Basin Catalog had just added automatic compaction. Since then, it has evolved to maintain healthy tables as your data scales:

  • Per-table compaction policies let you choose target file sizes based on each table's access pattern.
  • Automatic snapshot expiration removes old Iceberg snapshots according to a retention policy, while preserving a minimum number of recent snapshots.
  • Unreferenced data-file cleanup reclaims storage when snapshots expire, without requiring a separate Spark maintenance job.
  • Manifest optimization consolidates and clusters fragmented manifests by partition before compaction, reducing metadata I/O during query planning.

We have some exciting features in the works for Basin Catalog including:

  • A new way for compaction to efficiently sort and cluster data for improved query performance
  • More granular auth controls for namespaces and tables
  • Jurisdiction support to adhere to data sovereignty and compliance requirements

Basin SQL

Basin SQL is our serverless, distributed query engine for Apache Iceberg tables stored in Basin Catalog. It’s designed for reading large datasets and automatically scales across Cloudflare's global network. There are no clusters or resources to provision, just a readily available API for you and your agents to immediately start querying your data.

At beta launch, Basin SQL was great at filtering and exploring large event and time-series tables. Over the last year, Basin SQL has evolved to support hundreds of functions including:

  • Standard and approximate aggregations, GROUP BY, HAVING, and schema-discovery commands
  • More than 190 scalar and aggregate functions across strings, timestamps, regular expressions, cryptography, statistics, arrays, maps, and structs
  • CASE expressions, common table expressions, casting, arithmetic, and EXPLAIN
  • Inner, outer, semi, and anti joins; subqueries; self-joins; and multi-table queries
  • DISTINCT, UNION, INTERSECT, and EXCEPT
  • Window functions, QUALIFY, grouping sets, rollups, and cubes
  • A suite of JSON functions

Suppose your Pipeline delivers application events into one table and account data into another. You can now join those tables, aggregate activity by customer, rank the results with a window function, and filter the ranking in one query:

WITH account_activity AS (
  SELECT
    a.plan,
    e.account_id,
    count(*) AS events,
    approx_distinct(e.user_id) AS active_users
  FROM analytics.events e
  JOIN analytics.accounts a
    ON e.account_id = a.account_id
  WHERE e.event_time >= '2026-09-01T00:00:00Z'
  GROUP BY a.plan, e.account_id
)
SELECT
  *,
  rank() OVER (PARTITION BY plan ORDER BY events DESC) AS activity_rank
FROM account_activity
QUALIFY rank() OVER (PARTITION BY plan ORDER BY events DESC) <= 10;

You can run Basin SQL from Wrangler or the API, or open the built-in editor in the Cloudflare dashboard. The editor provides syntax highlighting and autocomplete, a browser for namespaces and tables, query statistics and plans, and exportable results. It makes the path from a new table to a useful answer a matter of seconds.

The team isn’t stopping here and is currently working on:

  • Advanced statistics and adaptive scheduling to improve performance and efficiency of queries
  • Full data definition language (DDL) support directly from Basin SQL
  • Iceberg V3 support including support for the VARIANT and geospatial types

What comes next

Our future vision is that data infrastructure is completely abstracted away. Storage formats, products, and resources are just implementation details — important ones that help enable the important outcomes — but tend to get in the way. We’re building towards a platform where developers start with questions rather than CREATE statements or CLI commands. We’ve laid the foundation for that vision, and now we’re building towards that vision including:

  • Support for the latest Apache Iceberg spec across the entire platform, unlocking more flexible ways to use your data
  • Push-button ingestion sources and destinations, with zero-configuration connections across Cloudflare's developer and observability products
  • Advanced adaptive table-maintenance strategies in Basin Catalog that automatically organize data around real query patterns
  • Continued expansion of SQL compatibility, performance, and observability for increasingly complex analytical workloads
  • More ways to continuously process data in real-time and trigger actions based on the signals within the data
  • Tools for adhering to data compliance and sovereignty rules across the platform

We will continue to build with open standards: using open formats and protocols, contributing improvements to the projects we depend on, and making sure your data remains available to the broader ecosystem.

Get started

Basin Pipelines, Basin Catalog, and Basin SQL are generally available today. You can use them together as an end-to-end platform or adopt the parts that fit your existing architecture.

Follow the getting started tutorial to ingest events, create an Apache Iceberg table in Basin Catalog, and query it with Basin SQL. Visit the Basin documentation for product guides, pricing, limits, and integrations.

Existing Cloudflare Pipelines, R2 Data Catalog, and R2 SQL configurations will continue to work.

We are excited to see what you build. Share your feedback with us in the Cloudflare Developer Discord.