Skip to content
Home » Google BigQuery Borderless Lakehouse: The Business Case

Google BigQuery Borderless Lakehouse: The Business Case

The pipeline you built is now the liability

The warehouse is full. The lake is full. But the analysts are still waiting — because the pipeline that was supposed to connect them broke again at 3 a.m., and nobody noticed until the Monday morning dashboard showed last week’s numbers.

That is not a tooling problem. It is an architectural assumption: that data must be moved before it can be used. Copy it, transform it, land it somewhere else, then query it. Repeat. The result is a network of brittle ETL jobs, duplicated storage costs, unpredictable egress fees, and a governance surface that grows every time a new cloud or business unit joins the picture.

Google’s borderless Lakehouse architecture challenges that assumption directly. Query data where it lives. Enforce policy at the catalog layer. Move compute to the data, not the other way around.

Whether that holds in practice depends on architecture, economics, and — critically — compliance constraints that Google’s own documentation does not minimise.

What the borderless Lakehouse actually does

The architecture decouples three things that traditional stacks bundle together: storage, compute, and governance. Storage stays where it is — Amazon S3, Azure Data Lake Storage Gen2, or Google Cloud Storage. Compute is selected per workload: BigQuery for gold-layer aggregations and exact-match federated queries, Managed Service for Apache Spark with Lightning Engine for memory-heavy vectorised joins and complex transformations. Governance travels with the data through a unified catalog, not through the pipeline.

The catalog layer is Apache Iceberg. As Hightouch noted in their June 2026 integration announcement, a single Iceberg table can be read and written by BigQuery, Spark, Trino, and other open engines without copies or proprietary formats — an open-storage model that keeps composable architectures viable as engines change.

The runtime catalog handles each query in six steps: the engine submits SQL; the catalog identifies the table and its metadata location; IAM and fine-grained security policies are validated; the catalog returns metadata and, when credential vending is enabled, a short-lived token; the engine uses that token to read data files directly from Cloud Storage; the engine processes and returns results. No data leaves its origin store — only the query result travels.

Three catalog endpoint options exist for Iceberg interoperability. The Apache Iceberg REST catalog endpoint is the recommended interface for new workloads, offering full read and write interoperability with Spark, Flink, and Trino. The Custom Apache Iceberg catalog for BigQuery endpoint is used primarily for Iceberg tables managed by BigQuery and for workloads transitioning to the Lakehouse architecture. The Apache Hive catalog endpoint is currently in Preview, providing compatibility for workloads that depend on the Apache Hive metastore interface.

Security extends beyond the catalog. Lakehouse extends BigQuery’s fine-grained row- and column-level security to tables on remote object stores — S3, Azure Data Lake Storage Gen2, and Cloud Storage — through access delegation. This decouples access to the table from the underlying cloud storage credential, so users and pipelines receive only the slice of data they are authorised to see, regardless of which engine they use to query it.

The economics

The borderless Lakehouse targets three cost levers: data transfer, compute, and — for AI workloads — token consumption.

On data transfer: Cross-Cloud Interconnect establishes a private connection between Google Cloud and AWS, replacing public internet egress with a flat, predictable rate. Google’s cross-cloud data access documentation states that private interconnect produces potentially lower egress charges from AWS compared to internet egress, alongside more predictable network latency and bandwidth. The qualifier “potentially” matters — actual savings depend on your interconnect capacity and query volume.

On compute: eliminating physical data movement removes the storage duplication and pipeline maintenance overhead that accumulates in traditional ETL architectures. Google’s Architecture Center documentation describes this as reducing operational complexity and lowering total cost of ownership by querying data directly where it resides, avoiding the storage costs and latency associated with physical data movement.

On token consumption: Google’s July 2026 announcement reports customers seeing a 230x reduction in token consumption for agentic workloads. That figure is attributed to Knowledge Catalog filtering the precise business context each prompt requires — preventing token bloat and eliminating unnecessary reasoning loops — combined with BigQuery AI’s built-in controls that let operators estimate token usage pre-query, enforce strict limits, and use an optimised mode that automatically selects smaller, distilled models. The 230x figure is presented as a customer observation, not a guaranteed outcome; workloads with different schema complexity or agent patterns will see different results.

The compliance constraints that cannot be footnoted

Eliminating physical data movement sounds clean on a diagram. In practice, it shifts complexity rather than removing it — and two constraints in particular require deliberate handling before deployment.

The first is cross-jurisdiction caching. Google’s own documentation is explicit: if your AWS S3 data resides in the EU and your BigQuery compute runs in us-east4, cross-cloud querying stores cached copies of that remote data in the US region. The user or administrator creating the connection must ensure that cross-jurisdiction caching complies with their organisation’s data residency, sovereignty, and compliance requirements. For GDPR-regulated organisations, this is a mapping exercise that must happen before a single connection is configured.

The second is encryption key custody. Customer-managed encryption keys (CMEK) are not supported for Lakehouse caching. All cached data blocks are encrypted at rest using Google-owned and Google-managed encryption keys. For organisations with contractual or regulatory key-custody requirements, that is a blocker, not a configuration option.

The honest summary: the borderless Lakehouse reduces pipeline complexity and can lower egress cost, but it introduces a new governance surface in the catalog and network layer. It is differently complex, not simply simpler. A lower capture rate on the cost savings — because interconnect capacity is constrained, or because compliance requirements force data to remain in a single region — lowers the return. That trade-off should be evaluated against your specific workload and regulatory profile, not assumed away.

What operators should do with this now

Map data residency exposure before configuring any cross-cloud connection. The cross-cloud data access documentation specifies exactly which Google Cloud region will hold cached copies of remote data. Run that against your compliance register before deployment, not after a data residency audit surfaces the gap.

Layer the medallion architecture correctly. Google’s Lakehouse documentation recommends structuring data into bronze (raw ingestion), silver (cleansed and conformed), and gold (curated business aggregations) layers, with BigQuery reserved for the gold consumption layer. BigQuery’s query performance and concurrency advantages are most pronounced there. Running BigQuery against raw bronze-layer data wastes both cost and speed; memory-heavy transformations belong in Managed Spark with Lightning Engine.

Treat Knowledge Catalog as a prerequisite for agentic workloads, not an optional add-on. The July 2026 Google Cloud announcement states that grounded, high-accuracy agent results depend on natural-language-to-SQL translations being anchored in curated business schemas inside the catalog. Agents querying uncontrolled schemas produce hallucinations. The catalog is the trust layer — and for cross-cloud deployments, it is also where Databricks authentication is federated via Lakehouse REST catalog federation with Secret Manager.

The borderless Lakehouse is an architectural posture, not a product you switch on. The Building a borderless open data lakehouse Codelab is the most practical starting point: it walks through provisioning the network infrastructure, executing PySpark workflows, and configuring AI agents end to end. Start there, with your residency constraints already mapped.

— Eagentix


Eagentix helps growth-focused enterprises redesign and automate manual business processes. We combine executive strategy, implementation support, and managed services to build dependable operations across Southeast Asia.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *