What’s new with Google Data Cloud

Source: Google Cloud Blog•

What’s new with Google Data Cloud

Recent product news and updates from our data analytics, database and business intelligence teams.

October 5 - October 9

  • Data Agent Kit is now generally available to bring Google Data Cloud to any coding agentData Agent Kit provides a free collection of Model Context Protocol (MCP) tools and Google-authored agent skills that connect more than 15 Google Data Cloud services directly to coding agents across VS Code, Antigravity, Cursor, Claude Code, and Codex. This release expands support for developers and data practitioners to inspect schemas, author queries, and build end-to-end data pipelines in natural language directly from their IDE or CLI. Read the announcement blog.

Data Agent Kit is now generally available to bring Google Data Cloud to any coding agentData Agent Kit provides a free collection of Model Context Protocol (MCP) tools and Google-authored agent skills that connect more than 15 Google Data Cloud services directly to coding agents across VS Code, Antigravity, Cursor, Claude Code, and Codex. This release expands support for developers and data practitioners to inspect schemas, author queries, and build end-to-end data pipelines in natural language directly from their IDE or CLI. Read the announcement blog.

  • Spanner Omni is now generally available for deploy-anywhere, multi-model workloadsSpanner Omni brings Spanner's distributed SQL, graph, vector, and key-value capabilities to any environment, including on-premises data centers, other public clouds via Kubernetes or VMs, and local developer laptops. Organizations can build and run multi-model applications with consistent APIs and query engines across hybrid and multi-cloud footprints while seamlessly scaling to managed Spanner on Google Cloud. Read the launch blog and explore how Spanner also removed cumulative mutation limits for DML transactions.

Spanner Omni is now generally available for deploy-anywhere, multi-model workloadsSpanner Omni brings Spanner's distributed SQL, graph, vector, and key-value capabilities to any environment, including on-premises data centers, other public clouds via Kubernetes or VMs, and local developer laptops. Organizations can build and run multi-model applications with consistent APIs and query engines across hybrid and multi-cloud footprints while seamlessly scaling to managed Spanner on Google Cloud. Read the launch blog and explore how Spanner also removed cumulative mutation limits for DML transactions.

  • AlloyDB delivers PostgreSQL for agents and native BM25 hybrid searchWe announced PostgreSQL for agents in AlloyDB, introducing an agentic database architecture engineered for real-time data access at agent scale with workload isolation. In addition, native BM25 ranking is now in Preview for AlloyDB and Cloud SQL, allowing developers to combine keyword search and vector embeddings directly in PostgreSQL without maintaining an external search engine. For hybrid deployments, the AlloyDB Omni RPM Orchestrator is now generally available, and you can dive deeper in our new Build AI apps and agents with AlloyDB eBook.

AlloyDB delivers PostgreSQL for agents and native BM25 hybrid searchWe announced PostgreSQL for agents in AlloyDB, introducing an agentic database architecture engineered for real-time data access at agent scale with workload isolation. In addition, native BM25 ranking is now in Preview for AlloyDB and Cloud SQL, allowing developers to combine keyword search and vector embeddings directly in PostgreSQL without maintaining an external search engine. For hybrid deployments, the AlloyDB Omni RPM Orchestrator is now generally available, and you can dive deeper in our new Build AI apps and agents with AlloyDB eBook.

  • Lakehouse runtime catalog adds regional endpoints across 22 regions and GA support for flexible column namesLakehouse runtime catalog regional endpoints are now live across 22 Google Cloud regions across the Americas, Europe, Asia-Pacific, and the Middle East, supporting both the Apache Iceberg REST catalog endpoint and the Apache Hive catalog endpoint to help customers meet strict data residency and sovereignty requirements. In addition, flexible column names for Apache Iceberg tables are now generally available by default across Iceberg external tables and managed tables, enabling seamless querying of columns containing spaces, hyphens, and Unicode characters. We also announced the Preview of cross-cloud caching and connections for the borderless Lakehouse.

Lakehouse runtime catalog adds regional endpoints across 22 regions and GA support for flexible column namesLakehouse runtime catalog regional endpoints are now live across 22 Google Cloud regions across the Americas, Europe, Asia-Pacific, and the Middle East, supporting both the Apache Iceberg REST catalog endpoint and the Apache Hive catalog endpoint to help customers meet strict data residency and sovereignty requirements. In addition, flexible column names for Apache Iceberg tables are now generally available by default across Iceberg external tables and managed tables, enabling seamless querying of columns containing spaces, hyphens, and Unicode characters. We also announced the Preview of cross-cloud caching and connections for the borderless Lakehouse.

  • Unlock predictive and agent-ready insights with BigQuery augmented analytics TVFs, TabFM, and identity columnsBigQuery introduced augmented analytics table-valued functions (TVFs): six built-in functions that automate metric diagnostics, trend analysis, and pattern discovery directly where your data lives and integrate as skills for AI agents. We also introduced TabFM in BigQuery to bring tabular foundation model capabilities to predictive analytics, alongside BigQuery identity columns to automatically generate sequential surrogate keys and simplify data warehouse pipelines.

Unlock predictive and agent-ready insights with BigQuery augmented analytics TVFs, TabFM, and identity columnsBigQuery introduced augmented analytics table-valued functions (TVFs): six built-in functions that automate metric diagnostics, trend analysis, and pattern discovery directly where your data lives and integrate as skills for AI agents. We also introduced TabFM in BigQuery to bring tabular foundation model capabilities to predictive analytics, alongside BigQuery identity columns to automatically generate sequential surrogate keys and simplify data warehouse pipelines.

  • Dataflow adds Pause/Resume and NVIDIA RTX PRO 6000 Blackwell GPU support for large-scale AI workloadsNew Dataflow capabilities make it easier to run high-throughput streaming and batch AI pipelines at scale, including Pause/Resume job controls and support for NVIDIA RTX PRO 6000 Blackwell GPUs. And for teams running Apache Spark workloads, learn how to mitigate compute stockouts and maximize availability using flexible VMs.

Dataflow adds Pause/Resume and NVIDIA RTX PRO 6000 Blackwell GPU support for large-scale AI workloadsNew Dataflow capabilities make it easier to run high-throughput streaming and batch AI pipelines at scale, including Pause/Resume job controls and support for NVIDIA RTX PRO 6000 Blackwell GPUs. And for teams running Apache Spark workloads, learn how to mitigate compute stockouts and maximize availability using flexible VMs.

  • Unlock up to 3x higher QPS and microsecond latency with Memorystore for Valkey 9.1Memorystore for Valkey 9.1 is now available, delivering up to 3x higher queries per second (QPS) and microsecond-level caching latencies to accelerate high-throughput applications and real-time AI agent state management. Read the blog.

Unlock up to 3x higher QPS and microsecond latency with Memorystore for Valkey 9.1Memorystore for Valkey 9.1 is now available, delivering up to 3x higher queries per second (QPS) and microsecond-level caching latencies to accelerate high-throughput applications and real-time AI agent state management. Read the blog.

  • Customer momentum across Google Data Cloud — Yahoo, PayPal, Surescripts, Airwallex, and Lucius AIOrganizations across industries are scaling their data and AI foundations on Google Data Cloud. Yahoo uses flexible VMs in Managed Service for Apache Spark to reduce provisioning failures by 85%, while PayPal migrated mission-critical analytics workloads to Managed Service for Apache Spark to scale and reduce operational overhead. On the database front, Surescripts migrated 30.5 billion annual healthcare transactions to AlloyDB with 99.998% uptime, Airwallex manages more than 5,000 databases with 99.99% availability on Cloud SQL with a team of four DBAs, and Lucius AI runs a 210,000-tender global platform on AlloyDB and MCP with 47x faster ScaNN vector search.

Customer momentum across Google Data Cloud — Yahoo, PayPal, Surescripts, Airwallex, and Lucius AIOrganizations across industries are scaling their data and AI foundations on Google Data Cloud. Yahoo uses flexible VMs in Managed Service for Apache Spark to reduce provisioning failures by 85%, while PayPal migrated mission-critical analytics workloads to Managed Service for Apache Spark to scale and reduce operational overhead. On the database front, Surescripts migrated 30.5 billion annual healthcare transactions to AlloyDB with 99.998% uptime, Airwallex manages more than 5,000 databases with 99.99% availability on Cloud SQL with a team of four DBAs, and Lucius AI runs a 210,000-tender global platform on AlloyDB and MCP with 47x faster ScaNN vector search.

September 7 - September 10

  • Pub/Sub SMTs can now AI Inference your Gemini Enterprise Agent Platform models!Pub/Sub AI Inference SMTs allow you to apply inference on an incoming stream of events using models hosted in Gemini Enterprise Agent Platform. The model’s prediction is appended to your event, making it available for downstream processing in your data warehouse (like BigQuery) or operational database (like BigTable). This feature, now generally available, can dramatically simplify or enhance anomaly detection systems you are operating.

Pub/Sub SMTs can now AI Inference your Gemini Enterprise Agent Platform models!Pub/Sub AI Inference SMTs allow you to apply inference on an incoming stream of events using models hosted in Gemini Enterprise Agent Platform. The model’s prediction is appended to your event, making it available for downstream processing in your data warehouse (like BigQuery) or operational database (like BigTable). This feature, now generally available, can dramatically simplify or enhance anomaly detection systems you are operating.

  • PostgreSQL Source Connector is now generally available in Managed Service for Apache Kafka!Managed Service for Apache Kafka’s PostgreSQL connector allows customers to capture changes from their PostgreSQL database and ingest them into their Kafka infrastructure with low latency. This source connector is compatible with Cloud SQL for Postgres, AlloyDB, and self-managed PostgreSQL databases. Try this along with our entire portfolio of managed connectors, including MirrorMaker 2.0, BigQuery, Cloud Storage, and Pub/Sub! E-mail kafka-hotline@google.com if you have questions or feedback!
  • Pause-on-failure for Dataflow batch jobs is GADataflow pause-on-failure enables you to preserve the state of a batch Dataflow job before it fails. By pausing your Dataflow job, you can address issues that are external to the pipeline and resume processing without losing completed work. This helps you better manage resource costs and improve job reliability when you face temporary outages or capacity constraints.

Pause-on-failure for Dataflow batch jobs is GADataflow pause-on-failure enables you to preserve the state of a batch Dataflow job before it fails. By pausing your Dataflow job, you can address issues that are external to the pipeline and resume processing without losing completed work. This helps you better manage resource costs and improve job reliability when you face temporary outages or capacity constraints.

  • The insertAll API is now the BigQuery Storage Write API (REST)The legacy insertAll streaming API is now rebranded as the BigQuery Storage Write API (REST). By dropping the "legacy" label, developers can confidently build long-term HTTP-based streaming workflows. This stateless JSON-over-HTTPS endpoint offers a lightweight alternative to heavy gRPC libraries—ideal for serverless web apps, IoT telemetry, and AI logging. The transition is seamless for existing users, requiring zero code changes and offering 100% backward compatibility. However, the Storage Write API (gRPC) version remains the recommended standard for high-throughput, continuous pipelines.

The insertAll API is now the BigQuery Storage Write API (REST)The legacy insertAll streaming API is now rebranded as the BigQuery Storage Write API (REST). By dropping the "legacy" label, developers can confidently build long-term HTTP-based streaming workflows. This stateless JSON-over-HTTPS endpoint offers a lightweight alternative to heavy gRPC libraries—ideal for serverless web apps, IoT telemetry, and AI logging. The transition is seamless for existing users, requiring zero code changes and offering 100% backward compatibility. However, the Storage Write API (gRPC) version remains the recommended standard for high-throughput, continuous pipelines.

August 31 - September 4

  • Stateful processing is available in BigQuery continuous queries in PreviewStateful operations significantly expand what’s possible with BigQuery continuous queries. This feature allows users to leverage functions like JOINs, aggregations, and windowing functions directly in their streaming queries. Now you can calculate metrics over time (for example, a 30-minute average) to power your downstream applications and AI agents with much richer, real-time signals.Try out our feature here and share your feedback with bq-continuous-queries-feedback@google.com!
  • Synthetic data generator tool is available for Managed Service for KafkaYou’ve launched your first Kafka cluster. Now what? The next thing to do is to produce some data to the cluster, but that involves modifying a client application somewhere or spinning up a virtual machine. The synthetic data generator tool, now generally available, can start sending mock data to your cluster in 3 clicks, and will get data streaming into your cluster in less than two minutes. The perfect utility for those moments you just want to test your cluster and new features. Try our quickstart today!

Synthetic data generator tool is available for Managed Service for KafkaYou’ve launched your first Kafka cluster. Now what? The next thing to do is to produce some data to the cluster, but that involves modifying a client application somewhere or spinning up a virtual machine. The synthetic data generator tool, now generally available, can start sending mock data to your cluster in 3 clicks, and will get data streaming into your cluster in less than two minutes. The perfect utility for those moments you just want to test your cluster and new features. Try our quickstart today!

  • Dataflow pipeline updates are faster & more flexibleDataflow pipeline updates can now stop-and-replace pipelines, a major addition to the existing in-place-update feature. The new parallel pipeline option accelerates the migration between the old & new pipeline, resulting in reduced disruption to your business. You can also set a timeout on drains that prevents runaway costs for your pipeliness in the event of stuck processing. This feature is generally available. Try it here!

Dataflow pipeline updates are faster & more flexibleDataflow pipeline updates can now stop-and-replace pipelines, a major addition to the existing in-place-update feature. The new parallel pipeline option accelerates the migration between the old & new pipeline, resulting in reduced disruption to your business. You can also set a timeout on drains that prevents runaway costs for your pipeliness in the event of stuck processing. This feature is generally available. Try it here!

July 6 - July 10

  • New Lakehouse managed tables now in preview Lakehouse tables for Apache Iceberg are now in preview and available in the console. By using Google-managed Apache Iceberg tables in Lakehouse, you can eliminate the overhead of maintaining duplicate data pipelines and complex synchronization logic between BigQuery and open-source engines. This unified table format delivers native, multi-engine read and write interoperability, allowing you to run concurrent DML/DDL operations across diverse analytics tools on a single, shared storage layer. Built-in automated table management handles painful background optimization tasks like compaction and partition tuning, freeing up your team to focus on building rather than managing storage maintenance.

New Lakehouse managed tables now in preview Lakehouse tables for Apache Iceberg are now in preview and available in the console. By using Google-managed Apache Iceberg tables in Lakehouse, you can eliminate the overhead of maintaining duplicate data pipelines and complex synchronization logic between BigQuery and open-source engines. This unified table format delivers native, multi-engine read and write interoperability, allowing you to run concurrent DML/DDL operations across diverse analytics tools on a single, shared storage layer. Built-in automated table management handles painful background optimization tasks like compaction and partition tuning, freeing up your team to focus on building rather than managing storage maintenance.

June 1 - June 5

  • Beyond the Query: Powering AI Agents with Bigtable, Firestore & Memorystore Discover the latest advancements in Google Cloud's NoSQL Database portfolio, including Bigtable, Firestore, and Memorystore. This series is designed for a broad audience: whether you are exploring these databases for the first time or are an existing user looking to leverage the new capabilities announced at Next '26. Register here to secure your spot!

Beyond the Query: Powering AI Agents with Bigtable, Firestore & Memorystore Discover the latest advancements in Google Cloud's NoSQL Database portfolio, including Bigtable, Firestore, and Memorystore. This series is designed for a broad audience: whether you are exploring these databases for the first time or are an existing user looking to leverage the new capabilities announced at Next '26. Register here to secure your spot!

  • Cloud Engineer's AI Toolkit Workshops: Solve data-driven challenges with BigQuery, AlloyDB, Gemini and more. Hosted by Google Cloud Labs, this highly technical event is built specifically for Platform Engineers, SREs, and cloud infrastructure teams ready to bridge the gap between AI prototypes and production-grade deployments. Look out for more locations coming soonToronto - June 25 (Data Cloud) | RSVP HereChicago - June 30 (Data Cloud) | RSVP Here

Cloud Engineer's AI Toolkit Workshops: Solve data-driven challenges with BigQuery, AlloyDB, Gemini and more. Hosted by Google Cloud Labs, this highly technical event is built specifically for Platform Engineers, SREs, and cloud infrastructure teams ready to bridge the gap between AI prototypes and production-grade deployments. Look out for more locations coming soonToronto - June 25 (Data Cloud) | RSVP HereChicago - June 30 (Data Cloud) | RSVP Here

  • Start a 10-day Bigtable free trial with a 1 node SSD cluster and up to 500GB of storage capacity. With no credit card required to start, you can easily ingest workloads and manage workloads that require low-latency, high-throughput, and predictable access. Plus, new Google Cloud customers get $300 in free credits on signup.

May 11 - May 15

  • Managed Service for Apache Airflow has launched a wave of new features, including the general availability of Airflow 3.1, AI-powered agentic troubleshooting, a new managed Airflow MCP Server for custom agent integration, and declarative YAML-based orchestration pipelines—discover all the details in thefull blog post.

April 20 - April 24

  • Google-built ODBC Driver for BigQuery is now available in PreviewWe are excited to announce the launch of the new, Google-built ODBC driver for BigQuery. This new open-source driver provides a direct, high-performance connection for applications to BigQuery and is developed entirely in-house by Google. Download a new driver and connect your application to BigQuery.

April 13 - April 17

  • We announced we are reintroducing Data Studio to play a significant role in the AI era, expanding from data visualizations and reports to host BigQuery conversational agents and data apps built in Colab notebooks.
  • We announced BigQuery Graph is now available in preview, offering an easy-to-use, highly scalable graph analytics solution, empowering data professionals to model, analyze and visualize massive-scale relationships in an entirely new way.

April 6 - April 10

  • We introduced Conversational Analytics for Looker Embedded environments, enabling users to add natural language experiences to their own custom data-driven applications, powered by Gemini.
  • We expanded Looker’s capabilities for faster ad-hoc analysis, with the introduction of self-service Explores, enabling you to bring your own data to Looker’s semantic layer and gain instant access to insights in a governed data environment.

March 23 - March 27

  • We showed you how you can scale your reads with Cloud SQL autoscaling read pools. This feature allows you to provision multiple read replicas that are accessible via a single read endpoint and to dynamically adjust your read capability based on real-time application needs.
  • Our customers are leveraging the full power of Conversational Analytics and Looker to drive major business and technical breakthroughs in the AI era. Companies like Telenor, Pet Circle, Fluent Commerce, Lighthouse Intelligence, Wego, and ROLLER are turning data into insights and actions, grounded by Looker’s semantic layer.

March 16 - March 20

  • We introduced an enhanced Gemini assistant in BigQuery Studio, transforming the agent from a code assistant into a fully context-aware analytics partner.

February 23 - February 27

  • We introduced managed and remote MCP support for Google Cloud databases, including AlloyDB, Spanner, Cloud SQL, Bigtable and Firestore, to power the next generation of agents. This announcement extends the ability for AI models to plan, build, and solve complex problems, connecting to the database tools our customers leverage daily as the backbone of their work environment.

We introduced managed and remote MCP support for Google Cloud databases, including AlloyDB, Spanner, Cloud SQL, Bigtable and Firestore, to power the next generation of agents. This announcement extends the ability for AI models to plan, build, and solve complex problems, connecting to the database tools our customers leverage daily as the backbone of their work environment.

  • We outlined how you can build a conversational agent in BigQuery using the Conversational Analytics API to help you build context-aware agents that can understand natural language, query your BigQuery data, and deliver answers in text, tables, and visual charts.

February 16 - February 20

  • Our customers are leveraging the full power of Looker to drive major business and technical breakthroughs. Companies like Arrive, Audika, Carousell, Framebridge, GumGum, Intel, Overdose Digital, Ocean Network Express, Subskribe and Promevo are leveraging Looker’s newest AI-driven capabilities, including Conversational Analytics, to transform data to insights and actions, and empower their entire organization with a single source of truth, powered by Looker’s semantic layer.

February 2 - February 6

  • Join us on March 4 for our webinar, Win Your AI Strategy with Cloud SQL Enterprise Plus, to learn how to power your generative AI workloads with 3x higher performance and 99.99% availability. Register today to discover how to build a scalable, enterprise-grade foundation for your most demanding AI applications.

January 26 - January 30

January 19 - January 23

January 12 - January 16

  • Data Analytics
  • Databases
  • Business Intelligence

What this article says