Enterprises and individual developers frequently run multiple AI inference models. The right architecture can simplify how the models are called while also providing centralized governance. In this post, we'll look at two reference architectures focused on networking AI inference model serving: one for Google Kubernetes Engine (GKE) and one all other backend types. First, we'll explore the commonalities between the reference architectures that you'll see later. Then we'll explore unique components of the architecture for GKE backends and finally, we'll go over the elements of the architecture for all backend types.
The entry point
You can expose your model deployment behind a stable, secure, and reliable entry point that acts as the front end for inference calls. This entry point also acts as a control zone where policy, security, and logic can be enforced. Both the Cloud Load balancer and the Inference Gateway provide entry point capability. These types of endpoints can terminate secure connections with TLS, integrate with API management components, extend functionally with service extensions, and capitalize on capabilities of Model Armor for added security.
Common services in the designs
Both reference designs use these services:
- Private Service Connect inference endpoint: Anchors the entry point inside your consumer Virtual Private Cloud (VPC) network. Traffic hits a private internal IP address, keeping inference calls in your private network.
Private Service Connect inference endpoint: Anchors the entry point inside your consumer Virtual Private Cloud (VPC) network. Traffic hits a private internal IP address, keeping inference calls in your private network.
- Apigee API Management (Optional): Integrates via an Apigee Extension Processor callout to handle client identity verification, rate limits, and quota enforcement before requests ever reach compute resources.
Apigee API Management (Optional): Integrates via an Apigee Extension Processor callout to handle client identity verification, rate limits, and quota enforcement before requests ever reach compute resources.
- Model Armor: Serves as an inline AI safety checkpoint, screening prompts and output completions against prompt injection and sensitive data leakage.
Model Armor: Serves as an inline AI safety checkpoint, screening prompts and output completions against prompt injection and sensitive data leakage.
Design pattern serving on GKE only
This section focuses on a GKE-only backend design. To understand the full end-to-end concept, please read the entire architecture document Networking for AI inference model serving on GKE. The design pattern is based on this diagram:
In addition to the common services identified in the previous section, the design for GKE uses these components:
- GKE Inference Gateway: Deployed as an internal Application Load Balancer (gke-l7-rilb). It acts as a specialized ingress engine that parses incoming request payloads, evaluates HTTPRoute rules, and steers queries to appropriate model-serving targets.
GKE Inference Gateway: Deployed as an internal Application Load Balancer (gke-l7-rilb). It acts as a specialized ingress engine that parses incoming request payloads, evaluates HTTPRoute rules, and steers queries to appropriate model-serving targets.
- Inference pools: A logical group containing replicas of the same model. When the Gateway receives a prompt, it evaluates HTTPRoute rules to select the appropriate inference pool based on the model identifier. Pools have an initial size and can be configured to autoscale dynamically.
Inference pools: A logical group containing replicas of the same model. When the Gateway receives a prompt, it evaluates HTTPRoute rules to select the appropriate inference pool based on the model identifier. Pools have an initial size and can be configured to autoscale dynamically.
- Model replica sets: Individual model replicas (inference server instances) deployed across single-node or multi-node GPU or TPU node pools. A replica set represents a uniform group of these model replicas.
Model replica sets: Individual model replicas (inference server instances) deployed across single-node or multi-node GPU or TPU node pools. A replica set represents a uniform group of these model replicas.
The traffic flow GKE example
A client application that uses this GKE-based architecture to call a backend model would go through a flow like this:
- Ingress: A client application in the consumer VPC issues an OpenAI-compatible API call to the local Private Service Connect endpoint, routing directly to the GKE Inference gateway using a regional internal Application Load Balancer.
Ingress: A client application in the consumer VPC issues an OpenAI-compatible API call to the local Private Service Connect endpoint, routing directly to the GKE Inference gateway using a regional internal Application Load Balancer.
- Payload inspection: The Gateway reads the target model parameter specified in the request body and adds it to the HTTP headers.
Payload inspection: The Gateway reads the target model parameter specified in the request body and adds it to the HTTP headers.
- Control plane validation: If Apigee is used, it checks client credentials and quotas. Model Armor screens the prompt for policy violations or data leakage.
Control plane validation: If Apigee is used, it checks client credentials and quotas. Model Armor screens the prompt for policy violations or data leakage.
- Backend selection: The Gateway evaluates HTTPRoute mappings to identify the target pool, matches shared prefix cache context, and routes to the lowest-load GPU or TPU replica based on real-time Prometheus data.
Backend selection: The Gateway evaluates HTTPRoute mappings to identify the target pool, matches shared prefix cache context, and routes to the lowest-load GPU or TPU replica based on real-time Prometheus data.
- Egress: The replica runs the inference workload. Output tokens pass through Model Armor for final response verification before streaming back over the private connection.
Egress: The replica runs the inference workload. Output tokens pass through Model Armor for final response verification before streaming back over the private connection.
Design pattern serving on all backends
This section focuses on multiple backend types which can be used for inference, and it provides an overview of the architecture in the following diagram. To understand the full end-to-end concept, please read the entire architecture document Networking for AI inference model serving on all backends.
For architectures spanning mixed environments such as GKE, Cloud Run, Agent Platform, on-premises data centers, or external clouds, these additional components are used:
- Regional internal Application Load Balancer: Serves as the central Layer 7 routing proxy that manages routing logic, SSL termination, and Service Extensions callouts.
Regional internal Application Load Balancer: Serves as the central Layer 7 routing proxy that manages routing logic, SSL termination, and Service Extensions callouts.
- Inference Payload Processor (Service Extensions): This is similar to body-based routing as used in the GKE Inference Gateway, but to enable the functionality on an Application Load Balancer a service extension is needed. A lightweight Cloud Run callout inspects the JSON body of incoming OpenAI API requests, extracts the target model identifier, and writes an X-Gateway-Model-Name header to drive URL map routing.
Inference Payload Processor (Service Extensions): This is similar to body-based routing as used in the GKE Inference Gateway, but to enable the functionality on an Application Load Balancer a service extension is needed. A lightweight Cloud Run callout inspects the JSON body of incoming OpenAI API requests, extracts the target model identifier, and writes an X-Gateway-Model-Name header to drive URL map routing.
- Network Endpoint Group (NEG): Delivers flexible routing to heterogeneous backends based on the injected model header.
Network Endpoint Group (NEG): Delivers flexible routing to heterogeneous backends based on the injected model header.
All-backends traffic flow example
A client application that uses this architecture to call a backend model would go through a flow like this:
- Private ingress: The client application targets the Private Service Connect endpoint over private IP address space. The regional internal Application Load Balancer receives the request.
Private ingress: The client application targets the Private Service Connect endpoint over private IP address space. The regional internal Application Load Balancer receives the request.
- Model name extraction: The load balancer sends the payload to the Cloud Run body-based router callout, which inspects the JSON payload and injects the X-Gateway-Model-Name header.
Model name extraction: The load balancer sends the payload to the Cloud Run body-based router callout, which inspects the JSON payload and injects the X-Gateway-Model-Name header.
- Policy and safety enforcement: The request passes to Apigee for identity and quota validation, then to Model Armor to scrub sensitive data and block malicious prompts.
Policy and safety enforcement: The request passes to Apigee for identity and quota validation, then to Model Armor to scrub sensitive data and block malicious prompts.
- NEG routing: The load balancer URL map inspects the model header and forwards the request to the matching backend NEG (Agent Platform, GKE, Cloud Run, Hybrid, or Internet).
NEG routing: The load balancer URL map inspects the model header and forwards the request to the matching backend NEG (Agent Platform, GKE, Cloud Run, Hybrid, or Internet).
- Private delivery: The target backend executes the model prompt, Model Armor screens the completion, and the result returns privately along the ingress path.
Private delivery: The target backend executes the model prompt, Model Armor screens the completion, and the result returns privately along the ingress path.
- Developers & Practitioners
- Networking





