Deploy voice agents with Pipecat, Deepgram and Amazon SageMaker AI

Фото: x_OI (Pixabay) — https://pixabay.com/photos/microphone-vintage-cromatic-mic-5340340/

Deploy voice agents with Pipecat, Deepgram and Amazon SageMaker AI

Source: Amazon Web Services, Inc.

Learn how to deploy a real-time voice agent on AWS using Pipecat to orchestrate Deepgram speech models on Amazon SageMaker AI bidirectional streaming and Claude on Amazon Bedrock, keeping latency low and audio inside your VPC.

•Updated: October 5, 2026

Pipecat uses a three-layer architecture to integrate speech models running on Amazon SageMaker AI. Understanding this pattern helps you wrap additional models beyond the Deepgram services included in the sample. The source code for all layers lives in the Pipecat repository under pipecat/services/aws/sagemaker/.

Layer 1: SageMakerBidiClient

At the base, Pipecat provides SageMakerBidiClient — a reusable HTTP/2 client that handles the SageMaker bidirectional streaming protocol. It manages SigV4 authentication, session lifecycle, and binary/text message framing:

The client connects to runtime.sagemaker.<region>.amazonaws.com:8443 and uses InvokeEndpointWithBidirectionalStream from the AWS SDK. The model_invocation_path and model_query_string parameters map to the container’s internal routing, so you can target different model APIs on the same endpoint.

Layer 2: Service wrapper (TTSService or STTService)

The service wrapper extends Pipecat’s base TTSService or STTService class and implements the model-specific protocol. For example, DeepgramSageMakerTTSService translates Pipecat’s run_tts(text) interface into Deepgram’s WebSocket protocol messages:

A background task (_process_responses) continuously reads from the BiDi stream, distinguishing binary audio payloads from JSON control messages (Flushed, Warning, Error, Close) and routing them to the appropriate handler.

The STT wrapper follows the same pattern with model_invocation_path="v1/listen", sending raw audio by using send_audio_chunk() and parsing Deepgram’s transcription JSON responses. It also sends KeepAlive messages every 5 seconds to maintain the connection during silence, and Finalize when VAD detects end-of-speech to flush partial results.

Layer 3: Pipeline integration

The service factory (backend/voice-agent/app/services/factory.py) selects the provider at runtime based on environment variables. Switch between cloud APIs and SageMaker endpoints without changing pipeline code:

Wrap your own model

To add a new speech model deployed on Amazon SageMaker AI with bidirectional streaming, follow this pattern:

  • Identify the container’s protocol. Determine the route path (for example, /v1/synthesize), query parameters, and message format your model container expects (JSON commands, binary audio, or both). Your container must accept WebSocket connections on port 8080 at the /invocations-bidirectional-stream path, which is where Amazon SageMaker AI forwards the bridged stream.
  • Extend the base service class. Subclass TTSService or STTService and implement the required methods: For TTS, run_tts(text) sends text, yields TTSAudioRawFrame chunks, and handles flush/completion signals. For STT, run_stt(audio) sends audio chunks, and a response processor emits TranscriptionFrame when transcriptions arrive.
  • Configure the BiDi client. Set model_invocation_path and model_query_string to match your container’s routing. The SageMakerBidiClient handles authentication and HTTP/2 framing.
  • Handle lifecycle events. Implement _connect() to start the session and launch background tasks, _disconnect() for graceful teardown, and handle_interruption() if your model supports mid-stream cancellation.

Any model container that exposes a bidirectional streaming interface on Amazon SageMaker AI can be wrapped using this pattern. This includes third-party marketplace models and custom models you have deployed with vLLM or Triton. To contribute a new service wrapper, fork the Pipecat repository, implement your service class following the preceding pattern, and open a pull request. You can use an AI coding agent such as Claude Code or Kiro to scaffold the implementation from the existing Deepgram wrappers as a reference.

What this article says

Something is unclear? Ask about the article — I will explain in plain words.

Do not want to dig deeper? We will sort it out for you.