Getting Speaker‑Labeled Transcripts on SageMaker with WhisperX

Getting Speaker‑Labeled Transcripts on SageMaker with WhisperX

According to Artificial Intelligence, AWS now ships a WhisperX Deep Learning Container (DLC) that adds per‑word timestamps and speaker diarization to OpenAI’s Whisper model. The container runs on SageMaker AI endpoints, letting teams turn raw audio—from call‑center recordings to meeting minutes—into searchable, caption‑ready text without building a custom Docker image.

What WhisperX actually does

Whisper is an open‑source automatic speech recognition (ASR) model that produces transcriptions with phrase‑level timestamps. WhisperX builds on Whisper in three ways that matter for production workloads:

  • Forced alignment using wav2vec2 gives timestamps for every word instead of just each segment.
  • Speaker diarization tags each spoken segment with a speaker identifier (e.g., SPEAKER_01).
  • Batched inference speeds up processing by handling multiple audio chunks in one GPU pass.

The result is a JSON (or SRT/VTT) payload that contains the transcript, word‑level timing, and speaker labels—all in a single request. In the AWS blog the authors demonstrate the model on a noisy, multi‑party air‑traffic‑control recording, showing that even with background radio noise the system can separate speakers and align words to the clock.

Real‑time vs. asynchronous endpoints on SageMaker

SageMaker AI supports two serving patterns. Which one you choose depends mostly on clip length and whether you need an immediate response.

Dimension Real‑time endpoint Asynchronous endpoint
Best for Short, interactive clips Long audio, high‑volume batch
Latency Must finish < 60 s (synchronous) Submit‑and‑poll, no response cap
Invocation InvokeEndpoint (inline body) InvokeEndpointAsync (S3 reference)
I/O Body in request / response Input + output in Amazon S3
Scaling Add instances (one request per container) Add instances, can autoscale to zero
Cost profile Bills while endpoint is up Scale‑to‑zero when idle saves cost

The real‑time pattern runs the whole WhisperX pipeline inside the 60‑second request window. It’s handy for quick voice commands or short call snippets. The asynchronous pattern uploads the audio to S3, triggers the container, and writes the result back to S3, allowing any length of audio to be processed without the 60‑second limit.

Deploying WhisperX on SageMaker

The DLC image (e.g., 763104351884.dkr.ecr.<region>.amazonaws.com/whisperx:3.8.6-cu128-amzn2023-sagemaker) already contains Whisper, the alignment model, and diarization weights, so you only need a GPU‑enabled instance type such as ml.g4dn.xlarge for cost‑effective testing or ml.g5.2xlarge for higher throughput. Two configuration details are easy to miss:

  1. GPU AMI pin – set InferenceAmiVersion to al2-ami-sagemaker-inference-gpu-3-1. Without this the container fails to start.
  2. Health‑check timeout – WhisperX loads weights lazily; a 900‑second startup timeout prevents premature failures.

A minimal deployment script registers the model, creates an endpoint config, and launches the endpoint. The same image works for both serving patterns; you only change the invocation API.

What the trade‑offs really mean for everyday teams

The headline claim is that WhisperX gives “production‑ready” speaker‑labeled transcripts. In practice the trade‑off is between latency and cost. A real‑time endpoint keeps an instance warm, so you pay for the whole hour even if you send a single 10‑second clip. That may be acceptable for a call‑center dashboard that needs sub‑second turn‑around, but it can become expensive if usage is sporadic.

An asynchronous endpoint lets you spin the GPU down when there are no jobs. SageMaker can even scale to zero, meaning you only pay for the minutes the container actually runs. The downside is added complexity: you must manage S3 buckets, poll for results, and handle failure locations. For teams that already store recordings in S3 (most do), the extra step is minor; for ad‑hoc users it can feel heavyweight.

Another subtle point is accuracy vs. noise. The ATC example shows WhisperX handling compression and overlapping speech, but the output still contains occasional mis‑recognitions (e.g., “stoppy to park”). If your compliance regime requires near‑perfect transcripts, you’ll need a human‑in‑the‑loop review step regardless of the model.

Finally, vendor lock‑in: the container runs on SageMaker and uses an Amazon‑specific GPU AMI. Moving the same image to another cloud would require rebuilding the container with the appropriate serving contract. Teams that are already on AWS will find the integration smooth; others may need to weigh the effort of migrating.

Quick start you can run today

  1. Create a SageMaker notebook (or use SageMaker Studio) and clone the AWS Samples repository linked in the blog.
  2. Pick an instance – start with ml.g4dn.xlarge for a cheap test.
  3. Deploy the async endpoint using the sample Python snippet (set InferenceAmiVersion as described).
  4. Upload a short audio file (e.g., a 30‑second meeting excerpt) to an S3 bucket whose name contains sagemaker.
  5. Invoke the endpoint asynchronously with diarize=true and timestamp_granularities[]=word.
  6. Download the JSON or SRT output from the provided S3 location and verify the speaker labels.

If the result meets your needs, you can scale the endpoint pool, adjust the instance type for higher throughput, or switch to a real‑time endpoint for ultra‑short clips. The same notebook also shows how to switch the response_format to srt for direct caption generation.

Sources

Read next

We count page views without cookies — no identifier, nothing stored on your device. Accept to allow cookies for analytics.