Speech recognition with speaker labels (WhisperX)#

This procedure requires access to the installation directory, server settings and Docker. It is necessary to save the previous Whisper parameter values for rollback. In Chat, the transcription Model assigned to the Account takes priority. If its active Instance is used, it must point to the local Whisper service and request large-v2. Changing .env does not automatically redirect an external Instance to the local service.

This guide explains how to enable local audio and video transcription with speaker separation (diarization) in Sherpa AI Server. Processing runs entirely on your server, without external services or internet access during operation.

Recognized speech segments with an identified speaker receive a speaker label (speaker_1, speaker_2, and so on) and timestamps. Labels start at 1 in the order speakers first appear in the recording. They are included in the API result and the LLM context as lines such as [0:00:05] speaker_1: Привет. The Chat interface has no separate speaker screen.

Requirements#

Before enabling this method, make sure that:

  • Sherpa AI Server is installed and the aiserver-whisper service runs with the whisper or full profile;
  • Whisper runs on an NVIDIA GPU (see Switching Whisper from CPU to GPU); this method is very slow on a CPU;
  • you have the model-whisperx-large-v2.tar.gz model archive (about 3 GB) and enough disk space;
  • your aiserver-whisper image version supports the whisperx method.

The large-v2.pt file used by the standard base method is not suitable for WhisperX: WhisperX uses a different model format. You do not need to delete the existing model files.

Step 1. It is necessary to extract the model files#

From the installation directory, run:

tar -xzvf model-whisperx-large-v2.tar.gz -C ./whisper/models
ls ./whisper/models/whisperx

The ./whisper/models/whisperx directory should appear alongside the existing model files.

Step 2. Enable the method in .env#

It is necessary to open .env and set the following values in the # !whisper block:

WHISPER_METHOD=whisperx
WHISPER_MODEL=large-v2
WHISPER_DEFAULT_MODEL=large-v2
WHISPERX_DIARIZE=true

Additional parameters:

Parameter Purpose
WHISPERX_DIARIZE true — separate speakers (default); false — transcription with timestamps only
WHISPERX_NUM_SPEAKERS The exact number of speakers, if known. It is necessary to leave empty to detect it automatically
WHISPERX_BATCH_SIZE Reduce this value if you receive a CUDA out of memory error

Step 3. Apply the server and Whisper settings#

It is necessary to recreate both the server and Whisper containers to apply the new .env settings, including WHISPER_DEFAULT_MODEL on the server. The command uses the images already available locally, without rebuilding or pulling them. It is necessary to schedule the server restart for a suitable time:

docker compose --profile whisper up -d --no-build --pull never --force-recreate aiserver aiserver-whisper

Step 4. It is necessary to check operation#

curl -sS http://127.0.0.1:3005/health

The response should contain "method":"whisperx" and "diarization_enabled":true.

This response confirms the service configuration. It does not prove that weights loaded or transcription succeeded: weights load on the first request. Verify the result by processing audio through the selected local route.

Then attach an audio or video recording of two people talking in the Chat and send a message. The transcription text passed to the LLM will contain speaker_1 and speaker_2 labels with timestamps.

A short recording can be submitted directly to the local Whisper API to inspect the result. The shell variable WHISPER_API_KEY must contain this service's valid key when key verification is enabled. If the service has no key configured, the example does not require the Authorization header.

curl -fsS http://127.0.0.1:3005/v1/audio/transcriptions \
  -H "Authorization: Bearer $WHISPER_API_KEY" \
  -F "file=@./meeting.wav" \
  -F "model=large-v2" \
  -F "response_format=verbose_json"

The response's segments array must be checked for recognized text, start and end timestamps, and speaker for segments whose speaker was identified. A missing speaker does not mean the segment's author was identified. A failed request requires checking the weights, GPU availability and the service message. A successful /health alone does not confirm this result.

Limitations#

  • Processing takes longer than standard transcription. It is necessary to check request timeouts for long recordings.
  • Speakers may be identified incorrectly when their voices overlap significantly.
  • The method does not determine real names: it assigns speaker_N labels only.
  • The prompt and temperature parameters are not used by this method.

Rollback#

To roll back, restore the saved previous Whisper parameter values in .env and repeat step 3 for both containers. The previous Model weights must be available.