Applied AI Engineering: From Prompt to Production

Course Content

Applied AI Engineering: From Prompt to Production

9 sections · 29 lessons

Speech in and out


About 180 of Harbourline's employees work in facilities: security, housekeeping, drivers and cafeteria staff. Most have no laptop and no company chat account. When they have a question about shift allowances, overtime or leave, they phone the HR helpline, which takes about 300 such calls a month. Many calls are in Hindi, Marathi or a mix of Hindi and English, and most are questions PolicyPal already answers well in text.

A voice version of PolicyPal meets them where they are: a phone number they call, a question asked in their own words, a short answer read aloud and the policy link sent by SMS. Everything behind the voice is the pipeline you have already built. What changes is the input, the output and, most of all, the timing. In a chat window, 4 seconds with streaming text feels fine. On a phone, 4 seconds of silence feels like the line has dropped.

This lesson builds the voice path, sets its latency budget and deals with the part that is harder than it looks: turning real speech, in real accents and mixed languages, into text the rest of the system can use.

Milliseconds of silence before the answer500350500700600250012345end of speechWhisperrewrite,retrievefirst tokenfirst sentencefirst audioAbout 2.9 s in total; the rest of the answer is written faster than it is spoken.
Speaking the first sentence while the model writes the rest keeps the silence short enough that the line never feels dead.

The voice pipeline

  1. Telephony — the call arrives as an audio stream, usually 8 kHz on a phone line.
  2. End-of-speech detection — voice activity detection decides when the caller has finished speaking.
  3. Speech-to-text — the audio becomes text, with the detected language.
  4. The existing pipeline — router, retrieval, answer and checks, with a voice-style prompt.
  5. Text-to-speech — the answer becomes audio, one sentence at a time.
  6. Follow-up by SMS — the policy link and any citations are sent as text.

Here is the latency budget for one turn, measured from the moment the caller stops speaking.

StepMedian time
End-of-speech detection (silence threshold)500 ms
Speech-to-text, a 6-second question350 ms
Route, rewrite and retrieve500 ms
Model time to first token700 ms
First sentence complete600 ms
Text-to-speech, first audio250 ms
Silence before the answer startsabout 2.9 s

The trick in the last three rows is that PolicyPal never waits for the whole answer. As soon as the first sentence is complete, it goes to text-to-speech and starts playing, while the model keeps writing. The rest of the answer arrives faster than it can be spoken. The biggest single lever in the table is the first row: a shorter silence threshold answers sooner but cuts off callers who pause mid-sentence. At 500 milliseconds, about 3% of turns were cut off, mostly older callers; the team added a short "Sorry, please go on" recovery rather than lengthening the threshold for everyone.

Speech-to-text with Whisper

Harbourline runs speech-to-text itself, with an open Whisper model through the faster-whisper library, for the same reason it runs the router itself: calls contain voices, names and health details, and a local model keeps the audio inside the company.

Python
# policypal/speech.pyfrom faster_whisper import WhisperModelstt = WhisperModel("large-v3", device="cuda", compute_type="float16")VOCAB = ("Harbourline, PolicyPal, earned leave, casual leave, sick leave, "         "shift allowance, overtime, HRMS, CHRO, Pune, Bengaluru")def transcribe(path: str) -> tuple[str, str, float]:    segments, info = stt.transcribe(path, beam_size=5, vad_filter=True,                                    initial_prompt=VOCAB)    text = " ".join(seg.text.strip() for seg in segments)   # segments is a generator    return text, info.language, info.language_probability

initial_prompt is a small but valuable detail. It gives the model a preview of words it should expect, and it fixed the most common transcription error in testing: "casual leave" heard as "causal leave". vad_filter skips silence, which speeds things up and reduces invented text during long pauses. Note that segments is a generator: the transcription runs as you iterate over it.

Measuring on your own audio

Published accuracy numbers come from clean recordings. Harbourline's calls are 8 kHz phone audio from noisy car parks and kitchens. The only number that matters is measured on your own calls, and the standard metric is word error rate (WER): the number of word substitutions, deletions and insertions needed to turn the transcript into the correct text, divided by the number of words in the correct text.

Python
import jiwerrefs = [c["human_transcript"] for c in calls]        # 120 calls, typed by a personhyps = [transcribe(c["audio_path"])[0] for c in calls]print("WER:", round(jiwer.wer(refs, hyps), 3))

On 120 recorded test calls, with consent, WER was 9% for calls in English, 14% for calls mixing Hindi and English, and 23% for the Marathi calls. A smaller Whisper model was three times faster but raised the mixed-language WER to 21%, and the team kept the large model on a GPU. They also measured something more useful than WER: task success, whether the caller got the right answer. It was 81% overall. A 14% WER sounds high, but most errors were small words that did not change the meaning. The errors that mattered were numbers and leave types, which is why the vocabulary prompt and the next step exist.

Mixed languages: rewrite, retrieve, reply in kind

Many questions arrive like this: "Mera earned leave kitna carry forward hoga next year?" (roughly "How much of my earned leave will carry forward next year?"). The policies are in English, and BM25 needs English words to match. So the voice path adds one step before retrieval: a small model call that rewrites the transcript into a short English search query, "earned leave carry forward limit", and records the caller's language. Retrieval runs on the English query. The answer model receives the original transcript as the question and is told to reply in the caller's language. This rewrite adds about 300 milliseconds, already inside the budget above, and raised task success on mixed-language calls from 64% to 79%.

Speaking the answer

An answer written for a screen is wrong for a phone. PolicyPal uses a voice variant of its prompt, answer-voice-v1, with three changes: answers under 60 words, no lists, and a closing offer to send the policy link by SMS. The text still goes through a small cleaning step before text-to-speech, because some things that read well sound terrible.

Python
import redef for_speech(answer: str) -> list[str]:    text = re.sub(r"\s*\[\d+\]", "", answer)                      # citations go by SMS    text = re.sub(r"₹\s?([\d,]+)", lambda m: m.group(1).replace(",", "") + " rupees", text)    text = text.replace("e.g.", "for example")    return [s for s in re.split(r"(?<=[.!?])\s+", text) if s]     # one sentence at a time

"₹1,800" becomes "1800 rupees", which a speech engine reads as a number followed by the currency. Citation markers are removed, because "open bracket two close bracket" helps nobody; the citations arrive in the SMS instead. The sentences are returned separately so the first one can be synthesised and played while the rest are still being written. Text-to-speech itself sits behind a small Speaker interface with one method, synthesize(text, language) -> bytes, so the team could compare a hosted voice API with a self-hosted open model on voice quality for Hindi before choosing.

One Speaker: a hosted voice

Harbourline chose a hosted service, Google Cloud Text-to-Speech, for the first version. It has Indian English, Hindi and Marathi voices, and it can return audio in the format a phone line uses, so nothing has to be converted on the way out.

Python
# policypal/tts.pyfrom typing import Protocolfrom google.cloud import texttospeechclass Speaker(Protocol):    def synthesize(self, text: str, language: str) -> bytes: ...LOCALES = {"en": "en-IN", "hi": "hi-IN", "mr": "mr-IN"}    # Whisper's code -> voice localeclass GoogleSpeaker:    """Hosted text-to-speech that returns 8 kHz mu-law audio for a phone line."""    def __init__(self, speaking_rate: float = 0.9, timeout_s: float = 2.0) -> None:        self._client = texttospeech.TextToSpeechClient()   # credentials from the environment        self._timeout_s = timeout_s        self._audio = texttospeech.AudioConfig(            audio_encoding=texttospeech.AudioEncoding.MULAW,            sample_rate_hertz=8000,            speaking_rate=speaking_rate)    def synthesize(self, text: str, language: str) -> bytes:        voice = texttospeech.VoiceSelectionParams(language_code=LOCALES.get(language, "en-IN"))        response = self._client.synthesize_speech(            input=texttospeech.SynthesisInput(text=text), voice=voice,            audio_config=self._audio, timeout=self._timeout_s)        return response.audio_content              # a WAV header, then mu-law samples

LOCALES maps the language Whisper detected to a voice locale, so a caller who asked in Marathi hears the answer in a Marathi voice. MULAW at 8,000 Hz is the standard phone-line format; the service returns it inside a WAV file, which many telephony platforms can play as it is. If yours streams raw samples, strip the header first. speaking_rate=0.9 slows the voice a little, so callers can catch numbers. The 2-second timeout matters because a sentence that has not arrived by then is already a long silence on the line; on timeout, the call plays a recorded "one moment, please" and sends the full answer by SMS.

Four trade-offs decided hosted over self-hosted, and each could go the other way for you.

  • Latency. One sentence comes back in about 250 ms, which is the row in the budget above. That is a network round trip, so a slow day at the provider becomes silence. A self-hosted model next to the telephony server avoids the network but needs a GPU to be as fast.
  • Voice quality. In a blind test with 15 facilities staff, the hosted Hindi voice was preferred for most answers. The open model the team tried sounded natural in English but stumbled on Marathi words and on numbers.
  • Cost. Hosted voices are billed per character, at a few dollars to a few tens of dollars per million characters depending on the voice tier; check the current price list. A 60-word answer is about 350 characters, so the voice line's few hundred calls a month cost a few dollars. A self-hosted voice could share the Whisper GPU, but the real cost would be the engineering time to run and update one more model.
  • Data leaving the company. Speech-to-text stays in-house because it hears the caller's voice, names and health details. Text-to-speech receives only PolicyPal's answer, which is mostly policy text. The exception is a leave balance, a personal figure, so the privacy team reviewed the provider's data-use terms before approving it.

If your answers carry health or salary details, or your contract forbids sending personal data to a vendor, the same Speaker interface takes a self-hosted engine instead. You pay in quality and operations work, not in code changes.

What the voice path does not do

Voice changes what is safe. A spoken "yes" is easily misheard, so the voice path never creates tickets or takes other actions; it answers and sends links. Checking a personal leave balance requires identity, so the caller must be calling from their registered number and enter a 4-digit PIN on the keypad, which speech-to-text never sees. And any call the router marks sensitive_hr is transferred immediately to a human on the HR line, with no attempt to answer.

Check your understanding

0 of 3 answered

1.Why does PolicyPal start speaking after the first sentence instead of waiting for the full answer?

2.A mixed-language transcript has a 14% word error rate, but callers get the right answer 81% of the time. How can both be true?

3.Why can a caller check their leave balance only after entering a PIN on the keypad, not by saying it?