Course Content
Applied AI Engineering: From Prompt to Production
9 sections · 29 lessons
Speech in and out
About 180 of Harbourline's employees work in facilities: security, housekeeping, drivers and cafeteria staff. Most have no laptop and no company chat account. When they have a question about shift allowances, overtime or leave, they phone the HR helpline, which takes about 300 such calls a month. Many calls are in Hindi, Marathi or a mix of Hindi and English, and most are questions PolicyPal already answers well in text.
A voice version of PolicyPal meets them where they are: a phone number they call, a question asked in their own words, a short answer read aloud and the policy link sent by SMS. Everything behind the voice is the pipeline you have already built. What changes is the input, the output and, most of all, the timing. In a chat window, 4 seconds with streaming text feels fine. On a phone, 4 seconds of silence feels like the line has dropped.
This lesson builds the voice path, sets its latency budget and deals with the part that is harder than it looks: turning real speech, in real accents and mixed languages, into text the rest of the system can use.
The voice pipeline
- Telephony — the call arrives as an audio stream, usually 8 kHz on a phone line.
- End-of-speech detection — voice activity detection decides when the caller has finished speaking.
- Speech-to-text — the audio becomes text, with the detected language.
- The existing pipeline — router, retrieval, answer and checks, with a voice-style prompt.
- Text-to-speech — the answer becomes audio, one sentence at a time.
- Follow-up by SMS — the policy link and any citations are sent as text.
Here is the latency budget for one turn, measured from the moment the caller stops speaking.
| Step | Median time |
|---|---|
| End-of-speech detection (silence threshold) | 500 ms |
| Speech-to-text, a 6-second question | 350 ms |
| Route, rewrite and retrieve | 500 ms |
| Model time to first token | 700 ms |
| First sentence complete | 600 ms |
| Text-to-speech, first audio | 250 ms |
| Silence before the answer starts | about 2.9 s |
The trick in the last three rows is that PolicyPal never waits for the whole answer. As soon as the first sentence is complete, it goes to text-to-speech and starts playing, while the model keeps writing. The rest of the answer arrives faster than it can be spoken. The biggest single lever in the table is the first row: a shorter silence threshold answers sooner but cuts off callers who pause mid-sentence. At 500 milliseconds, about 3% of turns were cut off, mostly older callers; the team added a short "Sorry, please go on" recovery rather than lengthening the threshold for everyone.
Speech-to-text with Whisper
Harbourline runs speech-to-text itself, with an open Whisper model through the faster-whisper library, for the same reason it runs the router itself: calls contain voices, names and health details, and a local model keeps the audio inside the company.
1# policypal/speech.py2from faster_whisper import WhisperModel34stt = WhisperModel("large-v3", device="cuda", compute_type="float16")56VOCAB = ("Harbourline, PolicyPal, earned leave, casual leave, sick leave, "7 "shift allowance, overtime, HRMS, CHRO, Pune, Bengaluru")89def transcribe(path: str) -> tuple[str, str, float]:10 segments, info = stt.transcribe(path, beam_size=5, vad_filter=True,11 initial_prompt=VOCAB)12 text = " ".join(seg.text.strip() for seg in segments) # segments is a generator13 return text, info.language, info.language_probabilityinitial_prompt is a small but valuable detail. It gives the model a preview of words it should expect, and it fixed the most common transcription error in testing: "casual leave" heard as "causal leave". vad_filter skips silence, which speeds things up and reduces invented text during long pauses. Note that segments is a generator: the transcription runs as you iterate over it.
Measuring on your own audio
Published accuracy numbers come from clean recordings. Harbourline's calls are 8 kHz phone audio from noisy car parks and kitchens. The only number that matters is measured on your own calls, and the standard metric is word error rate (WER): the number of word substitutions, deletions and insertions needed to turn the transcript into the correct text, divided by the number of words in the correct text.
1import jiwer23refs = [c["human_transcript"] for c in calls] # 120 calls, typed by a person4hyps = [transcribe(c["audio_path"])[0] for c in calls]5print("WER:", round(jiwer.wer(refs, hyps), 3))On 120 recorded test calls, with consent, WER was 9% for calls in English, 14% for calls mixing Hindi and English, and 23% for the Marathi calls. A smaller Whisper model was three times faster but raised the mixed-language WER to 21%, and the team kept the large model on a GPU. They also measured something more useful than WER: task success, whether the caller got the right answer. It was 81% overall. A 14% WER sounds high, but most errors were small words that did not change the meaning. The errors that mattered were numbers and leave types, which is why the vocabulary prompt and the next step exist.
Mixed languages: rewrite, retrieve, reply in kind
Many questions arrive like this: "Mera earned leave kitna carry forward hoga next year?" (roughly "How much of my earned leave will carry forward next year?"). The policies are in English, and BM25 needs English words to match. So the voice path adds one step before retrieval: a small model call that rewrites the transcript into a short English search query, "earned leave carry forward limit", and records the caller's language. Retrieval runs on the English query. The answer model receives the original transcript as the question and is told to reply in the caller's language. This rewrite adds about 300 milliseconds, already inside the budget above, and raised task success on mixed-language calls from 64% to 79%.
Speaking the answer
An answer written for a screen is wrong for a phone. PolicyPal uses a voice variant of its prompt, answer-voice-v1, with three changes: answers under 60 words, no lists, and a closing offer to send the policy link by SMS. The text still goes through a small cleaning step before text-to-speech, because some things that read well sound terrible.
1import re23def for_speech(answer: str) -> list[str]:4 text = re.sub(r"\s*\[\d+\]", "", answer) # citations go by SMS5 text = re.sub(r"₹\s?([\d,]+)", lambda m: m.group(1).replace(",", "") + " rupees", text)6 text = text.replace("e.g.", "for example")7 return [s for s in re.split(r"(?<=[.!?])\s+", text) if s] # one sentence at a time"₹1,800" becomes "1800 rupees", which a speech engine reads as a number followed by the currency. Citation markers are removed, because "open bracket two close bracket" helps nobody; the citations arrive in the SMS instead. The sentences are returned separately so the first one can be synthesised and played while the rest are still being written. Text-to-speech itself sits behind a small Speaker interface with one method, synthesize(text, language) -> bytes, so the team could compare a hosted voice API with a self-hosted open model on voice quality for Hindi before choosing.
One Speaker: a hosted voice
Harbourline chose a hosted service, Google Cloud Text-to-Speech, for the first version. It has Indian English, Hindi and Marathi voices, and it can return audio in the format a phone line uses, so nothing has to be converted on the way out.
1# policypal/tts.py2from typing import Protocol34from google.cloud import texttospeech56class Speaker(Protocol):7 def synthesize(self, text: str, language: str) -> bytes: ...89LOCALES = {"en": "en-IN", "hi": "hi-IN", "mr": "mr-IN"} # Whisper's code -> voice locale1011class GoogleSpeaker:12 """Hosted text-to-speech that returns 8 kHz mu-law audio for a phone line."""1314 def __init__(self, speaking_rate: float = 0.9, timeout_s: float = 2.0) -> None:15 self._client = texttospeech.TextToSpeechClient() # credentials from the environment16 self._timeout_s = timeout_s17 self._audio = texttospeech.AudioConfig(18 audio_encoding=texttospeech.AudioEncoding.MULAW,19 sample_rate_hertz=8000,20 speaking_rate=speaking_rate)2122 def synthesize(self, text: str, language: str) -> bytes:23 voice = texttospeech.VoiceSelectionParams(language_code=LOCALES.get(language, "en-IN"))24 response = self._client.synthesize_speech(25 input=texttospeech.SynthesisInput(text=text), voice=voice,26 audio_config=self._audio, timeout=self._timeout_s)27 return response.audio_content # a WAV header, then mu-law samplesLOCALES maps the language Whisper detected to a voice locale, so a caller who asked in Marathi hears the answer in a Marathi voice. MULAW at 8,000 Hz is the standard phone-line format; the service returns it inside a WAV file, which many telephony platforms can play as it is. If yours streams raw samples, strip the header first. speaking_rate=0.9 slows the voice a little, so callers can catch numbers. The 2-second timeout matters because a sentence that has not arrived by then is already a long silence on the line; on timeout, the call plays a recorded "one moment, please" and sends the full answer by SMS.
Four trade-offs decided hosted over self-hosted, and each could go the other way for you.
- Latency. One sentence comes back in about 250 ms, which is the row in the budget above. That is a network round trip, so a slow day at the provider becomes silence. A self-hosted model next to the telephony server avoids the network but needs a GPU to be as fast.
- Voice quality. In a blind test with 15 facilities staff, the hosted Hindi voice was preferred for most answers. The open model the team tried sounded natural in English but stumbled on Marathi words and on numbers.
- Cost. Hosted voices are billed per character, at a few dollars to a few tens of dollars per million characters depending on the voice tier; check the current price list. A 60-word answer is about 350 characters, so the voice line's few hundred calls a month cost a few dollars. A self-hosted voice could share the Whisper GPU, but the real cost would be the engineering time to run and update one more model.
- Data leaving the company. Speech-to-text stays in-house because it hears the caller's voice, names and health details. Text-to-speech receives only PolicyPal's answer, which is mostly policy text. The exception is a leave balance, a personal figure, so the privacy team reviewed the provider's data-use terms before approving it.
If your answers carry health or salary details, or your contract forbids sending personal data to a vendor, the same Speaker interface takes a self-hosted engine instead. You pay in quality and operations work, not in code changes.
What the voice path does not do
Voice changes what is safe. A spoken "yes" is easily misheard, so the voice path never creates tickets or takes other actions; it answers and sends links. Checking a personal leave balance requires identity, so the caller must be calling from their registered number and enter a 4-digit PIN on the keypad, which speech-to-text never sees. And any call the router marks sensitive_hr is transferred immediately to a human on the HR line, with no attempt to answer.
Check your understanding
0 of 3 answered
1.Why does PolicyPal start speaking after the first sentence instead of waiting for the full answer?
2.A mixed-language transcript has a 14% word error rate, but callers get the right answer 81% of the time. How can both be true?
3.Why can a caller check their leave balance only after entering a PIN on the keypad, not by saying it?