Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your streaming chatbot feels slow even though total response time is only 4 seconds. How do you optimize perceived latency instead of raw latency?
What you need to know
What "feels slow" means
| Measure | What the user feels |
|---|---|
| Time to first token (TTFT) | "Is it working?" |
| Longest inter-token gap | "It froze" |
| Pace of text | Bursty text reads as stuttering |
| Time to the useful sentence | "Did it answer me?" |
Total time matters least. A 4-second answer that starts at 300 ms and flows steadily feels faster than a 3-second answer that starts at 2.5 seconds.
Cut TTFT: parallel, not sequential
Many pipelines do pre-work one step at a time. Most of those steps do not depend on each other.
1import asyncio23async def prepare(user, question):4 profile, chunks, safe = await asyncio.gather(5 load_profile(user.id), # 300 ms6 retrieve(question), # 400 ms7 moderate_input(question), # 250 ms8 )9 if not safe:10 raise Rejected()11 return build_prompt(profile, chunks, question)1213async def answer(user, question, background_tasks):14 prompt = await prepare(user, question)15 background_tasks.add_task(log_analytics, user.id, question) # off the critical path16 async for token in llm.stream(prompt):17 yield tokenRun in sequence, the three steps take 950 ms; in parallel, about 400 ms. Analytics and logging run after the response starts, not before. Prefix caching and a shorter system prompt cut prefill time further.
Make the wait and the stream feel good
- Progress within 500 ms — stream stage events: "Searching 3 sources", tool names, document titles. Honest signals, not a spinner.
- Citation cards early — show sources as soon as retrieval ends.
- Answer first — prompt the model to give the direct answer in the first sentence, then the details.
- Steady pace — buffer a little and release text word by word at a steady rate, so bursts and partial words do not flicker.
- No forced buffering — moderate on a sliding window while streaming, and parse partial JSON with a tolerant parser instead of waiting for the whole object.
Stalls in the middle of a stream usually come from a tool call, a slow moderation step, or a server that pauses decoding for another request's long prefill. Log the longest gap per response to find them.
A real-life example
Scenario, numbers made up. A bank's chatbot takes 4 seconds per answer, but users rate it "slow" and 22% abandon before the answer ends. Traces show TTFT of 2.6 seconds: authentication, profile fetch, retrieval and input moderation run one after another, then an analytics write waits for a database.
The team runs profile, retrieval and moderation in parallel and moves analytics to a background task: TTFT drops to 0.7 seconds. They stream "Checking your account details" at 200 ms and pace text at a steady rate. The longest gap, caused by a mid-answer balance lookup, is fixed by fetching the balance up front. Total time barely changes, at 3.8 seconds, but abandonment falls to 8%.
Follow-up questions to expect
- "Would a faster model fix it?" — Only partly. If TTFT is mostly sequential pre-work, a faster model changes little; measure where the time goes first.
- "What target would you set?" — TTFT p95 under about 800 ms and no gap over roughly a second, then check abandonment, which reflects perception directly.
- "How do you stream structured output?" — Use a parser that accepts incomplete JSON and render fields as they complete.