Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your streaming chatbot feels slow even though total response time is only 4 seconds. How do you optimize perceived latency instead of raw latency?


Gaps between tokens in one answer, in milliseconds4038429004139444001234567stall:mid-answer lookupFetching the balance before generation removed the stall without changing total time.
One 900 ms freeze in a steady stream reads as broken, even when the whole answer finishes on time.

What you need to know

What "feels slow" means

MeasureWhat the user feels
Time to first token (TTFT)"Is it working?"
Longest inter-token gap"It froze"
Pace of textBursty text reads as stuttering
Time to the useful sentence"Did it answer me?"

Total time matters least. A 4-second answer that starts at 300 ms and flows steadily feels faster than a 3-second answer that starts at 2.5 seconds.

Cut TTFT: parallel, not sequential

Many pipelines do pre-work one step at a time. Most of those steps do not depend on each other.

Python
import asyncioasync def prepare(user, question):    profile, chunks, safe = await asyncio.gather(        load_profile(user.id),          # 300 ms        retrieve(question),             # 400 ms        moderate_input(question),       # 250 ms    )    if not safe:        raise Rejected()    return build_prompt(profile, chunks, question)async def answer(user, question, background_tasks):    prompt = await prepare(user, question)    background_tasks.add_task(log_analytics, user.id, question)   # off the critical path    async for token in llm.stream(prompt):        yield token

Run in sequence, the three steps take 950 ms; in parallel, about 400 ms. Analytics and logging run after the response starts, not before. Prefix caching and a shorter system prompt cut prefill time further.

Make the wait and the stream feel good

  1. Progress within 500 ms — stream stage events: "Searching 3 sources", tool names, document titles. Honest signals, not a spinner.
  2. Citation cards early — show sources as soon as retrieval ends.
  3. Answer first — prompt the model to give the direct answer in the first sentence, then the details.
  4. Steady pace — buffer a little and release text word by word at a steady rate, so bursts and partial words do not flicker.
  5. No forced buffering — moderate on a sliding window while streaming, and parse partial JSON with a tolerant parser instead of waiting for the whole object.

Stalls in the middle of a stream usually come from a tool call, a slow moderation step, or a server that pauses decoding for another request's long prefill. Log the longest gap per response to find them.

A real-life example

Scenario, numbers made up. A bank's chatbot takes 4 seconds per answer, but users rate it "slow" and 22% abandon before the answer ends. Traces show TTFT of 2.6 seconds: authentication, profile fetch, retrieval and input moderation run one after another, then an analytics write waits for a database.

The team runs profile, retrieval and moderation in parallel and moves analytics to a background task: TTFT drops to 0.7 seconds. They stream "Checking your account details" at 200 ms and pace text at a steady rate. The longest gap, caused by a mid-answer balance lookup, is fixed by fetching the balance up front. Total time barely changes, at 3.8 seconds, but abandonment falls to 8%.

Follow-up questions to expect

  • "Would a faster model fix it?" — Only partly. If TTFT is mostly sequential pre-work, a faster model changes little; measure where the time goes first.
  • "What target would you set?" — TTFT p95 under about 800 ms and no gap over roughly a second, then check abandonment, which reflects perception directly.
  • "How do you stream structured output?" — Use a parser that accepts incomplete JSON and render fields as they complete.