Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You stream responses token-by-token, but a moderation check on the final output forces you to buffer everything. How do you stream safely without holding back the whole response?
What you need to know
The tension: streaming shows words as they are generated, but a moderation check wants to see the whole answer. The compromise is to hold back a small window, not the whole response.
The release loop
1async def safe_stream(stream, send, send_event):2 buf = ""3 async for delta in stream:4 buf += delta5 while (seg := take_complete_sentence(buf)) is not None:6 buf = buf[len(seg):]7 if not quick_rules_ok(seg): # regex: PII, secrets, banned terms8 await send_event("blocked"); return9 await send(seg)10 if buf and quick_rules_ok(buf):11 await send(buf) # flush the final fragmentTime to first visible text is the time to generate the first sentence, usually well under a second or two, instead of the full 5 to 10 seconds.
Layer the checks by speed
| Layer | Speed | Runs | Catches |
|---|---|---|---|
| Input moderation | Before generation | Once per request | Most abuse, before any cost |
| Rule checks per segment | Sub-millisecond | Every sentence, blocking | PII patterns, secrets, competitor names, banned terms |
| Safety classifier (for example Llama Guard or a provider moderation endpoint) | Tens to hundreds of ms | On the growing text, in parallel | Harmful content that rules miss |
| Full-response check | After completion | Once | Anything that needs the whole answer |
The classifier runs alongside the stream rather than blocking each sentence. It usually finishes before the response does.
The trade-off you must accept
With progressive release, a sentence can reach the user before the classifier flags it. So the client must support a retraction event: remove the displayed text and show a safety message.
- Moderate input — block or redirect before generating.
- Stream by sentence — rules check each sentence before release.
- Classify in parallel — on the accumulating text.
- Retract if needed — the client replaces the message with a notice; the event is logged for review.
For some products, this trade-off is not acceptable. A medical or financial advice feature may buffer fully and show a "checking the answer" indicator. That is a product decision, and saying so in the interview shows judgement.
Metrics
Time to first token, retraction rate, and false-block rate on a labelled set of safe and unsafe outputs. A rising retraction rate means the input check or the prompt needs work.
A real-life example
Scenario (illustrative numbers). An edtech company's homework helper buffers every answer for a full moderation check before showing it. Answers take 7 to 9 seconds to appear, and students leave. The company must still block unsafe content for minors.
They switch to sentence-level release with rule checks, a classifier running in parallel, and strict input moderation. Time to first text drops to 1.1 seconds. Over a month of 3 million answers, 0.02% trigger a retraction, almost all caught by the classifier within the first two sentences. For the "health and body" topic, detected at input, they keep full buffering, because the risk is higher and those questions are rare.
Follow-up questions to expect
- "Why sentences and not tokens?" — Checks on single tokens have no context; a sentence is the smallest unit a rule or classifier can judge reliably.
- "What if the classifier is slower than generation?" — Hold release when the classifier falls more than a sentence or two behind, or buffer the final sentence until it catches up.
- "How does the client handle a retraction?" — A server-sent event type such as
retract, which the UI handles by replacing the message and logging the id.