Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

You stream responses token-by-token, but a moderation check on the final output forces you to buffer everything. How do you stream safely without holding back the whole response?


Release one sentence behind generationSentence 1:rules pass, sentSentence 2:rules pass, sentSentence 3:generatingClassifier:checking 1 to 2nulluser seestext by nowruns in parallelA late flag sends a retract event to the client.
Users wait for one sentence instead of the whole answer, at the price of an occasional retraction the client must handle.

What you need to know

The tension: streaming shows words as they are generated, but a moderation check wants to see the whole answer. The compromise is to hold back a small window, not the whole response.

The release loop

Python
async def safe_stream(stream, send, send_event):    buf = ""    async for delta in stream:        buf += delta        while (seg := take_complete_sentence(buf)) is not None:            buf = buf[len(seg):]            if not quick_rules_ok(seg):              # regex: PII, secrets, banned terms                await send_event("blocked"); return            await send(seg)    if buf and quick_rules_ok(buf):        await send(buf)                              # flush the final fragment

Time to first visible text is the time to generate the first sentence, usually well under a second or two, instead of the full 5 to 10 seconds.

Layer the checks by speed

LayerSpeedRunsCatches
Input moderationBefore generationOnce per requestMost abuse, before any cost
Rule checks per segmentSub-millisecondEvery sentence, blockingPII patterns, secrets, competitor names, banned terms
Safety classifier (for example Llama Guard or a provider moderation endpoint)Tens to hundreds of msOn the growing text, in parallelHarmful content that rules miss
Full-response checkAfter completionOnceAnything that needs the whole answer

The classifier runs alongside the stream rather than blocking each sentence. It usually finishes before the response does.

The trade-off you must accept

With progressive release, a sentence can reach the user before the classifier flags it. So the client must support a retraction event: remove the displayed text and show a safety message.

  1. Moderate input — block or redirect before generating.
  2. Stream by sentence — rules check each sentence before release.
  3. Classify in parallel — on the accumulating text.
  4. Retract if needed — the client replaces the message with a notice; the event is logged for review.

For some products, this trade-off is not acceptable. A medical or financial advice feature may buffer fully and show a "checking the answer" indicator. That is a product decision, and saying so in the interview shows judgement.

Metrics

Time to first token, retraction rate, and false-block rate on a labelled set of safe and unsafe outputs. A rising retraction rate means the input check or the prompt needs work.

A real-life example

Scenario (illustrative numbers). An edtech company's homework helper buffers every answer for a full moderation check before showing it. Answers take 7 to 9 seconds to appear, and students leave. The company must still block unsafe content for minors.

They switch to sentence-level release with rule checks, a classifier running in parallel, and strict input moderation. Time to first text drops to 1.1 seconds. Over a month of 3 million answers, 0.02% trigger a retraction, almost all caught by the classifier within the first two sentences. For the "health and body" topic, detected at input, they keep full buffering, because the risk is higher and those questions are rare.

Follow-up questions to expect

  • "Why sentences and not tokens?" — Checks on single tokens have no context; a sentence is the smallest unit a rule or classifier can judge reliably.
  • "What if the classifier is slower than generation?" — Hold release when the classifier falls more than a sentence or two behind, or buffer the final sentence until it catches up.
  • "How does the client handle a retraction?" — A server-sent event type such as retract, which the UI handles by replacing the message and logging the id.