Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI assistant must reason over constantly changing real-time data like stock prices, deliveries, and IoT events. How do you combine streaming systems with LLM reasoning reliably?
What you need to know
Split the work
Stream layer
- Handles thousands of events per second
- Maintains current state per entity
- Deterministic rules fire alerts
- Handles late and out-of-order events
LLM layer
- Explains, summarises, answers questions
- Reads state through tools when asked
- Drafts alerts for humans
- Never blocks the stream
Fresh reads with an as-of time
Anything time-sensitive is fetched at answer time, never baked into the prompt. A prompt built 90 seconds ago already has a stale price.
1from datetime import datetime, timezone, timedelta23MAX_AGE = {"price": timedelta(seconds=15), "shipment": timedelta(minutes=5)}45def get_shipment(shipment_id: str) -> dict:6 row = state_store.get(f"shipment:{shipment_id}") # kept current by the stream job7 age = datetime.now(timezone.utc) - row["updated_at"]8 if age > MAX_AGE["shipment"]:9 return {"error": "stale", "as_of": row["updated_at"].isoformat(),10 "hint": "Tell the user the latest data is delayed; do not guess."}11 return {"status": row["status"], "eta": row["eta"], "as_of": row["updated_at"].isoformat()}The tool returns the value with its as-of time, and refuses when data is older than the allowed age. The system prompt tells the model to state the as-of time in the answer.
The event-triggered path
- Stream job detects a condition — a freezer above 4 °C for 5 minutes, a delivery 20 minutes late.
- Deduplicate — one open incident per device or order, not one per event.
- Rate-limit the LLM lane — a flapping sensor must not create 500 calls.
- LLM drafts the explanation or message, reading state through tools.
- Execute by rule or human — the model proposes; deterministic code or a person acts on anything financial or safety-critical.
If the LLM lane falls behind, it sheds load and drops to template messages; it must never slow the stream.
Measure the p95 age of the oldest fact used in an answer, alert precision and recall, and cost per event.
A real-life example
Scenario, numbers made up. A quick-commerce company's assistant answers "Where is my order?" and alerts store managers about freezer problems. The first version puts a snapshot of order status in the prompt when the chat opens. Customers who chat for a few minutes get ETAs up to 8 minutes out of date, and one faulty freezer sensor flapping around its threshold triggers 600 LLM alerts in an hour.
The team moves order data behind a get_shipment tool with a 5-minute max age, adds "as of 7:42 pm" to answers, and deduplicates alerts to one open incident per device with a 10-minute cooldown. Complaints about wrong ETAs fall by 70%, and alert volume from that store drops to 3 incidents a day, each with a useful LLM-written summary.
Follow-up questions to expect
- "Why not stream events straight into the model?" — It is costly, slow, and models are poor at precise arithmetic over long sequences of numbers. Aggregate in the stream layer; give the model conclusions and current state.
- "What if data is stale when the user asks?" — Say so with the as-of time and do not guess; staleness is a normal answer, not an error to hide.
- "How do you handle late events?" — Watermarks and allowed lateness in the stream job; the state store records event time, not just arrival time.