Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Leadership wants 'ChatGPT for everything' across the company, but most workflows have unclear ROI. How do you identify which business problems are actually worth solving with LLMs?
What you need to know
Where LLMs pay off
LLMs are good at work made of language: reading long documents, drafting, summarising, classifying, and turning one format into another. They pay off where people currently spend many hours on that work. They do not fix a broken process or bad data; they put a fluent surface on the same problem, and the ROI never appears.
A simple value estimate
monthly value = volume x minutes saved per case x achievable success rate x cost per minute - (error rate x cost per error) - (model + review + running costs)Example: 20,000 claims a month x 12 minutes saved x 0.8 success x Rs 8 per minute = Rs 15.4 lakh gross - errors: 2% x 20,000 x Rs 500 to fix = Rs 2 lakh - model, review time, hosting = Rs 3 lakh = about Rs 10 lakh a monthThe numbers will be rough, and that is fine. The point is that each term must exist: a proposal with no volume, no baseline time or no error cost cannot be judged.
The screen
- Language-heavy and expensive? — people spend hours reading or writing for this task today.
- Errors tolerable or checkable? — best if a human already reviews the output, so the model drafts and the reviewer verifies.
- Baseline measurable? — current handle time, error rate, cost per case.
- Feasible? — data accessible, evaluation definable, a named business owner, and a place inside an existing tool where it fits.
- Pick two or three — with a six-week proof and a pre-agreed kill threshold.
The one-page scorecard
| Field | Example |
|---|---|
| Workflow and owner | Motor-claim document summaries; Head of Claims |
| Baseline | 25 minutes per claim; 3% rework |
| Expected lift | 12 minutes saved; rework not worse |
| Cost | Tokens, review time, integration, running |
| Risk tier | Medium: a human approves every claim |
| Proof and kill rule | 6 weeks on 2,000 claims; stop if saving is under 6 minutes |
That turns a slogan into a portfolio of bets, each with evidence.
A real-life example
Scenario, numbers made up. An insurer's leadership collects 41 AI ideas from departments, from "chat with the HR policy" to "auto-approve claims". The AI lead runs each through the screen in two weeks.
Twenty-two fail at step 3: nobody knows current time or error rate. Auto-approving claims fails step 2: an error costs lakhs and nothing catches it. Three survive: summarising claim documents for adjusters, drafting replies to broker emails, and classifying inbound complaints. After six weeks, claim summaries save 11 minutes per claim and go to production; broker drafts save only 2 minutes because brokers rewrite them, and are stopped as agreed. Leadership gets one clear win with numbers, and a cheap, documented "no".
Follow-up questions to expect
- "How do you estimate accuracy before building?" — Run a quick test on 100 past cases with a strong model and a simple prompt; it gives a realistic upper bound in days.
- "When is an LLM the wrong tool?" — When rules or a database query already answer it, when errors are costly and hard to catch, or when the real problem is missing or messy data.
- "Build or buy?" — Buy for common tasks with good products (meeting notes, coding help); build where the workflow, data or integration is specific to the company.