Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Leadership wants 'ChatGPT for everything' across the company, but most workflows have unclear ROI. How do you identify which business problems are actually worth solving with LLMs?


What you need to know

Where LLMs pay off

LLMs are good at work made of language: reading long documents, drafting, summarising, classifying, and turning one format into another. They pay off where people currently spend many hours on that work. They do not fix a broken process or bad data; they put a fluent surface on the same problem, and the ROI never appears.

A simple value estimate

Text
monthly value = volume x minutes saved per case x achievable success rate x cost per minute              - (error rate x cost per error)              - (model + review + running costs)Example: 20,000 claims a month x 12 minutes saved x 0.8 success x Rs 8 per minute       = Rs 15.4 lakh gross       - errors: 2% x 20,000 x Rs 500 to fix = Rs 2 lakh       - model, review time, hosting    = Rs 3 lakh       = about Rs 10 lakh a month

The numbers will be rough, and that is fine. The point is that each term must exist: a proposal with no volume, no baseline time or no error cost cannot be judged.

The screen

  1. Language-heavy and expensive? — people spend hours reading or writing for this task today.
  2. Errors tolerable or checkable? — best if a human already reviews the output, so the model drafts and the reviewer verifies.
  3. Baseline measurable? — current handle time, error rate, cost per case.
  4. Feasible? — data accessible, evaluation definable, a named business owner, and a place inside an existing tool where it fits.
  5. Pick two or three — with a six-week proof and a pre-agreed kill threshold.

The one-page scorecard

FieldExample
Workflow and ownerMotor-claim document summaries; Head of Claims
Baseline25 minutes per claim; 3% rework
Expected lift12 minutes saved; rework not worse
CostTokens, review time, integration, running
Risk tierMedium: a human approves every claim
Proof and kill rule6 weeks on 2,000 claims; stop if saving is under 6 minutes

That turns a slogan into a portfolio of bets, each with evidence.

A real-life example

Scenario, numbers made up. An insurer's leadership collects 41 AI ideas from departments, from "chat with the HR policy" to "auto-approve claims". The AI lead runs each through the screen in two weeks.

Twenty-two fail at step 3: nobody knows current time or error rate. Auto-approving claims fails step 2: an error costs lakhs and nothing catches it. Three survive: summarising claim documents for adjusters, drafting replies to broker emails, and classifying inbound complaints. After six weeks, claim summaries save 11 minutes per claim and go to production; broker drafts save only 2 minutes because brokers rewrite them, and are stopped as agreed. Leadership gets one clear win with numbers, and a cheap, documented "no".

Follow-up questions to expect

  • "How do you estimate accuracy before building?" — Run a quick test on 100 past cases with a strong model and a simple prompt; it gives a realistic upper bound in days.
  • "When is an LLM the wrong tool?" — When rules or a database query already answer it, when errors are costly and hard to catch, or when the real problem is missing or messy data.
  • "Build or buy?" — Buy for common tasks with good products (meeting notes, coding help); build where the workflow, data or integration is specific to the company.