Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your AI feature demos beautifully in meetings but real users barely use it after launch. How do you measure whether an AI feature creates genuine user value versus novelty?


Share of triers still using it, week 1 to week 8100%41%24%16%13%12%12%12%01234567week 4floorformsNovelty keeps falling toward zero; value falls and then flattens.
The number that matters is where the curve stops falling, not how high it started.

What you need to know

Adoption is not value

Trial rate is a vanity number: a new button with a sparkle icon gets clicked. What matters is whether people come back and whether the output gets used.

Novelty

  • High trial in week 1
  • Return rate keeps falling toward zero
  • Outputs viewed, rarely accepted
  • Usage spread thinly across everyone

Value

  • Trial may be modest
  • Return rate falls, then flattens on a floor
  • Outputs accepted with small edits
  • Heavy, steady use in some segments

The measures

MeasureWhat it tells you
Week-2 and week-4 return rate, by cohortNovelty or lasting habit
Acceptance rateDid they use the output?
Edit distance (draft vs what was sent)How much work the output saved
Time to finish the underlying taskReal efficiency, compared with the non-AI path
Holdout comparisonDid the business metric move because of the feature?
SQL
-- week-N return rate among users who first tried the feature in each weekly cohortWITH first_use AS (  SELECT user_id, date_trunc('week', min(used_at)) AS cohort  FROM feature_events GROUP BY user_id)SELECT f.cohort,       count(DISTINCT f.user_id)                                    AS triers,       count(DISTINCT e.user_id) FILTER (WHERE e.used_at >= f.cohort + interval '28 days'                                          AND e.used_at <  f.cohort + interval '35 days')         * 1.0 / count(DISTINCT f.user_id)                          AS week4_returnFROM first_use fLEFT JOIN feature_events e USING (user_id)GROUP BY f.cohort ORDER BY f.cohort;

This gives, for each weekly cohort of first-time users, the share who used the feature again in week 4. Plot several cohorts: if the curves flatten above zero, you have a floor of real users.

Diagnose the gap

Demos succeed on hand-picked inputs. To see where real use fails, instrument each step of the funnel separately:

  1. Impression — did users see the feature? If not, it is discoverability.
  2. Invocation — did they try it? If not, the value is unclear.
  3. Output shown — did it finish in time? If not, latency.
  4. Accepted — did they use the result? If not, quality or trust.

Trust failures have a distinctive look: the output is correct, and the user redoes the work anyway, because checking it costs as much as doing it. Showing sources or a diff often fixes that.

A real-life example

Scenario, numbers made up. A CRM company launches AI-drafted follow-up emails. 60% of active users try it in week 1; by week 4 only 6% use it. Leadership calls it a failure.

Segmenting shows sales reps who send many similar emails keep using it: their week-4 return rate is 38% and flat through week 8, with drafts sent after light edits. Account managers, writing few long emails, abandon it. Five recorded sessions show account managers copying the draft into a document to check facts against the deal record. The team repositions the feature for high-volume senders, adds the deal facts used beside each draft, and a holdout test shows reps with drafts send 22% more follow-ups per week.

Follow-up questions to expect

  • "What is edit distance telling you?" — How much of the draft survived. Small edits mean the output saved work; a full rewrite means it did not, even if the user clicked "accept".
  • "What if you can't run a holdout?" — Compare similar users before and after, or ask "how disappointed would you be if this disappeared?" — weaker evidence, but better than usage counts.
  • "How long before you judge?" — At least four to six weeks of cohorts, because novelty takes that long to wear off.