Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your AI feature demos beautifully in meetings but real users barely use it after launch. How do you measure whether an AI feature creates genuine user value versus novelty?
What you need to know
Adoption is not value
Trial rate is a vanity number: a new button with a sparkle icon gets clicked. What matters is whether people come back and whether the output gets used.
Novelty
- High trial in week 1
- Return rate keeps falling toward zero
- Outputs viewed, rarely accepted
- Usage spread thinly across everyone
Value
- Trial may be modest
- Return rate falls, then flattens on a floor
- Outputs accepted with small edits
- Heavy, steady use in some segments
The measures
| Measure | What it tells you |
|---|---|
| Week-2 and week-4 return rate, by cohort | Novelty or lasting habit |
| Acceptance rate | Did they use the output? |
| Edit distance (draft vs what was sent) | How much work the output saved |
| Time to finish the underlying task | Real efficiency, compared with the non-AI path |
| Holdout comparison | Did the business metric move because of the feature? |
1-- week-N return rate among users who first tried the feature in each weekly cohort2WITH first_use AS (3 SELECT user_id, date_trunc('week', min(used_at)) AS cohort4 FROM feature_events GROUP BY user_id5)6SELECT f.cohort,7 count(DISTINCT f.user_id) AS triers,8 count(DISTINCT e.user_id) FILTER (WHERE e.used_at >= f.cohort + interval '28 days'9 AND e.used_at < f.cohort + interval '35 days')10 * 1.0 / count(DISTINCT f.user_id) AS week4_return11FROM first_use f12LEFT JOIN feature_events e USING (user_id)13GROUP BY f.cohort ORDER BY f.cohort;This gives, for each weekly cohort of first-time users, the share who used the feature again in week 4. Plot several cohorts: if the curves flatten above zero, you have a floor of real users.
Diagnose the gap
Demos succeed on hand-picked inputs. To see where real use fails, instrument each step of the funnel separately:
- Impression — did users see the feature? If not, it is discoverability.
- Invocation — did they try it? If not, the value is unclear.
- Output shown — did it finish in time? If not, latency.
- Accepted — did they use the result? If not, quality or trust.
Trust failures have a distinctive look: the output is correct, and the user redoes the work anyway, because checking it costs as much as doing it. Showing sources or a diff often fixes that.
A real-life example
Scenario, numbers made up. A CRM company launches AI-drafted follow-up emails. 60% of active users try it in week 1; by week 4 only 6% use it. Leadership calls it a failure.
Segmenting shows sales reps who send many similar emails keep using it: their week-4 return rate is 38% and flat through week 8, with drafts sent after light edits. Account managers, writing few long emails, abandon it. Five recorded sessions show account managers copying the draft into a document to check facts against the deal record. The team repositions the feature for high-volume senders, adds the deal facts used beside each draft, and a holdout test shows reps with drafts send 22% more follow-ups per week.
Follow-up questions to expect
- "What is edit distance telling you?" — How much of the draft survived. Small edits mean the output saved work; a full rewrite means it did not, even if the user clicked "accept".
- "What if you can't run a holdout?" — Compare similar users before and after, or ask "how disappointed would you be if this disappeared?" — weaker evidence, but better than usage counts.
- "How long before you judge?" — At least four to six weeks of cohorts, because novelty takes that long to wear off.