Phase 5: Evaluation & Quality for AI Products · 50 min · Eval dashboards · Ship/no-ship frameworks · Braintrust
Reading an Eval Dashboard & Making Ship/No-Ship Calls
A real dashboard never says 'ship' or 'don't ship.' It says 'here are eight numbers, three of them disagree with each other, decide.'
Hiring signal: This lesson is the synthesis of every eval concept in the phase, and it's the closest the course gets to simulating the exact moment interviewers are testing for when they ask 'walk me through how you'd know if this AI feature is working' -- a real dashboard with mixed signals, and a PM who has to make and defend one call. Candidates who can't hold multiple conflicting signals at once and still produce a decision (not just 'more testing needed') get filtered out of AI PM loops fast.
What you will learn
- Read a real eval dashboard containing multiple, sometimes conflicting, quality and business signals
- Weigh a quality regression against a latency or cost improvement using an explicit framework, not a gut call
- Distinguish 'ship,' 'hold,' and 'ship with guardrails' as distinct decisions with different justification requirements
- Write a launch decision memo that a cross-functional team would accept as well-reasoned, not just decisive
The Problem
It's launch review day for "prompt v7" of an AI shopping assistant. The dashboard, built from everything in this phase, shows: faithfulness up 3 points, relevance flat, latency down 40% (a new smaller model), cost down 55%, hallucination rate up from 2.1% to 2.8% overall but down in the highest-severity category (payment/pricing claims), demographic parity difference unchanged, and the A/B test shows a statistically significant +2.1% increase in add-to-cart rate with a confidence interval that excludes zero. Engineering wants to ship today — the cost savings alone are material. A cautious teammate points at the hallucination rate going up and says hold. Both are reading the same dashboard correctly. The disagreement isn't about the numbers — it's about how to weigh them against each other, and that weighing is a PM decision, not a data science one.
This is the moment every earlier lesson in this phase has been building toward: not reading one metric, but reading a dashboard where several metrics point in different directions and a real decision is due.
The Three Real Decisions
Most AI PMs default to thinking in binary: ship or don't ship. A mature eval program actually has three distinct decisions available, and picking the right one — not just "yes" or "no" — is often the strongest move:
- Ship — all key metrics clear their bar, no unresolved red flags from Lessons 1-5, and the online test result is trustworthy (adequately powered, statistically significant in the intended direction, confidence interval excludes zero).
- Hold — a metric that matters more than the metrics that improved is regressing, a fairness or safety flag from Lesson 5 is unresolved, or the online test isn't yet trustworthy (underpowered, not significant, or still running).
- Ship with guardrails — the aggregate case for shipping is real, but a specific, identified risk needs a mitigation that doesn't require blocking the whole launch: routing a specific high-risk category to human review, capping rollout to a percentage of traffic while monitoring the regressing metric, or shipping to a subset of use cases while excluding the one that regressed.
"Ship with guardrails" is frequently the correct answer and frequently the one candidates fail to reach for in an interview — the instinct is to collapse a nuanced situation into a binary because binary feels more decisive. It usually isn't; it's just less work.
Weigh metrics against what they're attached to, not against each other in the abstract
An aggregate hallucination rate going up while cost drops 55% is not resolvable by asking "which number is bigger." It's resolved by asking what each number is attached to: does the hallucination increase concentrate in a high-severity category (per Lesson 5's category-breakdown practice), and does the cost savings translate into something users or the business actually value enough to accept that specific risk. In the opening scenario, the hallucination increase is overall but the highest-severity category (payment/pricing claims) actually improved -- that reframes "quality got worse" into "aggregate quality metric moved in a way that doesn't track the category that matters most," which is a very different, more shippable situation.
Unlock the full lesson
You've read the first 2 sections. The rest of this lesson covers A Framework for Weighing Conflicting Signals, Writing the Decision Memo, Build It, What to Practice — plus a hands-on lab, quiz, and project artifact.
Create a free account to unlock Phase 0 and Phase 1 of every course — no credit card.
Browse all courses · View pricing · DeVenture Academy