Phase 5: Evaluation & Quality for AI Products · 140 min · Golden datasets · Eval rubrics · DeepEval
Project: Evaluation Plan & Rubric
A PRD says what you're building. An evaluation plan says how you'll know, every single day after launch, whether it's still working.
Hiring signal: This project is the direct, portfolio-grade deliverable for the single highest-leverage AI PM interview signal identified in the research: 'how you evaluate a model that's sometimes wrong.' A candidate who can produce a real evaluation plan, a scored rubric, and a sample scoring pass for a real feature -- and explain every choice in it -- has demonstrable, defensible proof of the skill, not just a claim about it on a resume. This artifact carries forward into Phase 7's responsible AI review and the Phase 9 capstone case study.
What you will learn
- Produce a complete evaluation plan for a real AI feature, covering golden dataset design, rubric, automated metrics, and offline/online testing strategy
- Build a full scoring rubric with anchors, weights, and an explicit pass/fail rule for the same feature
- Run a sample scoring pass against the rubric and report results the way a data science partner would expect
- Translate the evaluation plan into ship-readiness criteria that connect directly to Phase 7's responsible AI review
The Problem
You've now worked through the full arc of an eval system: why probabilistic products need one at all (Lesson 1), how to design a human rubric that measures something real (Lesson 2), how automated pipelines extend that judgment to scale (Lesson 3), how offline and online testing answer different questions (Lesson 4), how to interpret hallucination, bias, and safety metrics without computing them (Lesson 5), and how to turn a dashboard full of mixed signals into a real ship/no-ship call (Lesson 6). Each of those lessons taught one piece in isolation. A real AI PM has to assemble all six pieces into a single artifact before a feature ever reaches a data science team for execution — because "we'll figure out evaluation once it's built" is exactly the eval-debt failure mode from Lesson 1, just deferred to a more expensive moment.
This project is that artifact. You will produce a complete evaluation plan, a scoring rubric, and a sample scoring pass for the same AI feature you carried through Phase 2's opportunity assessment and Phase 4's PRD. If you haven't been carrying a single feature through the course, pick one now — a specific, real or realistic AI product feature with enough surface area to need genuine evaluation (a support-ticket summarizer, a resume screener, a shopping assistant, a meeting-notes generator, or similar). This deliverable becomes the direct input to Phase 7's responsible AI review, and eventually one full section of the Phase 9 capstone case study — so build it as if a real data science team is going to receive it and start executing against it next week, because in the course's continuity, that's exactly what happens next.
What a Complete Evaluation Plan Contains
A data science team receiving a PM-authored evaluation plan should be able to start building against it without a clarifying-questions meeting. That means the plan has to specify, concretely, not vaguely:
- Golden dataset design (Lesson 1): what categories of input it must cover (happy path, edge case, at least one known or anticipated failure mode, adversarial if relevant), a target size, and who owns adding new examples over time as production surfaces new cases.
- Human eval rubric (Lesson 2): 3-5 weighted, anchored criteria specific to this feature — not generic "quality" — plus an explicit pass/fail rule for any catastrophic failure type, and a sampling strategy for ongoing human review once the feature is live.
- Automated metrics (Lesson 3): which DeepEval/Confident-AI-style metrics apply (faithfulness, relevance, bias, toxicity, or others specific to the feature), what threshold constitutes a pass, and a calibration cadence — how often the automated judge gets checked against human scores.
- Offline/online testing strategy (Lesson 4): what the golden-set eval delta process looks like for routine changes, whether shadow testing is warranted before a live test, and the A/B test design for a full rollout decision — including a pre-committed sample size, so nobody is tempted to peek and stop early.
- Hallucination/bias/safety measurement plan (Lesson 5): which metrics will be tracked ongoing, which are one-time pre-launch checks, and — critically — which categories of this specific feature are high-severity enough that a regression there should function as a near-automatic hard blocker in the ship/no-ship framework.
- Ship/no-ship criteria (Lesson 6): the actual thresholds and hard-blocker conditions that will be applied at the next launch review, stated concretely enough that two different people reading the plan would reach the same decision given the same dashboard.
The plan is worthless if the thresholds are vague
"We'll evaluate quality before launch" is not a plan -- it's a restatement of intent with the actual decision deferred to whoever happens to be in the room later. A real evaluation plan states numbers and conditions in advance: what hallucination rate on what category triggers a hold, what confidence interval width is required before an A/B result counts as trustworthy, what fairness gap on what protected characteristic is a hard blocker versus a monitor-and-report item. Writing these numbers down before you have a horse in the race (before a specific launch is on the calendar and a specific number is inconveniently close to the line) is what makes the plan a genuine commitment device instead of a document that gets quietly reinterpreted under launch pressure.
Unlock the full lesson
You've read the first 2 sections. The rest of this lesson covers Building the Rubric and Running a Sample Scoring Pass, Connecting Forward: Why This Feeds the Rest of the Course, Build It, What to Practice — plus a hands-on lab, quiz, and project artifact.
Create a free account to unlock Phase 0 and Phase 1 of every course — no credit card.
Browse all courses · View pricing · DeVenture Academy