When the Domain Expert Writes the Exam, the Answer Key, and the Grade
SMEs are building AI products that pass their own evals and fail the users they were built for. Five things that keep an eval honest.
With all these subject matter experts becoming builders, I’m starting to see AI products pass their evals and fail the users they’re built for. The basics of data-driven product development matter more than ever.
I’m excited that domain experts can now play a bigger role in building products they have spent years wishing existed. They know the workflow better than almost anyone. They can prototype ideas, expose bad assumptions early, and show product teams what good output actually looks like.
But I’m also seeing a concerning version of this shift.
Writing the exam and grading your own work
The domain expert builds the product, writes the eval, decides what a correct answer looks like, and then grades the result. They write the exam, make an answer key from their own beliefs, and grade their own work.
The product passes their evals, but the answer key may have very little to do with the real users. A domain expert may know the workflow better than anyone, but that does not mean their workflow represents every user, or that they know how to protect an analysis from bias, confounding, bad proxies, or their own assumptions.
A lot of the evals I’m seeing test whether the feature works as designed, not whether it was worth building. Like measuring how often a button takes me to the expected page, rather than its effect on a conversion rate.
In the work I’ve seen, the best results come when domain experts, product, design, engineering, and data science work together as a unit. The domain expert is a strong new addition to the product team, not a replacement for it.
Five things that keep an eval honest
Define the user impact before building. What should change for the user, for their behavior, and eventually for the business?
Evaluate more than technical correctness. Test whether the system worked, whether the answer was correct and useful, and whether a real outcome changed.
Divide the judgment across the team. Let domain experts define excellent work, product and design represent the user, engineering own reliability, and data science determine what the evidence supports.
Protect the evaluation from the people optimizing against it. Use unseen holdout cases and multiple reviewers. Examine disagreements. Do not change the rubric because the latest version got a bad score.
Validate the eval against the real world. Test whether better eval scores correspond to better user behavior and business outcomes. Use experiments when you can and causal inference when you can’t.
AI is making products easier to build. That makes knowing whether we built the right thing more valuable, not less. The domain expert should be inside the evaluation loop. They should not be the entire loop.
This is why I think data science is becoming more important, not less.
10+ years in product data science, causal inference and AI evaluation at Stripe, Nextdoor and Ontra.