Last month, PwC Middle East published four thought leadership reports on AI and electric vehicles. The reports contained fabricated citations, invented academic papers, and one footnote with ChatGPT’s tracking parameter still attached to the URL. The Financial Times picked it up. The problem with PwC’s situation isn’t that they’re using AI to produce reports; it’s releasing those AI outputs without proper testing.
On yesterday’s live session, I walked through why every AI workflow at your company needs a test layer, and introduced the two kinds of evals every team using AI should run. Capability evals measure what the AI does well and where it struggles. Regression evals measure whether the AI still handles the outputs your team already relies on. Together they answer whether the AI is safe to use, and whether the work you’re putting into improving your processes is paying off.
Today I’m giving you the entire testing system. A workbook with both tests (capability and regression) already set up, so you can drop your test cases straight in. A starter list of 21 conditions an AI output can be required to meet, sorted into six categories, that your team can copy from and adapt to your workflows. An AI grader prompt that scores each output against each condition, so one person can run both tests in the time it used to take to manually review a single output.
👋🏿 Hey, I’m Hodman. I write The Data Letter for senior managers, operators, and technical builders rolling out AI. Here are some recent popular articles you may have missed:
➡ A four-artifact workflow for calculating AI ROI in a format your CFO will approve, built in 90 minutes.
➡ The intake form, screening guide, vendor checklist, and documentation workflow every enterprise AI project now runs through under U.S. state AI regulations.

