Thank you to Jason Ives and everyone else who tuned in this evening!
PwC Middle East published four AI reports last month with fabricated citations, invented academic papers, and one footnote still carrying ChatGPT's tracking tag in the URL. The Financial Times even covered it. The problem is that they published AI outputs without testing them first, not that they used AI in the first place.
I spent this session walking through the testing system every team using AI needs, and how to apply both to your AI outputs.
Tomorrow I’m publishing the full build on The Data Letter. It catches broken AI workflows in a single review cycle and measures whether your changes to your prompts or review process are working. Inside:
The exact spreadsheet template, columns filled in
An AI grader prompt
A regression suite you can adapt to any workflow your team uses
A capability suite that shows where the AI is currently falling short
A weekly review cadence: who runs the evals, who reviews the failures, and what happens when a regression breaks
Three specific failure patterns to watch for, drawn from the PwC reports and other public examples
See you tomorrow!









