0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Test your AI before it embarrasses you

What PwC Middle East missed, and the AI testing framework your team can copy this week

Thank you to Jason Ives and everyone else who tuned in this evening!

PwC Middle East published four AI reports last month with fabricated citations, invented academic papers, and one footnote still carrying ChatGPT's tracking tag in the URL. The Financial Times even covered it. The problem is that they published AI outputs without testing them first, not that they used AI in the first place.

I spent this session walking through the testing system every team using AI needs, and how to apply both to your AI outputs.

Tomorrow I’m publishing the full build on The Data Letter. It catches broken AI workflows in a single review cycle and measures whether your changes to your prompts or review process are working. Inside:

  • The exact spreadsheet template, columns filled in

  • An AI grader prompt

  • A regression suite you can adapt to any workflow your team uses

  • A capability suite that shows where the AI is currently falling short

  • A weekly review cadence: who runs the evals, who reviews the failures, and what happens when a regression breaks

  • Three specific failure patterns to watch for, drawn from the PwC reports and other public examples

See you tomorrow!

Discussion about this video

User's avatar

Ready for more?