0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Test your AI before it embarrasses you

What PwC Middle East missed, and the AI testing framework your team can copy this week

Thank you to Jason Ives and everyone else who tuned in this evening!

PwC Middle East published four AI reports last month with fabricated citations, invented academic papers, and one footnote still carrying ChatGPT's tracking tag in the URL. The Financial Times even covered it. The problem is that they published AI outputs without testing them first, not that they used AI in the first place.

I spent this session walking through the testing system every team using AI needs, and how to apply both to your AI outputs.

Tomorrow I’m publishing the full build on The Data Letter. It catches broken AI workflows in a single review cycle and measures whether your changes to your prompts or review process are working. Inside:

  • The exact spreadsheet template, columns filled in

  • An AI grader prompt

  • A regression suite you can adapt to any workflow your team uses

  • A capability suite that shows where the AI is currently falling short

  • A weekly review cadence: who runs the evals, who reviews the failures, and what happens when a regression breaks

  • Three specific failure patterns to watch for, drawn from the PwC reports and other public examples

Get it here:

The Eval System for AI Workflows That Non-ML Teams Can Run

The Eval System for AI Workflows That Non-ML Teams Can Run

Last month, PwC Middle East published four thought leadership reports on AI and electric vehicles. The reports contained fabricated citations, invented academic papers, and one footnote with ChatGPT’s tracking parameter still attached to the URL. The Financial Times picked it up. The problem with PwC’s situation isn’t that they’re using AI to produce reports; it’s releasing those AI outputs without proper testing.

Discussion about this video

User's avatar

Ready for more?