Why OpenAI's Model Hacked Hugging Face | Tom McGrath, Goodfire

South Park Commons

South Park Commons

4,748 views

Why did OpenAI’s model hack into Hugging Face?

Tom McGrath, co-founder and Chief Scientist at Goodfire, joins South Park Commons Partner Jonathan Brebner to discuss OpenAI’s recent reward-hacking incident and why it highlights a major challenge in AI: understanding what happens inside these models.

They explore how interpretability could help researchers move beyond trial-and-error development, the role of neural geometry in understanding model behavior, and how Goodfire is building tools to make frontier AI systems more understandable and controllable.

1. Tom McGrath: LinkedIn: tom-mcgrath-7337bb151
2. Jonathan Brebner: LinkedIn: jonathan-brebner
3. South Park Commons: LinkedIn: southparkcommons

Apply to SPC: https://www.southparkcommons.com/apply

Chapters:
00:00:30 - Why AI Interpretability Matters Now
00:02:15 - When AI Reward Hacking Becomes Real
00:05:07 - The Core Problem of Interpretability
00:07:23 - Intentional Design: Steering What Models Learn
00:12:43 - Neural Geometry: The Shapes Inside AI Models
00:22:29 - Turning Interpretability Research Into a Product
00:25:28 - How Research Changes Inside a Startup
00:32:48 - Science Fiction, AI Risk, and the Future