Exploring How Reasoning Models Break Mechanistic Interpretability Techniques
Exploring How Reasoning Models Break Mechanistic Interpretability Techniques reveals several interesting facts.
- Join Prof. Subbarao Kambhampati and host Tim Scarfe for a deep dive into OpenAI's O1
- http://80000hours.org/mlst Visit our sponsor 80000 hours - grab their free career guide and check out their podcast! Use our ...
- 0:00 Introduction and Agenda 0:40 What is
- Episode 70 of the Stanford MLSys Seminar “Foundation
- In this episode, we host Jonas Geiping from ELLIS Institute & Max-Planck Institute for Intelligent Systems, Tübingen AI Center, ...
In-Depth Information on How Reasoning Models Break Mechanistic Interpretability Techniques
A talk I gave to my MATS 9.0 training program about A discussion on the philosophy of deep learning, EuroPython 2025 — South Hall 2B on 2025-07-17] *Hacking LLMs: An Introduction to This is a talk I gave to my MATS 9.0 training scholars about the big picture of mech interp - as of Oct 2025, what had changed?
For more information about Stanford's graduate programs, visit: https://online.stanford.edu/graduate-education November 7, 2025 ...
Stay tuned for more updates related to How Reasoning Models Break Mechanistic Interpretability Techniques.