WorkshopPredicting Future Behaviors in Reasoning Models Enables Better Steering
Evgenii Kortukov, Piotr Komorowski, Florian Klein, Paula Engl, Gabriele Sarti, Seong Joon Oh, Sebastian Lapuschkin, Wojciech Samek
ICML Workshop Mechanistic Interpretability, 2026Reasoning models represent distributions over future behaviors; probing these lets us steer the outcome by selecting reasoning sentences.
