Predicting Future Behaviors in Reasoning Models Enables Better Steering
Published in ICML Workshop Mechanistic Interpretability, 2026
Evgenii Kortukov, Piotr Komorowski, Florian Klein, Paula Engl, Gabriele Sarti, Seong Joon Oh, Sebastian Lapuschkin, Wojciech Samek
Interpretability tools that predict and control model behaviors typically rely on contrastive input pairs. This binary data hides the probabilistic nature of language model decision making. During reasoning, LLMs can keep track of multiple behavioral options before committing. We find that these future behavior distributions are reflected in their representations. We can steer the behavioral outcome by choosing the reasoning sentences that maximize the estimated probability of a future behavior. We propose a simple algorithm — Future Probe Controlled Generation (FPCG) — that samples multiple candidate sentences at each reasoning step and selects the one that maximizes the activation of a probe predicting future behavior likelihoods, enabling steering at the text level without modifying activations and with less quality degradation than standard activation steering.
Full paper Project page Code OpenReview Demo
title={Predicting Future Behaviors in Reasoning Models Enables Better Steering},
author={Evgenii Kortukov and Piotr Komorowski and Florian Klein and Paula Engl and Gabriele Sarti and Seong Joon Oh and Sebastian Lapuschkin and Wojciech Samek},
booktitle={Mechanistic Interpretability Workshop at ICML 2026},
year={2026},
url={https://openreview.net/forum?id=48NnVTsirb}
}
