Are Latent Reasoning Models Easily Interpretable?
摘要
这篇论文研究潜在推理模型(LRM)的可解释性,考察了两个前沿 LRM。主要发现有三点:一是在逻辑推理数据集上,LRM 几乎总能在不使用潜在推理 token 的情况下得出相同答案,说明这些 token 常被闲置,这可能解释了 LRM 为何未必稳定优于显式推理方法;二是当潜在推理 token 确实影响性能时,对预测正确的实例,有 65% 到 93% 的概率能解码出金标准推理轨迹,说明模型常实现预期解法而非不可解释的推理过程;三是提出一种无需预先知道金标准轨迹、即可从潜在 token 解码出经核验的自然语言推理轨迹的方法,该方法对多数正确预测能找到核验轨迹,对错误预测则只能找到少数。作者据此认为,当前 LRM 编码的推理过程大体可解释,且可解释性本身可作为预测正确性的信号。
荐读理由
这篇研究用实验数据告诉你:潜推理模型的推理标记常常是摆设,去掉也能答对,且可解码出自然语言推理链,这能纠正你对这类模型可解释性的既有判断。
原文
Computer Science > Machine Learning
[Submitted on 6 Apr 2026 (v1), last revised 10 Aug 2026 (this version, v2)]
Title:Are Latent Reasoning Models Easily Interpretable?
Authors:Connor Dilgren, Sarah Wiegreffe
View a PDF of the paper titled Are Latent Reasoning Models Easily Interpretable?, by Connor Dilgren and Sarah Wiegreffe
Abstract:Latent reasoning models (LRMs) have attracted significant research interest due to their low inference cost (relative to explicit reasoning models) and theoretical ability to explore multiple reasoning paths in parallel. However, these benefits come at the cost of reduced interpretability: LRMs are difficult to monitor because they do not reason in natural language. This paper presents an investigation into LRM interpretability by examining two state-of-the-art LRMs. First, we find that latent reasoning tokens are often unnecessary for LRMs' predictions; on logical reasoning datasets, LRMs can almost always produce the same final answers without using latent reasoning at all. This underutilization of reasoning tokens may partially explain why LRMs do not consistently outperform explicit reasoning methods and raises doubts about the stated role of these tokens in prior work. Second, we demonstrate that when latent reasoning tokens are necessary for performance, we can decode gold reasoning traces up to 65-93% of the time for correctly predicted instances. This suggests LRMs often implement the expected solution rather than an uninterpretable reasoning process. Finally, we present a method to decode a verified natural language reasoning trace from latent tokens without knowing a gold reasoning trace a priori, demonstrating that it is possible to find a verified trace for a majority of correct predictions but only a minority of incorrect predictions. Our findings highlight that current LRMs largely encode interpretable processes, and interpretability itself can be a signal of prediction correctness.
https://doi.org/10.48550/arXiv.2604.04902
arXiv-issued DOI via DataCite
| Comments: | |
|---|---|
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.04902 [cs.LG] |
| (or arXiv:2604.04902v2 [cs.LG] for this version) | |
Submission history
From: Connor Dilgren [view email] [v1] Mon, 6 Apr 2026 17:50:06 UTC (651 KB) [v2] Mon, 10 Aug 2026 13:30:13 UTC (699 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Are Latent Reasoning Models Easily Interpretable?, by Connor Dilgren and Sarah Wiegreffe
Current browse context:
cs.LG
Change to browse by:
References & Citations
Loading...
BibTeX formatted citation
Data provided by:
Bookmark
Bibliographic and Citation Tools
Bibliographic Explorer (What is the Explorer?)
Connected Papers (What is Connected Papers?)
Litmaps (What is Litmaps?)
scite Smart Citations (What are Smart Citations?)
Code, Data and Media Associated with this Article
alphaXiv (What is alphaXiv?)
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub (What is DagsHub?)
Gotit.pub (What is GotitPub?)
Hugging Face (What is Huggingface?)
ScienceCast (What is ScienceCast?)
Demos
Replicate (What is Replicate?)
Hugging Face Spaces (What is Spaces?)
TXYZ.AI (What is TXYZ.AI?)
Recommenders and Search Tools
Influence Flower (What are Influence Flowers?)
CORE Recommender (What is CORE?)
IArxiv Recommender (What is IArxiv?)
Author
Venue
Institution
Topic
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)
这条对你有帮助吗?