Arnav Kakani. Thanks to my coauthor Mrinal Agarwal for the help writing the paper. Poster at the 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026.
My paper got accepted to the MATH-AI workshop at NeurIPS this year. The result in it is a negative one. The method I studied does not work, and the paper says so in the title. I want to write down how the project went, because the path to that sentence taught me more than a positive result would have.
The idea I liked
If you have used a language model for math, you have seen this. The solution reads fine. Every step sounds confident. The final number is wrong. Somewhere early there was a slip, and everything after it was built on top.
The idea that pulled me in was simple. A model keeps an internal running state while it writes, called the residual stream. Other researchers had found that some of these internal states carry information related to whether a statement is true. So maybe a solution that is going wrong looks different on the inside. If it does, you could watch the state while the model writes, notice when it drifts too far from where it started, back up a few tokens, and tell the model to recheck.
I built exactly that. The monitor takes the first 10 tokens of a solution as an anchor and measures how far each later state has moved from it. I tried two ways of measuring the distance, a plain one and a fancier one called a Sinkhorn distance that compares groups of states. When the distance crosses a threshold, the monitor rolls back 15 tokens, inserts “Wait, let me recalculate this step to be sure:”, and lets the model continue.
I thought it would work. It is the kind of idea that feels like it should.
The first version of the evidence looked great
One way to score a method like this is to ask the model afterward whether its reasoning stayed coherent. When I scored my runs that way, the method looked like a success. On Mistral, the model’s own judgment reported 82 to 95 percent “recovery.”
Then I checked the answers against the answer key. On those same paths, the real rate of wrong answers turning right was 5 to 20 percent. Worse, the model rated the runs where I did nothing at all at 0.76 to 0.91. On one setting it gave 0.91 to a run whose accuracy was 0.44.
That was the moment the project changed for me. The question stopped being “how well does my monitor work” and became “how would I know if it worked at all.”
Rebuilding the test
I ended up caring about three things.
The first was scoring by the final answer and nothing else. The extracted number matches the reference or it does not.
The second was prompting the models the way they were meant to be prompted. This sounds minor. It moved Qwen’s accuracy on GSM8K from 0.40 to 0.90. Without the chat format, most of the “reasoning failures” were the model printing formatting tokens and giving up. A monitor that catches those is catching a prompting bug.
The third was controls. Every result had to be compared with doing nothing, with each half of the intervention on its own, and with the same intervention applied at a random position that never looks at the hidden state. If my monitor knows something, it has to beat a coin.
I ran two open 7B models, Qwen2.5-7B-Instruct and Mistral-7B-Instruct, on GSM8K and SVAMP. Eleven runs, about 19 hours of GPU time on rented H100 and A100 machines over two days in June.
What came back
The drift score did not predict wrong answers. Qwen was too good at these problems to tell me much, with only 7 to 10 errors per hundred. Mistral got 56 of 100 wrong, so that is where the test had teeth. The detection score there was 0.531 for the plain distance and 0.514 for Sinkhorn. A coin flip is 0.5. The monitor also fired on 48 to 98 percent of the solutions that were correct.
Acting on the score did not help either. The full method was unchanged or worse than doing nothing in five of six settings and 0.02 better in the sixth. It fixed some answers, 9, 7, and 6 across the three Mistral runs. It also broke 14, 5, and 13 answers that had been right.
The random trigger was the result that stung. On Mistral, my monitor’s net recovery was +3 problems out of 56. Intervening at a random spot also got +3.
When I looked at where the monitor was firing, it made sense. On one Mistral setting the median trigger was token 17, right after the monitor finishes setting itself up and before the model has done any arithmetic. It was not finding a mistake. It was restarting the solution.
The odd part is that the signal is real. It is very good at telling a math problem apart from a summarization task. It cannot tell correct math from incorrect math. It was measuring what the model was doing, not whether the model was doing it right.
Deciding to publish a “no”
I could have kept tuning until something looked positive. I wrote up the negative result instead, with every control in the paper and every run in a public repository.
Review was humbling in a useful way. One reviewer went through my released code and run files and found things I had missed. The problems I used to pick which layer to monitor were also among the problems I evaluated on. Two different baselines had ended up with the same name. I fixed all of it for the final version and reported what changed. With the overlapping problems removed, the detection scores were 0.517 and 0.500, so the conclusion held. I am glad someone checked. I am also glad the artifacts were there to be checked, because a claim like mine is only worth something if other people can recompute it.
The paper ends with five conditions a monitor like this should meet before anyone trusts it. Score by the final answer in the model’s own prompt format. Beat doing nothing. Beat a random trigger. Rank failures above chance where there are enough failures to measure. Still work when each piece is removed. Mine clears none.
What I took from it
A lenient evaluation would have let me believe my own idea. The self-judging score was sitting right there, and it said I had built something that worked. The only reason I know otherwise is that I kept adding controls until the result had nowhere to hide.
I also do not think the idea is dead. My result covers one drift score, two small models, and short arithmetic problems, and a test that cannot reject chance does not prove chance. The next things to try are a supervised probe on the same activations, to see whether the information is in there at all, picking the layer by correctness instead of by task, and a harder benchmark that strong small models still fail.
Thank you to Mrinal Agarwal, my coauthor, for the help writing the paper, and to the reviewers who read the code.
Links and citation
- Paper: OpenReview · PDF · Supplementary material
- Code, runs, and paper sources: github.com/ArnavKakani/DCRF
- Venue: The 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026 (poster)
@inproceedings{kakani2026detecting,
title = {Detecting and Correcting Reasoning Failures: A Controlled Negative Result for a Residual-Stream Monitor},
author = {Kakani, Arnav and Agarwal, Mrinal},
booktitle = {The 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026},
year = {2026},
url = {https://openreview.net/forum?id=AxOXuIaLWj}
}