Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

aaailarge reasoning modelslinear probesmathematical reasoningmean ablationreinforcement learningrepresentational qualitysupervised fine-tuningtoken allocation

Computer Science > Artificial Intelligence


arXiv:2607.26119 (cs)

Submitted on 28 Jul 2026


Title: Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models


Authors: Antyabha Rahman, Akshaj Gurugubelli, Omar Ankit, Kevin Zhu, and Aishwarya Balwani


Abstract


Large reasoning models trained via reinforcement learning (RL) have increasingly demonstrated superior performance on mathematical reasoning tasks compared to their supervised fine-tuned (SFT) counterparts. However, the mechanistic basis for this advantage remains unclear. In this work, we investigate the internal representational differences that enable RL models' improved performance. Our study presents two converging lines of evidence:


  1. Linear Probe Analysis: Linear probes trained on layer-wise hidden states reveal that RL models achieve higher accuracy in predicting answer correctness relative to SFT models, indicating more linearly separable and structured internal representations.

    1. Mean Ablation Studies: RL models develop a hierarchical architecture where deeper layers become progressively more critical for performance, whereas SFT models distribute importance more uniformly across layers.

    2. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While some RL-tuned models exhibit higher variability than their SFT counterparts, others show strong consistency, suggesting that token allocation may depend more on the overall training pipeline than on the RL versus SFT distinction alone. We interpret this token-allocation variability as reflecting the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that display under-determined, potentially non-identifiable solution behavior.


      Comments


      • Presented at the Second Workshop on XAI4Science, AAAI 2026.

      Subjects


      • Artificial Intelligence (cs.AI)
      • Computation and Language (cs.CL)

      Citation


      via ArXiv AI

Related