Submitted by Li Lyna Zhang 63 LoongRL:Reinforcement Learning for Advanced Reasoning over Long Contexts Microsoft Research 5
Submitted by Junpeng Liu 27 DocReward: A Document Reward Model for Structuring and Stylizing Microsoft Research 14 3
Submitted by Martina Vilas 2 Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning Microsoft Research 2
Submitted by Zijian Li 5 PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images Microsoft Research 32 2
Submitted by Yifei Shen 10 Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective Microsoft Research 2