Submitted by Fangxu 13 Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning Microsoft Research 5 2
Submitted by Akshay Nambi 13 Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scale Microsoft Research 37 2
Submitted by ytz 16 LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Microsoft Research 2
Submitted by Baohao Liao 13 Multi-Turn On-Policy Distillation with Prefix Replay Microsoft Research 25 2
Submitted by Yif Yang 143 RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Microsoft Research 469 2
Submitted by Sukmin Cho 32 ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving Microsoft Research 3
Submitted by Kinam Kim 7 Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement Microsoft Research 2