SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 7 days ago • 157
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 6 days ago • 92
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 11 days ago • 277
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Paper • 2608.11752 • Published 12 days ago • 22
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Paper • 2608.08975 • Published 15 days ago • 48
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries Paper • 2608.05604 • Published 19 days ago • 79
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 24 days ago • 261
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 14 days ago • 10
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published 19 days ago • 23
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 25 days ago • 97
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Paper • 2607.29613 • Published 25 days ago • 28
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 22 days ago • 156
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 26 days ago • 309