Collections
Discover the best community collections!
Collections including paper arxiv:2608.27454
-
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Paper • 2606.29538 • Published • 145 -
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Paper • 2608.02287 • Published • 31 -
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Paper • 2608.05139 • Published • 27 -
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Paper • 2608.11079 • Published • 17
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level
Paper • 2411.03562 • Published • 70 -
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
Paper • 2502.06060 • Published • 37 -
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196 -
SurveyX: Academic Survey Automation via Large Language Models
Paper • 2502.14776 • Published • 100
-
Demystifying Agent Skills: Why They Work-Until They Don't
Paper • 2608.14036 • Published • 169 -
Modular Cognitive Architecture Emerges in Large Language Models
Paper • 2608.13567 • Published • 15 -
Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
Paper • 2607.26179 • Published • 1 -
Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept
Paper • 2607.20995 • Published • 1
-
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper • 2605.20025 • Published • 192 -
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 89 -
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 47 -
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Paper • 2605.13941 • Published • 25
-
EvoMaster: A Foundational Agent Framework for Building Evolving Autonomous Scientific Agents at Scale
Paper • 2604.17406 • Published • 7 -
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
Paper • 2605.15301 • Published • 23 -
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Paper • 2605.11739 • Published • 61 -
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Paper • 2605.18401 • Published • 132
-
Demystifying Agent Skills: Why They Work-Until They Don't
Paper • 2608.14036 • Published • 169 -
Modular Cognitive Architecture Emerges in Large Language Models
Paper • 2608.13567 • Published • 15 -
Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
Paper • 2607.26179 • Published • 1 -
Where Animacy Lives in Large Language Models: Tracing the Circuits of the Animacy Concept
Paper • 2607.20995 • Published • 1
-
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Paper • 2606.29538 • Published • 145 -
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
Paper • 2608.02287 • Published • 31 -
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning
Paper • 2608.05139 • Published • 27 -
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Paper • 2608.11079 • Published • 17
-
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Paper • 2605.20025 • Published • 192 -
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
Paper • 2605.19769 • Published • 89 -
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
Paper • 2605.10912 • Published • 47 -
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
Paper • 2605.13941 • Published • 25
-
EvoMaster: A Foundational Agent Framework for Building Evolving Autonomous Scientific Agents at Scale
Paper • 2604.17406 • Published • 7 -
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
Paper • 2605.15301 • Published • 23 -
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Paper • 2605.11739 • Published • 61 -
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Paper • 2605.18401 • Published • 132
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level
Paper • 2411.03562 • Published • 70 -
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
Paper • 2502.06060 • Published • 37 -
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196 -
SurveyX: Academic Survey Automation via Large Language Models
Paper • 2502.14776 • Published • 100