StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 22 days ago • 11
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 6 days ago • 78
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 8 days ago • 113
Resolving Conflicting Evidence in Automated Fact-Checking: A Study on Retrieval-Augmented LLMs Paper • 2505.17762 • Published May 23, 2025 • 1
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 9 days ago • 38
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 8 days ago • 259
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 8 days ago • 99
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces Paper • 2605.29288 • Published May 28 • 9
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Paper • 2604.11297 • Published Apr 13 • 144
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification Paper • 2603.15726 • Published Mar 16 • 187
IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse Paper • 2603.12201 • Published Mar 12 • 69
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs Paper • 2603.05890 • Published Mar 6 • 93
From Perception to Action: An Interactive Benchmark for Vision Reasoning Paper • 2602.21015 • Published Feb 24 • 26
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities Paper • 2601.21937 • Published Jan 29 • 20