Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity Paper • 2608.13430 • Published 7 days ago • 12
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 19 days ago • 259
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Paper • 2608.06352 • Published 14 days ago • 23
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay Paper • 2608.05784 • Published 14 days ago • 28