Reasoning Quality Emerges Early: Data Curation for Reasoning Models Paper • 2606.26797 • Published 25 days ago • 1
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs Paper • 2606.32032 • Published 20 days ago • 29
AutoMem: Automated Learning of Memory as a Cognitive Skill Paper • 2607.01224 • Published 19 days ago • 21
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published 24 days ago • 9
MCP Server Architecture Patterns for LLM-Integrated Applications Paper • 2606.30317 • Published 21 days ago • 1
A Sovereign, Open-Source Foundation Model for German and English Paper • 2607.09424 • Published 10 days ago • 9
MemoHarness: Agent Harnesses That Learn from Experience Paper • 2607.14159 • Published 6 days ago • 1
From Foundation to Application: Improving VLA Models in Practice Paper • 2607.06403 • Published 13 days ago • 20
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 11 days ago • 74
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 10 days ago • 80
UniVR: Thinking in Visual Space for Unified Visual Reasoning Paper • 2607.12800 • Published 6 days ago • 28
Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors Paper • 2607.00447 • Published 19 days ago • 1
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published 27 days ago • 64
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems Paper • 2606.18837 • Published Jun 17 • 58
Human-like autonomy emerges from self-play and a pinch of human data Paper • 2606.19370 • Published Jun 11 • 1
Agent-as-a-Router: Agentic Model Routing for Coding Tasks Paper • 2606.22902 • Published 28 days ago • 38