When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 14 days ago • 110
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 17 days ago • 215
Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work Paper • 2609.11977 • Published 27 days ago • 117
Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement Paper • 2609.13406 • Published 20 days ago • 84
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 20 days ago • 264
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 17 days ago • 250
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 26 days ago • 165
Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy Paper • 2609.07470 • Published 24 days ago • 23
trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 Text Generation • 2.44M • Updated 23 days ago • 11.1M • 54
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 23 days ago • 327