q1716523669/cogrpo-n3-strict-qwen25-3b-x-llama32-3b-x-qwen3-1p7b-math345-groupB-llama32-end Reinforcement Learning • 175k • Updated 18 days ago • 13 • 1
electricsheepafrica/africa-ghana-agriculture-inputs-outputs-22b2f1a4 Viewer • Updated 18 days ago • 28 • 32 • 1
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment Paper • 2607.02471 • Published Jul 2 • 6
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluation Paper • 2606.31292 • Published Jun 30 • 9
Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories Paper • 2606.02060 • Published Jun 1 • 58
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 243
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search Paper • 2605.20244 • Published May 18 • 4
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation Paper • 2605.23271 • Published May 22 • 82