Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 7 days ago • 378
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 7 days ago • 232
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 7 days ago • 86
view article Article LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation LiquidAI • 21 days ago • 35
view reply Have you experimented with or tested different vocabulary sizes to see if the vocab size directly impacts the overtraining threshold and benchmark peak?
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 149
view article Article Distillation in 2026 (so far): which frontier models use it and how sergiopaniego • Jul 8 • 21
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • about 1 month ago • 40