Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 12 days ago • 149
Language Models Can Control Their Own Attention Paper • 2609.02737 • Published 13 days ago • 75
Gemma 4 Collection Our most intelligent open models to date • 16 items • Updated 26 days ago • 1.13k
REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects Paper • 2510.21585 • Published Oct 24, 2025 • 8
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Paper • 2603.19146 • Published Jun 5
Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training Paper • 2602.14759 • Published Feb 16
Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers Paper • 2602.14760 • Published Feb 16