UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations Paper • 2608.10835 • Published Aug 11 • 4
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 184
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models Paper • 2607.26056 • Published Jul 28 • 18
Latent-Identity Tuning in Text-to-Image Personalization Models Paper • 2607.11885 • Published Jul 13 • 15
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published Jul 2 • 71
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents Paper • 2606.24893 • Published May 29 • 9