text-embeddings Tarka-AIR/Tarka-Embedding-150M-V1 Feature Extraction • 0.2B • Updated Nov 18, 2025 • 534 • 7
LLMs Qwen/Qwen2-VL-2B-Instruct Image-Text-to-Text • 2B • Updated Jan 12, 2025 • 1.83M • 518 Qwen/QwQ-32B-Preview Text Generation • 33B • Updated Jan 12, 2025 • 19.2k • • 1.74k MiniMaxAI/MiniMax-M1-80k Text Generation • 456B • Updated Jul 7, 2025 • 1.17k • • 692 EssentialAI/essential-web-v1.0 Preview • Updated Oct 2, 2025 • 12.3k • 245
text-to-image Wan-AI/Wan2.2-T2V-A14B Text-to-Video • Updated Aug 7, 2025 • 6.85k • • 549 QuantStack/Wan2.2-T2V-A14B-GGUF Text-to-Video • 14B • Updated Jul 29, 2025 • 361k • 273
VLM HuggingFaceM4/Idefics3-8B-Llama3 Image-Text-to-Text • 8B • Updated Dec 2, 2024 • 98.4k • 305 HuggingFaceTB/SmolVLM-Instruct Image-Text-to-Text • 2B • Updated Apr 8, 2025 • 30.6k • 597 HuggingFaceTB/SmolLM3-3B Text Generation • 3B • Updated Sep 10, 2025 • 586k • 1.02k
LLMs-optimizations Prompt Cache: Modular Attention Reuse for Low-Latency Inference Paper • 2311.04934 • Published Nov 7, 2023 • 33 Qwen/Qwen2-VL-2B-Instruct Image-Text-to-Text • 2B • Updated Jan 12, 2025 • 1.83M • 518
Prompt Cache: Modular Attention Reuse for Low-Latency Inference Paper • 2311.04934 • Published Nov 7, 2023 • 33
text-embeddings Tarka-AIR/Tarka-Embedding-150M-V1 Feature Extraction • 0.2B • Updated Nov 18, 2025 • 534 • 7
text-to-image Wan-AI/Wan2.2-T2V-A14B Text-to-Video • Updated Aug 7, 2025 • 6.85k • • 549 QuantStack/Wan2.2-T2V-A14B-GGUF Text-to-Video • 14B • Updated Jul 29, 2025 • 361k • 273
VLM HuggingFaceM4/Idefics3-8B-Llama3 Image-Text-to-Text • 8B • Updated Dec 2, 2024 • 98.4k • 305 HuggingFaceTB/SmolVLM-Instruct Image-Text-to-Text • 2B • Updated Apr 8, 2025 • 30.6k • 597 HuggingFaceTB/SmolLM3-3B Text Generation • 3B • Updated Sep 10, 2025 • 586k • 1.02k
LLMs Qwen/Qwen2-VL-2B-Instruct Image-Text-to-Text • 2B • Updated Jan 12, 2025 • 1.83M • 518 Qwen/QwQ-32B-Preview Text Generation • 33B • Updated Jan 12, 2025 • 19.2k • • 1.74k MiniMaxAI/MiniMax-M1-80k Text Generation • 456B • Updated Jul 7, 2025 • 1.17k • • 692 EssentialAI/essential-web-v1.0 Preview • Updated Oct 2, 2025 • 12.3k • 245
LLMs-optimizations Prompt Cache: Modular Attention Reuse for Low-Latency Inference Paper • 2311.04934 • Published Nov 7, 2023 • 33 Qwen/Qwen2-VL-2B-Instruct Image-Text-to-Text • 2B • Updated Jan 12, 2025 • 1.83M • 518
Prompt Cache: Modular Attention Reuse for Low-Latency Inference Paper • 2311.04934 • Published Nov 7, 2023 • 33