ERNIE-Image The serieas of image generation models, including text2img、img2img. baidu/ERNIE-Image Text-to-Image • 8B • Updated Apr 17 • 68.3k • • 665 baidu/ERNIE-Image-Turbo Text-to-Image • 8B • Updated Apr 17 • 2.53k • • 400 baidu/ERNIE-Image-Aes 8B • Updated May 20 • 453 • 17 baidu/ERIA-1K-Benchmark Preview • Updated May 20 • 64 • 4
Qianfan-VL Qianfan-vl model series. The models are mainly domain enhanced vision language model, targeting enterprise level multi modal understanding scenarios. baidu/Qianfan-OCR Image-Text-to-Text • 5B • Updated Apr 29 • 281k • 1.19k baidu/Qianfan-VL-70B Image-Text-to-Text • 72B • Updated Apr 19 • 64 • 39 baidu/Qianfan-VL-8B Image-Text-to-Text • 9B • Updated Apr 19 • 2.09k • 41 baidu/Qianfan-VL-3B Image-Text-to-Text • 4B • Updated Sep 19, 2025 • 190 • 30
ERNIE 4.5 collection of ERNIE 4.5 models. baidu/ERNIE-4.5-VL-28B-A3B-Thinking Image-Text-to-Text • 30B • Updated Mar 6 • 720 • 541 baidu/ERNIE-4.5-21B-A3B-Thinking Text Generation • 22B • Updated Nov 26, 2025 • 14.5k • 787 baidu/ERNIE-4.5-VL-424B-A47B-Base-Paddle Image-Text-to-Text • 424B • Updated Aug 19, 2025 • 45 • 68 baidu/ERNIE-4.5-VL-424B-A47B-Base-PT Image-Text-to-Text • 424B • Updated Jan 16 • 264 • • 87
ERNIE-Image The serieas of image generation models, including text2img、img2img. baidu/ERNIE-Image Text-to-Image • 8B • Updated Apr 17 • 68.3k • • 665 baidu/ERNIE-Image-Turbo Text-to-Image • 8B • Updated Apr 17 • 2.53k • • 400 baidu/ERNIE-Image-Aes 8B • Updated May 20 • 453 • 17 baidu/ERIA-1K-Benchmark Preview • Updated May 20 • 64 • 4
ERNIE 4.5 collection of ERNIE 4.5 models. baidu/ERNIE-4.5-VL-28B-A3B-Thinking Image-Text-to-Text • 30B • Updated Mar 6 • 720 • 541 baidu/ERNIE-4.5-21B-A3B-Thinking Text Generation • 22B • Updated Nov 26, 2025 • 14.5k • 787 baidu/ERNIE-4.5-VL-424B-A47B-Base-Paddle Image-Text-to-Text • 424B • Updated Aug 19, 2025 • 45 • 68 baidu/ERNIE-4.5-VL-424B-A47B-Base-PT Image-Text-to-Text • 424B • Updated Jan 16 • 264 • • 87
Qianfan-VL Qianfan-vl model series. The models are mainly domain enhanced vision language model, targeting enterprise level multi modal understanding scenarios. baidu/Qianfan-OCR Image-Text-to-Text • 5B • Updated Apr 29 • 281k • 1.19k baidu/Qianfan-VL-70B Image-Text-to-Text • 72B • Updated Apr 19 • 64 • 39 baidu/Qianfan-VL-8B Image-Text-to-Text • 9B • Updated Apr 19 • 2.09k • 41 baidu/Qianfan-VL-3B Image-Text-to-Text • 4B • Updated Sep 19, 2025 • 190 • 30