Qwen/Qwen2.5-Omni-7B
Any-to-Any • 11B • Updated • 424k • 1.92k
This collection includes all the models, datasets and Spaces mentioned in the blog Vision Language Models: 2025 Update
Chat with text, audio, images, and video, get spoken replies
A unified multimodal understanding and generation model.
Chat with images, videos, PDFs and get thoughtful responses
Answer questions about images with AI chat
Ask questions about images or videos and get answers
Chat with an AI using text and images for visual answers
Annotate and describe images with text prompts
Generate text answers or segment objects from images
Demo for ShieldGemma 2, multimodal safety model
Check if text and images are safe
Chat with a multimodal AI using text, images, or video
Answer questions about uploaded videos or images