Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
jason500
's Collections
audio
siweilian
text_gen_img
duomotai&tuxiangbianji
grounding
Quant
video_preprocess
mutilmodal_video2text
mutil big modal image2text
caption
MMLM
MMLM
updated
Mar 24, 2025
Upvote
-
Sort: Collection
allenai/Molmo-7B-O-0924
Image-Text-to-Text
•
8B
•
Updated
Oct 9, 2025
•
1.65k
•
165
allenai/Molmo-7B-D-0924
Image-Text-to-Text
•
8B
•
Updated
Dec 15, 2025
•
34.6k
•
566
zai-org/cogvlm2-llama3-caption
Video-Text-to-Text
•
13B
•
Updated
May 14, 2025
•
308
•
119
mistralai/Pixtral-12B-Base-2409
Updated
Jul 28, 2025
•
21
•
108
mistralai/Pixtral-12B-2409
Updated
Jun 2
•
7.39k
•
696
zai-org/glm-4v-9b
14B
•
Updated
Mar 3, 2025
•
19.8k
•
268
OpenGVLab/InternVL-Chat-V1-2-SFT-Data
Viewer
•
Updated
Sep 20, 2024
•
573k
•
545
•
29
weic22/InstructSeg
3B
•
Updated
Dec 19, 2024
•
20
•
4
mistralai/Mistral-Small-3.1-24B-Instruct-2503
24B
•
Updated
Dec 22, 2025
•
303k
•
1.38k
Upvote
-
Sort: Collection
Share collection
View history
Collection guide
Browse collections