On this page
Open Source Model
Overview
Vimoo-1.0 is an open-source video generation model built on the MMDiT architecture. It supports text-to-video (T2V) and image-to-video (I2V), using prompts or reference images to create videos with control over subjects, actions, scenes, and camera movement.
Model Capabilities
- Text and image inputs: create a scene from a description or animate a reference image.
- HD video generation: generate 720p video at 24 fps, with 5–10-second outputs.
- In-video text conditioning: generate specified text within the scene.
- Multiple model variants: Vimoo-1.0-Distilled adds a distilled option. The planned Lite TI2V release uses one base for text-to-video and image-to-video, with a shared distilled adapter and an INT8 base option.
See Vimoo-1.0 for architecture, variants, and capabilities. Find weights and matching LoRAs in Resources.
Get Started
Vimoo provides model weights, inference examples, and ComfyUI workflows for local creation, programmatic integration, and research.
Local Deployment
Choose a workflow and generate your first video with ComfyUI or MLX.
Developer Integration
Use the native Diffusers Python API or CLI on a local computer or server.
Open-Source Resources
| Resource | Entry |
|---|---|
| Model and inference examples | Vimoo-1.0Coming soon |
| Model weights | Vimoo-AI on Hugging Face |
See Local Deployment for hardware, device timings, and local generation examples.