On this page
Open Source Model
Vimoo 1.0
Vimoo 1.0 is an open-source video generation model built on the MMDiT architecture, supporting both text-to-video and image-to-video generation. Its lightweight variant, Vimoo 1.0 - Lite, unifies both tasks in a single model designed for efficient inference.
Model Overview
The model combines text and image conditioning with video latents to generate subject motion, scene changes, and camera effects.
| Generation mode | Input | Use case |
|---|---|---|
| Text-to-Video (T2V) | Prompt | Create scenes, subject motion, and camera movement from a description |
| Image-to-Video (I2V) | Prompt and one first-frame reference image | Set the opening image, then generate subsequent motion and scene changes |
| Text/Image-to-Video (TI2V) | Prompt, or prompt and one first-frame reference image | A lightweight path for both modes |
| E-commerce Video (I2V-E) | Prompt and product images | Set the product image, then generate storyline product video |
Output Specifications
| Item | T2V | I2V | TI2V | I2V-E |
|---|---|---|---|---|
| Output resolution | 720p | 720p | 720p | 720p |
| Output duration | 5 seconds | 5 seconds | 5 seconds | 20 seconds |
| Aspect ratio | 16:9 | Auto, based on the reference image | 16:9 without an image; Auto with a reference image | 9:16 |
These specifications use 24 fps. For I2V and image-conditioned TI2V, the reference image determines the closest supported output ratio and is resized and center-cropped to match.
Architecture and Highlights
- Prompt adherence: a dedicated text encoder turns prompts into precise generation conditions. An optional prompt enhancer can add visual, cinematic, temporal, and contextual detail, or be disabled to preserve the original wording.
- Cinematic visual quality: training on carefully curated aesthetic data, with detailed annotations for lighting, composition, contrast, color grading, and more, enables precise control over cinematic styles and a wide range of aesthetic preferences.
- Advanced motion generation: a larger and more diverse training corpus improves generalization across motion, semantics, and visual aesthetics, delivering leading performance among open-source and proprietary video generation models.
- Efficient high-definition generation: the lightweight variant uses 32×32×4 spatiotemporal compression and supports T2V and I2V at 720p and 24 fps.
Known Limitations
- Output videos have no audio track. Music or dialogue descriptions do not generate sound.
- I2V and image-conditioned TI2V reference images are resized and center-cropped to the selected output ratio. Check that the subject remains within the retained area.
- The generated video's first frame is not guaranteed to match the reference image pixel for pixel. First-frame conditioning operates in latent space, and preprocessing, VAE reconstruction, and video export can alter the pixels.