Skip to documentation
On this page

Open Source Model

Vimoo 1.0

Vimoo 1.0 is an open-source video generation model built on the MMDiT architecture, supporting both text-to-video and image-to-video generation. Its lightweight variant, Vimoo 1.0 - Lite, unifies both tasks in a single model designed for efficient inference.

Model Overview

The model combines text and image conditioning with video latents to generate subject motion, scene changes, and camera effects.

Generation modeInputUse case
Text-to-Video (T2V)PromptCreate scenes, subject motion, and camera movement from a description
Image-to-Video (I2V)Prompt and one first-frame reference imageSet the opening image, then generate subsequent motion and scene changes
Text/Image-to-Video (TI2V)Prompt, or prompt and one first-frame reference imageA lightweight path for both modes
E-commerce Video (I2V-E)Prompt and product imagesSet the product image, then generate storyline product video

Output Specifications

ItemT2VI2VTI2VI2V-E
Output resolution720p720p720p720p
Output duration5 seconds5 seconds5 seconds20 seconds
Aspect ratio16:9Auto, based on the reference image16:9 without an image; Auto with a reference image9:16

These specifications use 24 fps. For I2V and image-conditioned TI2V, the reference image determines the closest supported output ratio and is resized and center-cropped to match.

Architecture and Highlights

  • Prompt adherence: a dedicated text encoder turns prompts into precise generation conditions. An optional prompt enhancer can add visual, cinematic, temporal, and contextual detail, or be disabled to preserve the original wording.
  • Cinematic visual quality: training on carefully curated aesthetic data, with detailed annotations for lighting, composition, contrast, color grading, and more, enables precise control over cinematic styles and a wide range of aesthetic preferences.
  • Advanced motion generation: a larger and more diverse training corpus improves generalization across motion, semantics, and visual aesthetics, delivering leading performance among open-source and proprietary video generation models.
  • Efficient high-definition generation: the lightweight variant uses 32×32×4 spatiotemporal compression and supports T2V and I2V at 720p and 24 fps.

Known Limitations

  • Output videos have no audio track. Music or dialogue descriptions do not generate sound.
  • I2V and image-conditioned TI2V reference images are resized and center-cropped to the selected output ratio. Check that the subject remains within the retained area.
  • The generated video's first frame is not guaranteed to match the reference image pixel for pixel. First-frame conditioning operates in latent space, and preprocessing, VAE reconstruction, and video export can alter the pixels.
Vimoo© 2026 Vimoo Team. All rights reserved.