Skip to documentation
On this page

Open Source Model

Overview

Vimoo-1.0 is an open-source video generation model built on the MMDiT architecture. It supports text-to-video (T2V) and image-to-video (I2V), using prompts or reference images to create videos with control over subjects, actions, scenes, and camera movement.

Model Capabilities

  • Text and image inputs: create a scene from a description or animate a reference image.
  • HD video generation: generate 720p video at 24 fps, with 5–10-second outputs.
  • In-video text conditioning: generate specified text within the scene.
  • Multiple model variants: Vimoo-1.0-Distilled adds a distilled option. The planned Lite TI2V release uses one base for text-to-video and image-to-video, with a shared distilled adapter and an INT8 base option.

See Vimoo-1.0 for architecture, variants, and capabilities. Find weights and matching LoRAs in Resources.

Get Started

Vimoo provides model weights, inference examples, and ComfyUI workflows for local creation, programmatic integration, and research.

Open-Source Resources

ResourceEntry
Model and inference examplesVimoo-1.0Coming soon
Model weightsVimoo-AI on Hugging Face

See Local Deployment for hardware, device timings, and local generation examples.

Vimoo© 2026 Vimoo Team. All rights reserved.