LongCat-Video
LongCat-Video is a 13.6B model for text-to-video, image-to-video, and video continuation. It uses a coarse-to-fine approach to generate high-quality, long videos.

Introduction
LongCat-Video is a 13.6B-parameter foundation video generation model developed by Meituan's "Long Cat Team." It supports three core tasks: text-to-video, image-to-video, and video continuation. The model natively generates minute-level long videos without color drift or quality degradation during generation. Built on a coarse-to-fine spatiotemporal generation strategy and block-sparse attention mechanism, it balances generation quality with inference efficiency, making it ideal for researchers and developers with high demands for long-form video generation.
Key Features
- Text-to-video: Generate high-quality dynamic videos from natural language descriptions
- Image-to-video: Create coherent motion from a single image while preserving original visual details
- Video continuation: Extend existing videos temporally while maintaining style and motion consistency
- Long video generation: Natively produce continuous videos lasting several minutes, suitable for narrative content creation
- Out-of-the-box multi-task capability with support for local deployment and GPU-accelerated inference
Highlights
- Industry-leading native long-video generation supporting 720p/30fps output at minute-level duration with no quality degradation
- Unified architecture covering three major video generation scenarios in a single model
- Optimized with multi-reward reinforcement learning (GRPO), delivering comprehensive performance comparable to mainstream open-source and commercial models
Who It's For
LongCat-Video is well-suited for short-video and AIGC content creators, teams working on automated content production, and professionals in gaming, film, and advertising who need dynamic asset generation and previsualization. Released under the MIT license, it is friendly for both academic and commercial use, making it a practical choice for teams looking to quickly integrate high-quality video generation capabilities. As AIGC evolves toward world models, LongCat-Video's long-sequence modeling provides forward-looking technical groundwork for developers seeking a competitive edge.






