Compare AIFind AIAI NewsAI How-To
About Us
PrivacyTermsFAQContactContact
AIB Inc.Company info
© 2026 AIB Inc.

ByteDance Research Introduces Lance Multimodal Model

ByteDance Research Introduces Lance Multimodal Model

HuggingFace·Wednesday, May 20, 2026
  • •ByteDance Research introduced Lance, a native unified multimodal model for image and video tasks.
  • •The 3B parameter model uses a dual-stream mixture-of-experts architecture trained on 128-A100 GPUs.
  • •Lance enables integrated understanding, generation, and editing performance by decoupling capability pathways via multi-task training.
  • •ByteDance Research introduced Lance, a native unified multimodal model for image and video tasks.
  • •The 3B parameter model uses a dual-stream mixture-of-experts architecture trained on 128-A100 GPUs.
  • •Lance enables integrated understanding, generation, and editing performance by decoupling capability pathways via multi-task training.

ByteDance Research released Lance on May 18, a unified multimodal model capable of understanding, generating, and editing both images and videos. The architecture is trained from scratch using a dual-stream mixture-of-experts (MoE) framework, which utilizes shared interleaved multimodal sequences. This design enables the system to perform joint context learning while keeping pathways for understanding and generation tasks decoupled. To manage heterogeneous visual tokens and improve cross-task alignment, the model incorporates modality-aware rotary positional encoding.

Lance operates efficiently at a scale of 3 billion active parameters. The training process followed a staged multi-task paradigm, utilizing capability-oriented objectives alongside adaptive data scheduling. This approach reportedly enhances both semantic comprehension and visual output quality. Developers completed the training within a budget of 128-A100 GPUs. Experimental results indicate that Lance outperforms existing open-source unified models in image and video generation benchmarks while maintaining strong multimodal understanding capabilities.

ByteDance Research released Lance on May 18, a unified multimodal model capable of understanding, generating, and editing both images and videos. The architecture is trained from scratch using a dual-stream mixture-of-experts (MoE) framework, which utilizes shared interleaved multimodal sequences. This design enables the system to perform joint context learning while keeping pathways for understanding and generation tasks decoupled. To manage heterogeneous visual tokens and improve cross-task alignment, the model incorporates modality-aware rotary positional encoding.

Lance operates efficiently at a scale of 3 billion active parameters. The training process followed a staged multi-task paradigm, utilizing capability-oriented objectives alongside adaptive data scheduling. This approach reportedly enhances both semantic comprehension and visual output quality. Developers completed the training within a budget of 128-A100 GPUs. Experimental results indicate that Lance outperforms existing open-source unified models in image and video generation benchmarks while maintaining strong multimodal understanding capabilities.

Read original (English)·May 20, 2026
#multimodal#bytedance#lance#image generation#video generation#mixture of experts