On August 12, Tencent Hunyuan introduced its cutting-edge multimodal understanding model, Hunyuan Large-Vision. Leveraging the Mixture of Experts (MoE) architecture, this model boasts an impressive 52 billion activation parameters, enabling it to support image, video, and 3D spatial inputs of any resolution. Additionally, Hunyuan Large-Vision significantly enhances comprehension capabilities in diverse language scenarios, showcasing a remarkable leap forward in artificial intelligence.
