Alibaba Cloud has officially announced the launch of the DeepSeek-V4.1-Flash model on its QianWen AI platform, complete with API services and a Token Plan that are now accessible. Developers have the flexibility to seamlessly integrate this model into their applications through APIs or directly utilize it within tools like Qoder and the QianWen APP. This versatile model excels in both text and image comprehension, boasting a remarkable maximum context length of 1 million Tokens. This feature renders it particularly well-suited for tasks such as code development and the handling of lengthy documents. Leveraging advanced caching and compression technologies, the model significantly diminishes its reliance on HBM and SSD resources, thereby enhancing its suitability for scenarios involving high-frequency calls. The deployment of this model is a collaborative effort between Alibaba Cloud Bailian and the vLLM open-source community.
