On October 16, Baidu officially released the open-source version of its self-developed multimodal document analysis model, PaddleOCR-VL. This model has clinched the top spot in comprehensive performance, achieving an impressive score of 92.6 on the globally recognized document analysis evaluation benchmark, OmniBenchDoc V1.5. Renowned for its lightweight design and high efficiency, the model boasts a mere 0.9 billion core model parameters. It excels in accurately recognizing intricate elements such as text, handwritten Chinese characters, tables, formulas, and charts, all while maintaining extremely low computational demands. Furthermore, the model supports an extensive range of 109 languages, encompassing multilingual scenarios like Chinese, English, French, Japanese, Russian, Arabic, and Spanish. Its versatility makes it an ideal choice for a wide array of document intelligence tasks, including government and enterprise document management, knowledge retrieval, archive digitization, and scientific research information extraction.
