On October 17, news emerged that, as per the official website of HuggingFace, Baidu's independently developed multimodal document analysis model, PaddleOCR-VL, which was unveiled on the evening of October 16, swiftly climbed to the pinnacle of the HuggingFace Trending global leaderboard within a mere 20 hours post-release. Boasting a compact set of core parameters totaling just 0.9 billion, this lightweight yet highly efficient model excels in accurately identifying intricate elements including text, handwritten Chinese characters, tables, mathematical formulas, and charts—all with minimal computational demands. It supports an impressive 109 languages.
On the authoritative OmniBenchDoc V1.5 benchmark, PaddleOCR-VL attained a comprehensive performance score of 92.6, securing the top spot globally. All four of its core capabilities have achieved SOTA (State-of-the-Art) performance, outperforming models such as GPT-4o and setting a new benchmark for OCR VL model efficacy. As an offshoot of Wenxin 4.5, PaddleOCR-VL seamlessly integrates the NaViT dynamic resolution visual encoder with the ERNIE-4.5-0.3B language model, marking significant advancements in both precision and operational efficiency.
