DeepSeek Unveils New 3B OCR Model: A Groundbreaking Leap in Streamlined Document Analysis
2025-10-21 / Read about 0 minute
Author:小编   

An AI technology firm has just launched its latest optical character recognition innovation, DeepSeek-OCR. This cutting-edge, end-to-end vision-language model is tailored for high-efficiency document parsing. During the Fox benchmark assessment, the model attained an impressive decoding accuracy rate of up to 97%. Remarkably, even when subjected to substantial compression, it continues to deliver top-notch performance. Additionally, it excels in the OmniDocBench benchmark test.

The model's architecture is a blend of a visual encoder, known as DeepEncoder, and a mixture-of-experts decoder. DeepEncoder leverages specialized mechanisms and algorithms, accommodating diverse modes like Tiny and Small, along with a range of resolution settings. A standout feature is its dynamic mode, which enables on-the-fly adjustments to token budgets, enhancing its versatility.

Training the model follows a phased approach. For real-world applications, it's advisable to begin with the Small mode, whereas the Gundam mode is the go-to choice for intricate scenarios. The introduction of the DeepSeek-OCR model heralds a major stride in document AI, distinguished by its efficiency, adaptability, and robust performance.