China Telecom's Xingchen Lab has unveiled TeleOCR, a streamlined visual language model boasting approximately 120 million parameters. This versatile model is adept at uniformly processing both digitally-created and photographed documents, adeptly transforming text, layouts, tables, formulas, and scientific charts into well-structured outputs. In the OmniDocBench v1.6 assessment, TeleOCR secured the top position with an impressive overall score of 96.87, outperforming competitors like MinerU2.5-Pro and PaddleOCR-VL-1.6. Additionally, it claimed the pinnacle spot on the PureDocBench leaderboard and emerged victorious in the ICDAR 2026 Scientific Chart Parsing Challenge.
