Team TeleOCR-VL’s Solution for Sci-ImageMiner 2026: Competition on Scientific Chart Understanding at ICDAR 2026
DOI:
https://doi.org/10.52825/ocp.v10i.3589Keywords:
Multimodal OCR, cientific Chart Understanding,, tructure-Aware Modeling, Data Table Extraction, Chart Summary GenerationAbstract
This report presents Team TeleOCR-VL’s system for Tasks 2 and 3 of the ICDAR 2026 Sci-ImageMiner Competition. Sci-ImageMiner focuses on scientific chart understanding in real ALD/ALE research papers, requiring systems to extract structured data from complex figures and generate summaries consistent with scientific context. To address the challenges of complex layouts, dense numerical information, strong structural dependencies, and domain specific visual patterns, we build a structure-aware multimodal OCR system based on TeleOCR-VL. The system combines large-scale synthetic data augmentation, structure-content decoupled training, a vision-plus-context modeling strategy, and multi-model voting. Experimental results show that our system achieves aweighted score of 41.81 in Task 2 Data TableExtraction, ranking first, and a weighted score of 0.54 in Task 3 Chart Summary Generation, ranking third. These
results indicate that a specialized OCR-VL system designed for scientific chart understanding is more stable than general-purpose vision-language models in table structure recovery, numerical extraction, and chart summary generation.
Downloads
References
[1] S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, and others, "Qwen3-VL Technical Report", arXiv preprint arXiv:2511.21631, 2025.
[2] H. Liu, C. Li, Y. Li et al., "LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge", Technical Report, 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/.
[3] J. D. Hunter, "Matplotlib: A 2D Graphics Environment", Computing in Sci. & Eng., vol. 9, no. 3, pp. 90–95, 2007. DOI: 10.1109/MCSE.2007.55.
[4] A. Nassar, A. Marafioti, M. Omenetti et al., "SmolDocling: An Ultra-Compact Vision-Language Model for End-to-End Multi-Modal Document Conversion", arXiv preprint arXiv:2503.11576, 2025.
[5] J. Devlin, M. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding", in Proc. of the 2019 Conf. of the North American Chapter of the Assoc. for Computational Linguistics: Human Language Technologies, 2019, pp. 4171–4186. DOI: 10.18653/v1/N19-1423.
[6] I. Loshchilov, and F. Hutter, "Decoupled Weight Decay Regularization", in Int. Conf. on Learning Representations, 2019.
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Yikun Wang, Peng Cai, Zhaofan Zou, Hao Sun

This work is licensed under a Creative Commons Attribution 4.0 International License.