Team TeleOCR-VL’s Solution for Sci-ImageMiner 2026: Competition on Scientific Chart Understanding at ICDAR 2026

Authors

DOI:

https://doi.org/10.52825/ocp.v10i.3589

Keywords:

Multimodal OCR, cientific Chart Understanding,, tructure-Aware Modeling, Data Table Extraction, Chart Summary Generation

Abstract

This report presents Team TeleOCR-VL’s system for Tasks 2 and 3 of the ICDAR 2026 Sci-ImageMiner Competition. Sci-ImageMiner focuses on scientific chart understanding in real ALD/ALE research papers, requiring systems to extract structured data from complex figures and generate summaries consistent with scientific context. To address the challenges of complex layouts, dense numerical information, strong structural dependencies, and domain specific visual patterns, we build a structure-aware multimodal OCR system based on TeleOCR-VL. The system combines large-scale synthetic data augmentation, structure-content decoupled training, a vision-plus-context modeling strategy, and multi-model voting. Experimental results show that our system achieves aweighted score of 41.81 in Task 2 Data TableExtraction, ranking first, and a weighted score of 0.54 in Task 3 Chart Summary Generation, ranking third. These
results indicate that a specialized OCR-VL system designed for scientific chart understanding is more stable than general-purpose vision-language models in table structure recovery, numerical extraction, and chart summary generation.

Downloads

Download data is not yet available.

References

[1] S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, and others, "Qwen3-VL Technical Report", arXiv preprint arXiv:2511.21631, 2025.

[2] H. Liu, C. Li, Y. Li et al., "LLaVA-NeXT: Improved Reasoning, OCR, and World Knowledge", Technical Report, 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/.

[3] J. D. Hunter, "Matplotlib: A 2D Graphics Environment", Computing in Sci. & Eng., vol. 9, no. 3, pp. 90–95, 2007. DOI: 10.1109/MCSE.2007.55.

[4] A. Nassar, A. Marafioti, M. Omenetti et al., "SmolDocling: An Ultra-Compact Vision-Language Model for End-to-End Multi-Modal Document Conversion", arXiv preprint arXiv:2503.11576, 2025.

[5] J. Devlin, M. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding", in Proc. of the 2019 Conf. of the North American Chapter of the Assoc. for Computational Linguistics: Human Language Technologies, 2019, pp. 4171–4186. DOI: 10.18653/v1/N19-1423.

[6] I. Loshchilov, and F. Hutter, "Decoupled Weight Decay Regularization", in Int. Conf. on Learning Representations, 2019.

Downloads

Published

2026-09-01

How to Cite

Wang, Y., Cai, P., Zou, Z., & Sun, H. (2026). Team TeleOCR-VL’s Solution for Sci-ImageMiner 2026: Competition on Scientific Chart Understanding at ICDAR 2026. Open Conference Proceedings, 10. https://doi.org/10.52825/ocp.v10i.3589