VLMinators at Sci-ImageMiner 2026 Tasks 1–4: QLoRA Fine-Tuning of Qwen2.5-VL With Selective Label Disambiguation and Cross-Task Context Injection
DOI:
https://doi.org/10.52825/ocp.v10i.3590Keywords:
Scientific Figure Understanding, Vision-Language Models, QLoRA, Chart Classification, Data Extraction, Summarization, Visual Question Answering, ALD/ALE, Selective PromptingAbstract
We describe our system for the ICDAR 2026 Sci-ImageMiner competition, covering figure panel classification (Task 1), data extraction (Task 2), summarization (Task 3), and visual question answering (Task 4) on scientific figures from atomic layer
deposition and etching literature. For Task 1, we progress from CNN and zero-shot VLM baselines through QLoRA fine-tuning of Qwen2.5-VL-7B-Instruct, finding that short disambiguation descriptions injected only for the most confusable class pairs improved fine-grained classification without degrading well-separated classes; a six model weighted ensemble achieved F1=0.8059 on the CodaBench evaluation phase. For Tasks 2–4, we fine-tune Qwen2.5-VL independently for each task and additionally propose a cross-task context-injection pipeline that prepends the structured outputs of earlier tasks to downstream prompts at inference time, grounding summarizations and VQA answers in explicitly extracted values rather than visual estimation, without additional training.
Downloads
References
[1] S. Organizers, "Sci-ImageMiner 2026: Overview of the Scientific Image Mining Challenge at ICDAR 2026", in Proc. of ICDAR 2026, 2026.
[2] K. V. Jobin, A. Mondal, and C. V. Jawahar, "DocFigure: A Dataset for Scientific Document Figure Classification", in Proc. of ICDAR 2019, 2019. DOI: 10.1109/ICDARW.2019.00018.
[3] Z. Karishma, S. Rohatgi, K. S. Puranik, J. Wu, and C. L. Giles, "ACL-Fig: A Dataset for Scientific Figure Classification", in Proc. of the Workshop on Scientific Document Understanding co-located with the 37th AAAI Conf. on Artificial Intelligence (AAAI 2023), ser. CEUR Workshop Proc., vol. 3656, 2023. [Online]. Available: https://ceur-ws.org/Vol-3656/paper2.pdf.
[4] S. E. Kahou, V. Michalski, A. Atkinson, and others, "FigureQA: An Annotated Figure Dataset for Visual Reasoning", arXiv preprint arXiv:1710.07300, 2018. [Online]. Available: https://arxiv.org/abs/1710.07300.
[5] K. Kafle, B. Price, S. Cohen, and C. Kanan, "DVQA: Understanding Data Visualizations via Question Answering", in Proc. of CVPR 2018, 2018. DOI: 10.1109/CVPR.2018.00592.
[6] A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque, "ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning", in Findings of the Assoc. for Computational Linguistics: ACL 2022, 2022, pp. 2263–2279. DOI: 10.18653/v1/2022.findings-acl.177. [Online]. Available: https://aclanthology.org/2022.findings-acl.177/.
[7] H. Liu, C. Li, Y. Li, and Y. J. Lee, "Improved Baselines with Visual Instruction Tuning", in Proc. of CVPR 2024, 2024. DOI: 10.1109/CVPR52733.2024.02484.
[8] H. Liu, C. Li, Y. Li et al., "Llava-next: Improved reasoning, ocr, and world knowledge, January 2024", 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next.
[9] S. Bai, K. Chen, X. Liu, and others, "Qwen2.5-VL Technical Report", arXiv preprint arXiv:2502.13923, 2025. [Online]. Available: https://arxiv.org/abs/2502.13923.
[10] M. Abdin, J. Aneja, H. Awadalla, and others, "Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone", arXiv preprint arXiv:2404.14219, 2024. [Online]. Available: https://arxiv.org/abs/2404.14219.
[11] A. Grattafiori, A. Dubey, A. Jauhri et al., "The Llama 3 herd of models", in Neural Information Processing Systems, 2024.
[12] E. J. Hu, Y. Shen, P. Wallis et al., "LoRA: Low-Rank Adaptation of Large Language Models", in Proc. of ICLR 2022, 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9.
[13] T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, "QLoRA: Efficient Finetuning of Quantized LLMs", in Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/hash/1feb87871436031bdc0f2beaa62a049b-Abstract-Conference.html.
[14] S. Mangrulkar, S. Gugger, L. Debut, and others, "PEFT: State-of-the-Art Parameter-Efficient Fine-Tuning", Hugging Face, 2022. [Online]. Available: https://github.com/huggingface/peft.
[15] S. Menon, and C. Vondrick, "Visual Classification via Description from Large Language Models", in Proc. of ICLR 2023, 2023. [Online]. Available: https://openreview.net/forum?id=jlAjNL8z5cs.
[16] W. Dai, J. Li, D. Li, and others, "InstructBLIP: Towards General-Purpose Vision-Language Models with Instruction Tuning", in Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/hash/9a6a435e75419a836fe47ab6793623e6-Abstract-Conference.html.
[17] M. Tan, and Q. V. Le, "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks", in Proc. of ICML 2019, 2019. [Online]. Available: http://proceedings.mlr.press/v97/tan19a.html.
[18] Z. Liu, H. Hu, Y. Lin, and others, "Swin Transformer V2: Scaling Up Capacity and Resolution", in Proc. of CVPR 2022, 2022. DOI: 10.1109/CVPR52688.2022.01170.
[19] C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi, "Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning", in Proc. of AAAI 2017, 2017. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/11231.
[20] R. Wightman, "PyTorch Image Models", https://github.com/rwightman/pytorch-image-models, GitHub, 2019. DOI: 10.5281/zenodo.4414861.
[21] J. Ezemba, C. McComb, and C. Tucker, "Simulation vs. hallucination: Assessing vision-language model question answering capabilities in engineering simulations", in Proc. of the 7th Workshop on Design Automation for CPS and IoT, 2025, pp. 1–9. DOI: 10.1145/3722573.3727826.
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Tirtha Vinchurkar, Rithika Sai Konakalla, Kareem Abdelmaqsoud

This work is licensed under a Creative Commons Attribution 4.0 International License.