IIT PATNA CV 1 at Sci-ImageMiner 2026 Task 1: LoRA-Enhanced Generative Vision–Language Models for Scientific Figure Panel Classification
DOI:
https://doi.org/10.52825/ocp.v10i.3599Keywords:
Scientific Figure Classification, Vision-Language Models, Low-Rank Adaptation, Generative ClassificationAbstract
This paper presents a solution to the Sci-ImageMiner Competition: Task 1 (classification) at ICDAR 2026. The objective is to accurately classify individual scientific figure panels from Atomic Layer Deposition and Etching (ALD/E) literature into a domain-specific taxonomy. We formulate this challenge as a conditional text generation problem rather than a traditional discriminative classification task. By employing Low Rank Adaptation (LoRA), we efficiently fine-tune the Qwen2.5-VL-7B-Instruct vision language model to directly generate taxonomy-aware class labels in natural language, conditioned on cropped panel images and a taxonomy-aware structured prompt. This generative approach retains the model’s pre-trained scientific visual reasoning capabilities without requiring architectural modifications. Our proposed system demonstrates highly competitive performance, achieving a micro F1 score of 0.77 and securing 3rd place on the official evaluation leaderboard.
Downloads
References
[1] Q. Team, and others, "Qwen2.5 Technical Report", ArXiv, 2024.
[2] E. J. Hu, and others, "LoRA: Low-rank adaptation of large language models.", ICLR, vol. 1, no. 2, p. 3, 2022.
[3] Z. Liu, and others, "A convnet for the 2020s", in CVPR, 2022, pp. 11976–11986.
[4] S. Zhang, and others, "BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs", arXiv preprint arXiv:2303.00915, 2023.
[5] P. Wang, and others, "Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution", arXiv preprint arXiv:2409.12191, 2024.
[6] J. Roberts, and others, "SciFiBench: Benchmarking large multimodal models for scientific figure interpretation", Advances in Neural Information Processing Systems, vol. 37, pp. 18695–18728, 2024.
[7] E. Borisova, and others, "SciVQA 2025: Overview of the first scientific visual question answering shared task", in Proc. of the Fifth Workshop on Scholarly Document Processing, 2025, pp. 182–210.
[8] S. Yin, and others, "A survey on multimodal large language models", National Sci. Review, vol. 11, no. 12, 2024.
[9] L. M. Allen, and others, "Publishing: Credit where credit is due", Nature, vol. 508, no. 7496, pp. 312–313, 2014. DOI: 10.1038/508312a.
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Soumyajyoti Mohanta, Shresth Kasyap, Sasmit Shashwat, Utsav Kumar Nareti, Saba Akram, Chandranath Adak

This work is licensed under a Creative Commons Attribution 4.0 International License.