IIT PATNA CV 1 at Sci-ImageMiner 2026 Task 1: LoRA-Enhanced Generative Vision–Language Models for Scientific Figure Panel Classification

Authors

DOI:

https://doi.org/10.52825/ocp.v10i.3599

Keywords:

Scientific Figure Classification, Vision-Language Models, Low-Rank Adaptation, Generative Classification

Abstract

This paper presents a solution to the Sci-ImageMiner Competition: Task 1 (classification) at ICDAR 2026. The objective is to accurately classify individual scientific figure panels from Atomic Layer Deposition and Etching (ALD/E) literature into a domain-specific taxonomy. We formulate this challenge as a conditional text generation problem rather than a traditional discriminative classification task. By employing Low Rank Adaptation (LoRA), we efficiently fine-tune the Qwen2.5-VL-7B-Instruct vision language model to directly generate taxonomy-aware class labels in natural language, conditioned on cropped panel images and a taxonomy-aware structured prompt. This generative approach retains the model’s pre-trained scientific visual reasoning capabilities without requiring architectural modifications. Our proposed system demonstrates highly competitive performance, achieving a micro F1 score of 0.77 and securing 3rd place on the official evaluation leaderboard.

Downloads

Download data is not yet available.

References

[1] Q. Team, and others, "Qwen2.5 Technical Report", ArXiv, 2024.

[2] E. J. Hu, and others, "LoRA: Low-rank adaptation of large language models.", ICLR, vol. 1, no. 2, p. 3, 2022.

[3] Z. Liu, and others, "A convnet for the 2020s", in CVPR, 2022, pp. 11976–11986.

[4] S. Zhang, and others, "BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs", arXiv preprint arXiv:2303.00915, 2023.

[5] P. Wang, and others, "Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution", arXiv preprint arXiv:2409.12191, 2024.

[6] J. Roberts, and others, "SciFiBench: Benchmarking large multimodal models for scientific figure interpretation", Advances in Neural Information Processing Systems, vol. 37, pp. 18695–18728, 2024.

[7] E. Borisova, and others, "SciVQA 2025: Overview of the first scientific visual question answering shared task", in Proc. of the Fifth Workshop on Scholarly Document Processing, 2025, pp. 182–210.

[8] S. Yin, and others, "A survey on multimodal large language models", National Sci. Review, vol. 11, no. 12, 2024.

[9] L. M. Allen, and others, "Publishing: Credit where credit is due", Nature, vol. 508, no. 7496, pp. 312–313, 2014. DOI: 10.1038/508312a.

Downloads

Published

2026-09-01

How to Cite

Mohanta, S., Kasyap, S., Shashwat, S., Nareti, U. K., Akram, S., & Adak, C. (2026). IIT PATNA CV 1 at Sci-ImageMiner 2026 Task 1: LoRA-Enhanced Generative Vision–Language Models for Scientific Figure Panel Classification. Open Conference Proceedings, 10. https://doi.org/10.52825/ocp.v10i.3599