DeepVitminC at Sci-ImageMiner 2026 Tasks 3&4: Structured Context Augmentation for Scientific Figure Reasoning
DOI:
https://doi.org/10.52825/ocp.v10i.3588Keywords:
Multimodal Large Language Model, Scientific Figure Understanding, Context Augmentation, Visual Question Answering, SummarizationAbstract
Scientific figures encode critical experimental findings that are often not fully captured in textual descriptions. Despite recent advances in MLLMs, their performance on domain-specific scientific figures remains limited due to insufficient grounding in relevant context. In this work, we focus on in Task 3 (summarization) and Task 4 (visual question answering) of the Sci-ImageMiner 2026 challenge. We propose a structured context augmentation approach that improves model performance by providing relevant and well-organized contextual inputs. Our method extracts figure-centric and cross-referenced textual evidence from scientific papers and integrates them into structured prompts for model training. Experimental results show that our approach consistently improves performance across both tasks. We achieve first place in both Task 3 and Task 4, demonstrating the effectiveness of structured contextual inputs for scientific figure reasoning. Code is available at https://github.com/yuanjun0416/SciImageMiner34Downloads
References
[1] X. Yue et al., “Instruction-augmented multimodal alignment for image-text and element matching,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 1379–1388. DOI: 10.1109/CVPRW67362.2025.00127.
[2] H. Fu, Y. Ouyang, K.-W. Chang, Y. Wang, Z. Huang, and Y. Cai, “Contextnav: Towards agentic multimodal in-context learning,” arXiv preprint arXiv:2510.04560, 2025. DOI: 10. 48550/arXiv.2510.04560.
[3] OpenAI, Gpt-4v(ision) system card, https://openai.com/research/gpt-4v-system card, Accessed: 2026-04-14, 2023.
[4] G. Team, “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530, 2024. DOI: 10.48550/arXiv.2403.05530.
[5] B. Y. Lin et al., “The unlocking spell on base llms: Rethinking alignment via in-context learning,” in International Conference on Learning Representations (ICLR), 2024. [On line]. Available: https://proceedings.iclr.cc/paper_files/paper/2024/file/6bcbb4a501dbad0eba1b660c1a55318c-Paper-Conference.pdf.
[6] OpenAI. “Toward understanding and preventing emergent misalignment,” Accessed: Apr. 14, 2026. [Online]. Available: https://openai.com/index/emergentmisalignment/.
[7] Y. Hu, Q.Li, D. Zhang, J. Yan, and Y. Chen, “Context-alignment: Activating and enhancing llm capabilities in time series,” arXiv preprint arXiv:2501.03747, 2026. DOI: 10.48550/arXiv.2501.03747
[8] E. J. Hu et al., “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685, 2021. DOI: 10.48550/arXiv.2106.09685.
[9] ModelScope Team. “Ms-swift: Efficient llm training and inference framework,” Accessed: Apr. 14, 2026. [Online]. Available: https://github.com/modelscope/ms-swift.
[10] Qwen Team. “Qwen3.5,” Accessed: Apr. 14, 2026. [Online]. Available: https://qwen.ai/blog?id=qwen3.5
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Xiaole Lv, Wei Wu, Yang Tan, Wenjie Wang, Yuan Zhao, Ying Liu, Liang Diao, Haisong Gu

This work is licensed under a Creative Commons Attribution 4.0 International License.