DeepVitminC at Sci-ImageMiner 2026 Tasks 3&4: Structured Context Augmentation for Scientific Figure Reasoning

Authors

DOI:

https://doi.org/10.52825/ocp.v10i.3588

Keywords:

Multimodal Large Language Model, Scientific Figure Understanding, Context Augmentation, Visual Question Answering, Summarization

Abstract

Scientific figures encode critical experimental findings that are often not fully captured in textual descriptions. Despite recent advances in MLLMs, their performance on domain-specific scientific figures remains limited due to insufficient grounding in relevant context. In this work, we focus on in Task 3 (summarization) and Task 4 (visual question answering) of the Sci-ImageMiner 2026 challenge. We propose a structured context augmentation approach that improves model performance by providing relevant and well-organized contextual inputs. Our method extracts figure-centric and cross-referenced textual evidence from scientific papers and integrates them into structured prompts for model training. Experimental results show that our approach consistently improves performance across both tasks. We achieve first place in both Task 3 and Task 4, demonstrating the effectiveness of structured contextual inputs for scientific figure reasoning. Code is available at https://github.com/yuanjun0416/SciImageMiner34

Downloads

Download data is not yet available.

References

[1] X. Yue et al., “Instruction-augmented multimodal alignment for image-text and element matching,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 1379–1388. DOI: 10.1109/CVPRW67362.2025.00127.

[2] H. Fu, Y. Ouyang, K.-W. Chang, Y. Wang, Z. Huang, and Y. Cai, “Contextnav: Towards agentic multimodal in-context learning,” arXiv preprint arXiv:2510.04560, 2025. DOI: 10. 48550/arXiv.2510.04560.

[3] OpenAI, Gpt-4v(ision) system card, https://openai.com/research/gpt-4v-system card, Accessed: 2026-04-14, 2023.

[4] G. Team, “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context,” arXiv preprint arXiv:2403.05530, 2024. DOI: 10.48550/arXiv.2403.05530.

[5] B. Y. Lin et al., “The unlocking spell on base llms: Rethinking alignment via in-context learning,” in International Conference on Learning Representations (ICLR), 2024. [On line]. Available: https://proceedings.iclr.cc/paper_files/paper/2024/file/6bcbb4a501dbad0eba1b660c1a55318c-Paper-Conference.pdf.

[6] OpenAI. “Toward understanding and preventing emergent misalignment,” Accessed: Apr. 14, 2026. [Online]. Available: https://openai.com/index/emergentmisalignment/.

[7] Y. Hu, Q.Li, D. Zhang, J. Yan, and Y. Chen, “Context-alignment: Activating and enhancing llm capabilities in time series,” arXiv preprint arXiv:2501.03747, 2026. DOI: 10.48550/arXiv.2501.03747

[8] E. J. Hu et al., “Lora: Low-rank adaptation of large language models,” arXiv preprint arXiv:2106.09685, 2021. DOI: 10.48550/arXiv.2106.09685.

[9] ModelScope Team. “Ms-swift: Efficient llm training and inference framework,” Accessed: Apr. 14, 2026. [Online]. Available: https://github.com/modelscope/ms-swift.

[10] Qwen Team. “Qwen3.5,” Accessed: Apr. 14, 2026. [Online]. Available: https://qwen.ai/blog?id=qwen3.5

Downloads

Published

2026-09-01

How to Cite

Lv, X., Wu, W., Tan, Y., Wang, W., Zhao, Y., Liu, Y., … Gu, H. (2026). DeepVitminC at Sci-ImageMiner 2026 Tasks 3&4: Structured Context Augmentation for Scientific Figure Reasoning. Open Conference Proceedings, 10. https://doi.org/10.52825/ocp.v10i.3588