A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

Authors

DOI:

https://doi.org/10.52825/ocp.v10i.3591

Keywords:

Scientific Images, Multimodal AI, Vision-Language Models, Digital Libraries,, Benchmarks, Scientific Reasoning

Abstract

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose “scientific conceptual understanding from images” as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.

Downloads

Download data is not yet available.

References

[1] N. Alampara, M. Schilling-Wilhelmi, M. Ríos-García et al., "Probing the limitations of multimodal language models for chemistry and materials research", Nature computational science, pp. 1–10, 2025.

[2] F. Ahmed, S. Auer, and J. D'Souza, "ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures", arXiv preprint arXiv:2607.26848, 2026. [Online]. Available: https://arxiv.org/abs/2607.26848.

[3] D. Circi, M. Bradley, S. Blouir et al., "Information Extraction from Diverse Charts In Materials Science", in LLM for Scientific Discovery: Reasoning, Assistance, and Collaboration.

[4] H. Hussein, F. Ahmed, A. Oelen, R. Ewerth, and S. Auer, "ORKGEx: Leveraging Language and Vision Models with Knowledge Graphs for Research Contribution Annotation", in Proc. of the Thirteenth Int. Conf. on Building and Exploring Web Based Environments (WEB 2025), P. Dini, Ed., Lisbon, Portugal: IARIA, Mar. 2025. ISBN: 978-1-68558-243-2, ISSN: 2308-4421.

[5] S. E. Kahou, V. Michalski, A. Atkinson, Á. Kádár, A. Trischler, and Y. Bengio, "Figureqa: An annotated figure dataset for visual reasoning", arXiv preprint arXiv:1710.07300, 2017.

[6] K. Kafle, B. Price, S. Cohen, and C. Kanan, "Dvqa: Understanding data visualizations via question answering", in Proc. of the IEEE conference on computer vision and pattern recognition, 2018, pp. 5648–5656.

[7] N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, "Plotqa: Reasoning over scientific plots", in Proc. of the ieee/cvf winter conference on applications of computer vision, 2020, pp. 1527–1536.

[8] R. Xia, H. Peng, H. Ye et al., "StructChart: On the schema, metric, and augmentation for visual chart understanding", arXiv e-prints, pp. arXiv–2309, 2023.

[9] A. Masry, X. L. Do, J. Q. Tan, S. Joty, and E. Hoque, "ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning", in Findings of the Assoc. for Computational Linguistics: ACL 2022, 2022, pp. 2263–2279.

[10] M. Jiang, J. Gao, J. Zhan, and D. Wang, "MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding", in Proc. of the Conf. on Language Modeling (COLM), 2025. [Online]. Available: https://arxiv.org/abs/2508.15802.

[11] D. Campbell, S. Rane, T. Giallanza et al., "Understanding the limits of vision language models through the lens of the binding problem", Advances in Neural Information Processing Systems, vol. 37, pp. 113436–113460, 2024.

[12] Z. Qi, Y. Fang, M. Zhang et al., "Gemini vs GPT-4V: A Preliminary Comparison and Combination of Vision-Language Models Through Qualitative Cases", CoRR, 2023.

[13] Y. Wu, L. Yan, L. Shen, Y. Wang, N. Tang, and Y. Luo, "Chartinsights: Evaluating multimodal large language models for low-level chart question answering", arXiv preprint arXiv:2405.07001, 2024.

[14] H. Pan, Q. Zhang, C. Caragea, E. Dragut, and L. Jan Latecki, "FlowLearn: Evaluating Large Vision-Language Models on Flowchart Understanding", in ECAI 2024, IOS Press, 2024, pp. 73–80.

[15] H. Liu, C. Li, Q. Wu, and Y. J. Lee, "Visual instruction tuning", Advances in neural information processing systems, vol. 36, pp. 34892–34916, 2023.

[16] C. Li, C. Wong, S. Zhang et al., "Llava-med: Training a large language-and-vision assistant for biomedicine in one day", Advances in Neural Information Processing Systems, vol. 36, pp. 28541–28564, 2023.

[17] S. Auer, A. Oelen, M. Haris et al., "Improving access to scientific literature with knowledge graphs", Bibliothek Forschung und Praxis, vol. 44, no. 3, pp. 516–529, 2020.

[18] D. R. Krathwohl, "A revision of Bloom's taxonomy: An overview", Theory into practice, vol. 41, no. 4, pp. 212–218, 2002.

[19] B. Wang, C. Xu, X. Zhao et al., "Mineru: An open-source solution for precise document content extraction", arXiv preprint arXiv:2409.18839, 2024.

[20] M. Mattinen, P. J. King, G. Popov et al., "Van der Waals epitaxy of continuous thin films of 2D materials using atomic layer deposition in low temperature and low vacuum conditions", 2D Materials, vol. 7, no. 1, p. 011003, 2020.

[21] A. Ghazy, J. Ylönen, N. Subramaniyam, and M. Karppinen, "Atomic/molecular layer deposition of europium–organic thin films on nanoplasmonic structures towards FRET-based applications", Nanoscale, vol. 15, no. 38, pp. 15865–15870, 2023.

[22] A. Andonian, S. G. Rodriques, A. D. White, and S. M. Narayanan, "MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition", arXiv preprint arXiv:2607.28532, 2026.

[23] S. Eger, Y. Cao, J. D'Souza et al., "Transforming science with large language models: A survey on ai-assisted scientific discovery, experimentation, content generation, and evaluation", arXiv preprint arXiv:2502.05151, 2025.

[24] M. Turpin, J. Michael, E. Perez, and S. Bowman, "Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting", Advances in Neural Information Processing Systems, vol. 36, pp. 74952–74965, 2023.

[25] V. Baherwani, T. Goldstein, and A. Panda, "Not All LLM Reasoning is Visible in the Chain-of-Thought", arXiv preprint arXiv:2607.22925, 2026.

[26] M. Ríos-García, N. Alampara, C. Gupta et al., "AI scientists produce results without reasoning scientifically", arXiv preprint arXiv:2604.18805, 2026.

[27] P. Kirgis, S. Kapoor, A. Schwartz et al., "Can AI agents conduct open-ended AI research? Early evidence from two case studies", arXiv preprint arXiv:2607.27191, 2026.

Downloads

Published

2026-09-01

How to Cite

D’Souza, J., Ahmed, F., Andrade, C. A. B., Frolova, L., Gnanasambandan, P., Hussain, D., … van Roeden, T. F. J. (2026). A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images. Open Conference Proceedings, 10. https://doi.org/10.52825/ocp.v10i.3591

Funding data