Vassilis Sioros at Sci-ImageMiner 2026 Tasks 2, 3, and 4
Extract, Summarize, Answer: A Staged Qwen3.5 Pipeline for Scientific VQA
DOI:
https://doi.org/10.52825/ocp.v10i.3511Keywords:
Scientific Figure Understanding, Multimodal Learning, Visual Question Answering, Table Extraction, Summarization, Qwen3.5-9B, LoRAAbstract
In this work, we address two core challenges in the ALD-E domain: auxiliary evidence construction via summarization and data table extraction, and visual question answering (VQA). Our approach first generates auxiliary summaries and structured tables to support test-time VQA prompting, using Qwen3.5-9B vision-language models with parameter-efficient fine-tuning (LoRA). The model operates on cropped subfigures alongside sliding-window document context, including captions and neighboring text blocks. It then answers Yes/No, Paragraph, Factoid, and List questions under a shared training and inference setup. This staged design reveals three key observations: (1) summarization performs strongly across both standard and contextual evaluation metrics; (2) table extraction achieves competitive structural accuracy; and (3) overall VQA performance is moderate. Yes/No questions appear largely saturated, while Factoid and List questions remain the primary bottlenecks. Overall, our findings highlight the effectiveness of staged multimodal evidence integration and answer-type-aware optimization. Code is available at https://github.com/insane-group/staged-qwen3.5-scivqa.
Downloads
References
[1] N. Methani, P. Ganguly, M. M. Khapra, and P. Kumar, “PlotQA: Reasoning over Scientific Plots,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2020, pp. 1527–1536. Accessed: Apr. 7, 2026.
[2] A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque, “ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning,” in Findings of the Association for Computational Linguistics: ACL 2022, S. Muresan, P. Nakov, and A. V. Sioros | Open Conf Proc X (2026) ”Sci-ImageMiner 2026” Villavicencio, Eds., Dublin, Ireland: Association for Computational Linguistics, Feb. 2022, pp. 2263–2279. DOI: 10.18653/v1/2022.findings-acl.177 Accessed: Apr. 7, 2026.
[3] E. Borisova, N. Rauscher, and G. Rehm, “SciVQA 2025: Overview of the First Scientific Visual Question Answering Shared Task,” in Proceedings of the Fifth Workshop on Scholarly Document Processing (SDP 2025), Vienna, Austria: Association for Computational Linguistics, 2025, pp. 182–210. DOI: 10.18653/v1/2025.sdp-1.18 Accessed: Mar. 11, 2026.
[4] F. Liu et al., “DePlot: One-shot visual language reasoning by plot-to-table translation, ”in Findings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds., Toronto, Canada: Association for Computational Linguistics, Apr. 2023, pp. 10381–10399. DOI: 10.18653/v1/2023.findings-acl.660 Accessed: Apr. 7, 2026.
[5] H. Singh and S. Shekhar, “STL-CQA: Structure-based Transformers with Localization and Encoding for Chart Question Answering,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online: Association for Computational Linguistics, 2020, pp. 3275–3284. DOI: 10.18653/v1/2020.emnlp main.264 Accessed: Mar. 12, 2026.
[6] M. Levy, R. Ben-Ari, and D. Lischinski, “Classification-Regression for Chart Comprehension,” in Computer Vision– ECCV 2022, S. Avidan, G. Brostow, M. Ciss´ e, G. M. Farinella, and T. Hassner, Eds., vol. 13696, Cham: Springer Nature Switzerland, 2022, pp. 469 484, ISBN: 978-3-031-20058-8 978-3-031-20059-5. DOI: 10.1007/978-3-031-200595_27 Accessed: Mar. 11, 2026.
[7] K.Yi, J. Wu, C. Gan,A.Torralba, P. Kohli, and J. Tenenbaum, “Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding,” in Advances in Neural Information Processing Systems, vol. 31, Curran Associates, Inc., 2018. Accessed: Mar. 12, 2026.
[8] K. Sanders, N. Weir, and B. Van Durme, “TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA: Association for Computational Linguistics, 2024, pp. 19009–19028. DOI: 10.18653/v1/2024.emnlp-main.1059 Accessed: Dec. 18, 2024.
[9] L. Pan, A. Albalak, X. Wang, and W. Wang, “Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore: Association for Computational Linguistics, 2023, pp. 3806–3824. DOI: 10.18653/v1/2023.findings-emnlp.248 Accessed: Mar. 12, 2026.
[10] R. R. Putra, R. S. P. Basuki, Y. Cheng, and P. Gao, Nl2logic: Ast-guided translation of natural language into first-order logic with large language models, 2026. arXiv: 2602. 13237 [cs.AI]. [Online]. Available: https://arxiv.org/abs/2602.13237
[11] S. Han et al., “FOLIO: Natural Language Reasoning with First-Order Logic,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA: Association for Computational Linguistics, 2024, pp. 22017–22031. DOI: 10.18653/v1/2024.emnlp-main.1229 Accessed: Feb. 11, 2025.
[12] C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” in Text Summarization Branches Out, Barcelona, Spain: Association for Computational Linguistics, Apr. 2004, pp. 74–81. Accessed: Apr. 7, 2026.
[13] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, BERTScore: Evaluating Text Generation with BERT, Feb. 2020. DOI: 10.48550/arXiv.1904.09675 arXiv: 1904.09675[cs]. Accessed: Apr. 7, 2026.
[14] B. Smock, R. Pesala, and R. Abraham, PubTables-1M: Towards comprehensive table extraction from unstructured documents, Nov. 2021. DOI: 10.48550/arXiv.2110.00061, arXiv: 2110.00061 [cs]. Accessed: Apr. 7, 2026.
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Vassilis Sioros

This work is licensed under a Creative Commons Attribution 4.0 International License.