[DocMiner] at Sci-ImageMiner 2026 (Tasks 1, 2, and 4): Skill-Based Agentic DocMiner
A Context-Enriched Multi-Agent System for Scientific Figure Understanding
DOI:
https://doi.org/10.52825/ocp.v10i.3510Keywords:
Scientific Figure Understanding, Multimodal Document Intelligence, Agentic Workflow, Chart Classification, Data Extraction, Visual Question AnsweringAbstract
We present Skill-Based Agentic DocMiner, our submission to Sci-ImageMiner 2026 for Task 1 (figure classification), Task 2 (data table extraction), and Task 4 (visual question answering). Our system follows a unified agentic workflow with three recurring components: multimodal context enrichment from captions and paper text, iterative skill refinement using development-set feedback, and collaborative verification by specialist reviewers. For classification, we use class prototype documents to distinguish visually similar figure types. For table extraction, we combine chart parsing with Markdown structure validation and content checking. For VQA, we route each question to a specialist agent aligned with the official reasoning categories. On the official leaderboard, our submission ranked fourth on classification, fifth on table extraction, and third on VQA. The framework is simple, modular, and well matched to structured scientific reasoning. Our code is publicly available at https://github.com/1gst/Sci-ImageMiner.
Downloads
References
[1] Sci-ImageMiner Organizing Committee. “Sci-imageminer– task description,” Accessed: Apr. 10, 2026. [Online]. Available: https://sites.google.com/view/sci-imageminer/task-description
[2] Google. “Gemini api documentation,” Accessed: Apr. 10, 2026. [Online]. Available: https://ai.google.dev/gemini-api/docs/models
[3] Alibaba Cloud. “Dashscope / qwen api documentation,” Accessed: Apr. 10, 2026. [Online]. Available: https://help.aliyun.com/zh/model-studio/getting-started/models
[4] S. Yao et al., “React: Synergizing reasoning and acting in language models,” arXiv preprint arXiv:2210.03629, 2022. DOI: 10.48550/arXiv.2210.03629 [Online]. Available: https://arxiv.org/abs/2210.03629
[5] T. Schick et al., “Toolformer: Language models can teach themselves to use tools,” arXiv preprint arXiv:2302.04761, 2023. DOI: 10.48550/arXiv.2302.04761 [Online]. Available: https://arxiv.org/abs/2302.04761
[6] Q. Wu et al., “Autogen: Enabling next-gen llm applications via multi-agent conversation,” arXiv preprint arXiv:2308.08155, 2023. DOI: 10.48550/arXiv.2308.08155 [Online]. Available: https://arxiv.org/abs/2308.08155
[7] Sci-ImageMiner Organizing Committee. “Sci-imageminer–task evaluation metrics,” Accessed: Apr. 10, 2026. [Online]. Available: https://sites.google.com/view/sci-imageminer/task-evaluation-metrics
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Shutao Gong, Jing Huang

This work is licensed under a Creative Commons Attribution 4.0 International License.