[DocMiner] at Sci-ImageMiner 2026 (Tasks 1, 2, and 4): Skill-Based Agentic DocMiner

A Context-Enriched Multi-Agent System for Scientific Figure Understanding

Authors

DOI:

https://doi.org/10.52825/ocp.v10i.3510

Keywords:

Scientific Figure Understanding, Multimodal Document Intelligence, Agentic Workflow, Chart Classification, Data Extraction, Visual Question Answering

Abstract

We present Skill-Based Agentic DocMiner, our submission to Sci-ImageMiner 2026 for Task 1 (figure classification), Task 2 (data table extraction), and Task 4 (visual question answering). Our system follows a unified agentic workflow with three recurring components: multimodal context enrichment from captions and paper text, iterative skill refinement using development-set feedback, and collaborative verification by specialist reviewers. For classification, we use class prototype documents to distinguish visually similar figure types. For table extraction, we combine chart parsing with Markdown structure validation and content checking. For VQA, we route each question to a specialist agent aligned with the official reasoning categories. On the official leaderboard, our submission ranked fourth on classification, fifth on table extraction, and third on VQA. The framework is simple, modular, and well matched to structured scientific reasoning. Our code is publicly available at https://github.com/1gst/Sci-ImageMiner.

Downloads

Download data is not yet available.

References

[1] Sci-ImageMiner Organizing Committee. “Sci-imageminer– task description,” Accessed: Apr. 10, 2026. [Online]. Available: https://sites.google.com/view/sci-imageminer/task-description

[2] Google. “Gemini api documentation,” Accessed: Apr. 10, 2026. [Online]. Available: https://ai.google.dev/gemini-api/docs/models

[3] Alibaba Cloud. “Dashscope / qwen api documentation,” Accessed: Apr. 10, 2026. [Online]. Available: https://help.aliyun.com/zh/model-studio/getting-started/models

[4] S. Yao et al., “React: Synergizing reasoning and acting in language models,” arXiv preprint arXiv:2210.03629, 2022. DOI: 10.48550/arXiv.2210.03629 [Online]. Available: https://arxiv.org/abs/2210.03629

[5] T. Schick et al., “Toolformer: Language models can teach themselves to use tools,” arXiv preprint arXiv:2302.04761, 2023. DOI: 10.48550/arXiv.2302.04761 [Online]. Available: https://arxiv.org/abs/2302.04761

[6] Q. Wu et al., “Autogen: Enabling next-gen llm applications via multi-agent conversation,” arXiv preprint arXiv:2308.08155, 2023. DOI: 10.48550/arXiv.2308.08155 [Online]. Available: https://arxiv.org/abs/2308.08155

[7] Sci-ImageMiner Organizing Committee. “Sci-imageminer–task evaluation metrics,” Accessed: Apr. 10, 2026. [Online]. Available: https://sites.google.com/view/sci-imageminer/task-evaluation-metrics

Downloads

Published

2026-09-01

How to Cite

Gong, S., & Huang, J. (2026). [DocMiner] at Sci-ImageMiner 2026 (Tasks 1, 2, and 4): Skill-Based Agentic DocMiner: A Context-Enriched Multi-Agent System for Scientific Figure Understanding. Open Conference Proceedings, 10. https://doi.org/10.52825/ocp.v10i.3510