Creating Business Documents With LLMs – Architecture & Evaluation

Authors

  • Mohammed Al-Hebshi Technical University of Applied Sciences Wildau image/svg+xml
  • Alexander Lübbe Technical University of Applied Sciences Wildau image/svg+xml

DOI:

https://doi.org/10.52825/th-wildau-ensp.v3i.3515

Keywords:

Large Language Models, LLM-Evaluation, Agentic LLM-Architektur

Abstract

Large Language Models (LLMs) offer the potential to automate knowledge-intensive document generation; however, selecting the right model for a specific application remains a methodological challenge, as generic benchmarks do not adequately reflect suitability for quality-critical documentation processes. This thesis therefore develops a system architecture for quality-assured document generation and evaluates three current models – GPT-5, Claude Sonnet 4.5 and LLaMA 4 Maverick – against seven application-specific criteria, combining automated system metrics, LLM-as-a-Judge and human evaluation. The results reveal a hierarchy of criteria that distinguishes between necessary basic requirements – factual accuracy and reliable integration of system components – and differentiating quality characteristics, thereby providing an empirically sound basis for model selection.

Downloads

Download data is not yet available.

References

[1] M. Brachman, A. El-Ashry, C. Dugan, and W. Geyer, “Current and Future Use of Large Language Models for Knowledge Work,” Mar. 21, 2025, arXiv: arXiv:2503.16774. doi: 10.48550/arXiv.2503.16774.

[2] S. K. Mohan, “Management Consulting in the Artificial Intelligence – LLM Era,” Management Consulting Journal, vol. 7, no. 1, pp. 9–24, Jan. 2024, doi: 10.2478/mcj-2024-0002.

[3] D. Lippold, Die Unternehmensberatung. Wiesbaden: Springer Fachmedien Wiesbaden, 2018. doi: 10.1007/978-3-658-21092-2.

[4] A. Kalenborn, Angebotserstellung und Planung von Internet-Projekten: Die werkzeugbasierte “Modeling by Example”-Methode. Wiesbaden: Springer Fachmedien Wiesbaden, 2014. doi: 10.1007/978-3-658-07950-5.

[5] N. Krishnan, “AI Agents: Evolution, Architecture, and Real-World Applications,” Mar. 16, 2025, arXiv: arXiv:2503.12687. doi: 10.48550/arXiv.2503.12687.

[6] Anthropic PBC, “Building effective agents.” Zugriff am: 2026-02-15 [Online]. Available: https://www.anthropic.com/engineering/building-effective-agents

[7] K. Bourne, Unlocking data with generative AI and RAG: enhance generative AI systems by integrating internal data with large language models using RAG. Place of publication not identified: Packt Publishing, 2024.

[8] E. Croxford et al., “Current and future state of evaluation of large language models for medical summarization tasks,” npj Health Syst., vol. 2, no. 1, p. 6, Feb. 2025, doi: 10.1038/s44401-024-00011-2.

[9] U. Kamath, K. Keenan, G. Somers, and S. Sorenson, Large Language Models: A Deep Dive: Bridging Theory and Practice. Cham: Springer Nature Switzerland, 2024. doi: 10.1007/978-3-031-65647-7.

[10] S. Pai, Designing Large Language Model Applications: A Holistic Approach to LLMs, First edition. Sebastopol, CA: O’Reilly, 2025.

[11] K. Busch and H. Leopold, “Towards a Benchmark for Large Language Models for Business Process Management Tasks,” Oct. 13, 2024, arXiv: arXiv:2410.03255. doi: 10.48550/arXiv.2410.03255.

[12] Langfuse GmbH, “Observability Overview – Langfuse Docs.” Zugriff am: 2026-02-15 [Online]. Available: https://langfuse.com/docs/observability/overview

[13] J. Gu et al., “A Survey on LLM-as-a-Judge,” Oct. 19, 2025, arXiv: arXiv:2411.15594. doi: 10.48550/arXiv.2411.15594.

[14] Y. Chang et al., “A Survey on Evaluation of Large Language Models,” Dec. 29, 2023, arXiv: arXiv:2307.03109. doi: 10.48550/arXiv.2307.03109.

[15] Dify.AI, “Introduction to Dify.” Zugriff am: 2026-02-15 [Online]. Available: https://docs.dify.ai/en/introduction

[16] Dify.ai, “Agent.” Zugriff am: 2026-02-15 [Online]. Available: https://docs.dify.ai/versions/3-0-x/en/user-guide/application-orchestrate/agent

[17] Dify.ai, “Integrate Knowledge within Apps.” Zugriff am: 2026-02-15 [Online]. Available: https://docs.dify.ai/en/use-dify/knowledge/integrate-knowledge-within-application

[18] OpenRouter, Inc, “Quickstart.” Zugriff am: 2026-02-15 [Online]. Available: https://openrouter.ai/docs/quickstart

[19] Langfuse GmbH, “Model Usage & Cost Tracking.” Zugriff am: 2026-02-15 [Online]. Available: https://langfuse.com/docs/observability/features/token-and-cost-tracking

[20] Google LLC, “Gemini 2.5 Pro.” Zugriff am: 2026-02-15 [Online]. Available: https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro

[21] J. Cook, T. Rocktäschel, J. Foerster, D. Aumiller, and A. Wang, “TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation,” Oct. 04, 2024, arXiv: arXiv:2410.03608. doi: 10.48550/arXiv.2410.03608.

[22] S. Ye et al., “FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets,” Apr. 14, 2024, arXiv: arXiv:2307.10928. doi: 10.48550/arXiv.2307.10928.

[23] C. Irugalbandara et al., “Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI’s LLM with Open Source SLMs in Production,” Apr. 16, 2024, arXiv: arXiv:2312.14972. doi: 10.48550/arXiv.2312.14972.

[24] Anthropic PBC, “Introducing Claude Sonnet 4.5.” Zugriff am: 2026-02-15 [Online]. Available: https://www.anthropic.com/news/claude-sonnet-4-5

[25] OpenAI, “Introducing GPT-5.” Zugriff am: 2026-02-15 [Online]. Available: https://openai.com/index/introducing-gpt-5/

[26] Meta Platforms Inc., “The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.” Zugriff am: 2026-02-15 [Online]. Available: https://ai.meta.com/blog/llama-4-multimodal-intelligence/

[27] Anthropic PBC, “Is my data used for model training?” Zugriff am: 2026-02-15 [Online]. Available: https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training

[28] OpenAI, “Enterprise privacy at OpenAI.” Zugriff am: 2026-02-15 [Online]. Available: https://openai.com/enterprise-privacy/

Published

2026-08-18

How to Cite

Al-Hebshi, M., & Lübbe, A. (2026). Creating Business Documents With LLMs – Architecture & Evaluation. TH Wildau Engineering and Natural Sciences Proceedings , 3. https://doi.org/10.52825/th-wildau-ensp.v3i.3515

Conference Proceedings Volume

Section

Contributions to the Wildau Conference on Artificial Intelligence 2026