Creating Business Documents With LLMs – Architecture & Evaluation
DOI:
https://doi.org/10.52825/th-wildau-ensp.v3i.3515Keywords:
Large Language Models, LLM-Evaluation, Agentic LLM-ArchitekturAbstract
Large Language Models (LLMs) offer the potential to automate knowledge-intensive document generation; however, selecting the right model for a specific application remains a methodological challenge, as generic benchmarks do not adequately reflect suitability for quality-critical documentation processes. This thesis therefore develops a system architecture for quality-assured document generation and evaluates three current models – GPT-5, Claude Sonnet 4.5 and LLaMA 4 Maverick – against seven application-specific criteria, combining automated system metrics, LLM-as-a-Judge and human evaluation. The results reveal a hierarchy of criteria that distinguishes between necessary basic requirements – factual accuracy and reliable integration of system components – and differentiating quality characteristics, thereby providing an empirically sound basis for model selection.
Downloads
References
[1] M. Brachman, A. El-Ashry, C. Dugan, and W. Geyer, “Current and Future Use of Large Language Models for Knowledge Work,” Mar. 21, 2025, arXiv: arXiv:2503.16774. doi: 10.48550/arXiv.2503.16774.
[2] S. K. Mohan, “Management Consulting in the Artificial Intelligence – LLM Era,” Management Consulting Journal, vol. 7, no. 1, pp. 9–24, Jan. 2024, doi: 10.2478/mcj-2024-0002.
[3] D. Lippold, Die Unternehmensberatung. Wiesbaden: Springer Fachmedien Wiesbaden, 2018. doi: 10.1007/978-3-658-21092-2.
[4] A. Kalenborn, Angebotserstellung und Planung von Internet-Projekten: Die werkzeugbasierte “Modeling by Example”-Methode. Wiesbaden: Springer Fachmedien Wiesbaden, 2014. doi: 10.1007/978-3-658-07950-5.
[5] N. Krishnan, “AI Agents: Evolution, Architecture, and Real-World Applications,” Mar. 16, 2025, arXiv: arXiv:2503.12687. doi: 10.48550/arXiv.2503.12687.
[6] Anthropic PBC, “Building effective agents.” Zugriff am: 2026-02-15 [Online]. Available: https://www.anthropic.com/engineering/building-effective-agents
[7] K. Bourne, Unlocking data with generative AI and RAG: enhance generative AI systems by integrating internal data with large language models using RAG. Place of publication not identified: Packt Publishing, 2024.
[8] E. Croxford et al., “Current and future state of evaluation of large language models for medical summarization tasks,” npj Health Syst., vol. 2, no. 1, p. 6, Feb. 2025, doi: 10.1038/s44401-024-00011-2.
[9] U. Kamath, K. Keenan, G. Somers, and S. Sorenson, Large Language Models: A Deep Dive: Bridging Theory and Practice. Cham: Springer Nature Switzerland, 2024. doi: 10.1007/978-3-031-65647-7.
[10] S. Pai, Designing Large Language Model Applications: A Holistic Approach to LLMs, First edition. Sebastopol, CA: O’Reilly, 2025.
[11] K. Busch and H. Leopold, “Towards a Benchmark for Large Language Models for Business Process Management Tasks,” Oct. 13, 2024, arXiv: arXiv:2410.03255. doi: 10.48550/arXiv.2410.03255.
[12] Langfuse GmbH, “Observability Overview – Langfuse Docs.” Zugriff am: 2026-02-15 [Online]. Available: https://langfuse.com/docs/observability/overview
[13] J. Gu et al., “A Survey on LLM-as-a-Judge,” Oct. 19, 2025, arXiv: arXiv:2411.15594. doi: 10.48550/arXiv.2411.15594.
[14] Y. Chang et al., “A Survey on Evaluation of Large Language Models,” Dec. 29, 2023, arXiv: arXiv:2307.03109. doi: 10.48550/arXiv.2307.03109.
[15] Dify.AI, “Introduction to Dify.” Zugriff am: 2026-02-15 [Online]. Available: https://docs.dify.ai/en/introduction
[16] Dify.ai, “Agent.” Zugriff am: 2026-02-15 [Online]. Available: https://docs.dify.ai/versions/3-0-x/en/user-guide/application-orchestrate/agent
[17] Dify.ai, “Integrate Knowledge within Apps.” Zugriff am: 2026-02-15 [Online]. Available: https://docs.dify.ai/en/use-dify/knowledge/integrate-knowledge-within-application
[18] OpenRouter, Inc, “Quickstart.” Zugriff am: 2026-02-15 [Online]. Available: https://openrouter.ai/docs/quickstart
[19] Langfuse GmbH, “Model Usage & Cost Tracking.” Zugriff am: 2026-02-15 [Online]. Available: https://langfuse.com/docs/observability/features/token-and-cost-tracking
[20] Google LLC, “Gemini 2.5 Pro.” Zugriff am: 2026-02-15 [Online]. Available: https://ai.google.dev/gemini-api/docs/models/gemini-2.5-pro
[21] J. Cook, T. Rocktäschel, J. Foerster, D. Aumiller, and A. Wang, “TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation,” Oct. 04, 2024, arXiv: arXiv:2410.03608. doi: 10.48550/arXiv.2410.03608.
[22] S. Ye et al., “FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets,” Apr. 14, 2024, arXiv: arXiv:2307.10928. doi: 10.48550/arXiv.2307.10928.
[23] C. Irugalbandara et al., “Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI’s LLM with Open Source SLMs in Production,” Apr. 16, 2024, arXiv: arXiv:2312.14972. doi: 10.48550/arXiv.2312.14972.
[24] Anthropic PBC, “Introducing Claude Sonnet 4.5.” Zugriff am: 2026-02-15 [Online]. Available: https://www.anthropic.com/news/claude-sonnet-4-5
[25] OpenAI, “Introducing GPT-5.” Zugriff am: 2026-02-15 [Online]. Available: https://openai.com/index/introducing-gpt-5/
[26] Meta Platforms Inc., “The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.” Zugriff am: 2026-02-15 [Online]. Available: https://ai.meta.com/blog/llama-4-multimodal-intelligence/
[27] Anthropic PBC, “Is my data used for model training?” Zugriff am: 2026-02-15 [Online]. Available: https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training
[28] OpenAI, “Enterprise privacy at OpenAI.” Zugriff am: 2026-02-15 [Online]. Available: https://openai.com/enterprise-privacy/
Downloads
Published
How to Cite
Conference Proceedings Volume
Section
License
Copyright (c) 2026 Mohammed Al-Hebshi, Alexander Lübbe

This work is licensed under a Creative Commons Attribution 4.0 International License.