LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch
J. Pfister, J. Wunderle, and A. Hotho. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page 2227--2246. Vienna, Austria, Association for Computational Linguistics, (July 2025)
Abstract
We transparently create two German-only decoder models, LLäMmlein 120M and 1B, from scratch and publish them, along with the training data, for the (German) NLP research community to use. The model training involved several key steps, including data preprocessing/filtering, the creation of a German tokenizer, the training itself, as well as the evaluation of the final models on various benchmarks, also against existing models. Throughout the training process, multiple checkpoints were saved in equal intervals and analyzed using the German SuperGLEBer benchmark to gain insights into the models' learning process.Compared to state-of-the-art models on the SuperGLEBer benchmark, both LLäMmlein models performed competitively, consistently matching or surpassing models with similar parameter sizes. The results show that the models' quality scales with size as expected, but performance improvements on some tasks plateaued early during training, offering valuable insights into resource allocation for future models.
%0 Conference Paper
%1 pfister-etal-2025-llammlein
%A Pfister, Jan
%A Wunderle, Julia
%A Hotho, Andreas
%B Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
%C Vienna, Austria
%D 2025
%E Che, Wanxiang
%E Nabende, Joyce
%E Shutova, Ekaterina
%E Pilehvar, Mohammad Taher
%I Association for Computational Linguistics
%K German LLM author:hotho author:pfister author:wunderle from:janpf myown nlp
%P 2227--2246
%T LLäMmlein: Transparent, Compact and Competitive German-Only Language Models from Scratch
%U https://aclanthology.org/2025.acl-long.111/
%X We transparently create two German-only decoder models, LLäMmlein 120M and 1B, from scratch and publish them, along with the training data, for the (German) NLP research community to use. The model training involved several key steps, including data preprocessing/filtering, the creation of a German tokenizer, the training itself, as well as the evaluation of the final models on various benchmarks, also against existing models. Throughout the training process, multiple checkpoints were saved in equal intervals and analyzed using the German SuperGLEBer benchmark to gain insights into the models' learning process.Compared to state-of-the-art models on the SuperGLEBer benchmark, both LLäMmlein models performed competitively, consistently matching or surpassing models with similar parameter sizes. The results show that the models' quality scales with size as expected, but performance improvements on some tasks plateaued early during training, offering valuable insights into resource allocation for future models.
%@ 979-8-89176-251-0
@inproceedings{pfister-etal-2025-llammlein,
abstract = {We transparently create two German-only decoder models, LL{\"a}Mmlein 120M and 1B, from scratch and publish them, along with the training data, for the (German) NLP research community to use. The model training involved several key steps, including data preprocessing/filtering, the creation of a German tokenizer, the training itself, as well as the evaluation of the final models on various benchmarks, also against existing models. Throughout the training process, multiple checkpoints were saved in equal intervals and analyzed using the German SuperGLEBer benchmark to gain insights into the models' learning process.Compared to state-of-the-art models on the SuperGLEBer benchmark, both LL{\"a}Mmlein models performed competitively, consistently matching or surpassing models with similar parameter sizes. The results show that the models' quality scales with size as expected, but performance improvements on some tasks plateaued early during training, offering valuable insights into resource allocation for future models.},
added-at = {2025-07-23T15:16:06.000+0200},
address = {Vienna, Austria},
author = {Pfister, Jan and Wunderle, Julia and Hotho, Andreas},
biburl = {https://www.bibsonomy.org/bibtex/21517e580b4316bccd2f9a5a738aad438/dmir},
booktitle = {Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
editor = {Che, Wanxiang and Nabende, Joyce and Shutova, Ekaterina and Pilehvar, Mohammad Taher},
interhash = {dd53e3b95ada629ac911cd077c3e42b1},
intrahash = {1517e580b4316bccd2f9a5a738aad438},
isbn = {979-8-89176-251-0},
keywords = {German LLM author:hotho author:pfister author:wunderle from:janpf myown nlp},
month = jul,
pages = {2227--2246},
publisher = {Association for Computational Linguistics},
timestamp = {2025-08-03T05:01:28.000+0200},
title = {LLäMmlein: Transparent, Compact and Competitive {G}erman-Only Language Models from Scratch},
url = {https://aclanthology.org/2025.acl-long.111/},
year = 2025
}