Arm-Gemma-E4B: First Open Armenian LLM with Full Training Recipe
September 4, 2026
Researchers released arm-gemma-e4b, a model created by continued pretraining of Gemma-4-E4B on the new ArmWeb and ArmSTEM datasets. The ArmSTEM dataset contains 373K English-Armenian parallel math and science problems, enabling the model to outperform existing open Armenian LLMs.
HOW THIS AFFECTS YOU
●
builderYou can use the released ArmWeb and ArmSTEM datasets to adapt models for Armenian-speaking users.
●
researcherThis provides a reproducible recipe for low-resource, morphologically rich language adaptation.