B1ade Architecture Delivers Efficient RAG with 335M Embedding and 1B SLM
July 31, 2026
B1ade uses a 335M parameter embedding model and a 1B parameter SLM trained via GRPO on 723M tokens. Despite no explicit citation training, the 1B model achieves a 42.4% emergent attribution rate in retrieving and citing passages.
HOW THIS AFFECTS YOU
●
builderYou can deploy highly efficient, low-latency RAG systems using these sub-1B parameter models.
●
researcherThe emergence of citation capabilities without explicit supervision provides a new direction for training small language models.