A 29B parameter Mixture-of-Experts (MoE) model activates 4B parameters per token and supports up to a 512K context length. Optimized on Ascend NPUs using the MindSpore framework, it utilizes mHC, MLA, and MTP architectures for agentic reasoning and tool calling.
HOW THIS AFFECTS YOU
●
builderYou can leverage its native support for long-context agentic workflows and complex planning.
●
researcherThe use of mHC and MLA architectures on Ascend hardware offers a new benchmark for MoE efficiency.