Minimax M3 model support, including Multi-Head Selective Attention (MSA), has been merged into the llama.cpp repository. This enables efficient local inference for Minimax M3 architectures.
HOW THIS AFFECTS YOU
●
builderYou can now run Minimax M3 models locally using llama.cpp for faster, private inference.