TEMPLAR
Run
Now Running
QWEN30_MOE_003
--model=qwen3-30b-a3b--arch=moe--experts=128--top-k=8
Pre-training with SparseLoCo and low-bandwidth pipelining. Read the paper →
Training Progress
Is the model learning, and how fast
Training Loss
Tokens Processed
MFU — compute vs effective
Network & Topology
Stage chains, replica sync and round health