TEMPLAR 

Run
Now Running

QWEN30_MOE_003

--model=qwen3-30b-a3b--arch=moe--experts=128--top-k=8

Pre-training with SparseLoCo and low-bandwidth pipelining. Read the paper →

Training Progress

Is the model learning, and how fast

Training Loss

Tokens Processed

MFU — compute vs effective

Network & Topology

Stage chains, replica sync and round health

Cross-Replica Bandwidth

Round Duration

Participants