Today, we’re releasing Ling-3.0-tiny: 7.9B total parameters, with only 1.3B active per token.
A native hybrid reasoning model built for real-world tasks, math, instruction following, and resource-sensitive deployment.
More intelligence with less compute. 🧵
显示更多
We recently released a paper showing that UFP4, our uniform-grid FP4 training recipe, stays closer to BF16 than strong E2M1 baselines across Dense 1.5B, MoE 7.9B, and MoE 124B long-run pretraining.
The key insight: FP4 training quality is not only about bit width, but also grid geometry.
显示更多