Home: Motoring > Ant Group Open-Sources Ling-3.0-Tiny, a Lightweight 7.9B-Parameter Inference Model

Ant Group Open-Sources Ling-3.0-Tiny, a Lightweight 7.9B-Parameter Inference Model

From:Internet Info Agency 2026-08-11 18:32:00

Today, Ant Bailing released the Ling-3.0-tiny model on the Hugging Face platform. This lightweight mixture-of-experts (MoE) model features a total of 7.9 billion parameters, with 1.3 billion parameters activated per token. It offers three weight precision versions—BF16, FP8, and INT4—to accommodate diverse deployment requirements across platforms. The model employs a hybrid linear architecture that alternately stacks KDA and MLA layers in a 3:1 ratio and incorporates a sparse MoE feedforward network with 128 routed experts. Ling-3.0-tiny supports both fast-response mode and multi-step reasoning mode, specifically optimized for local deployment. Official validation has been completed on devices including NVIDIA DGX Spark, Apple Silicon MacBooks, and Mac mini. On an M1 Pro MacBook, the model achieves an inference speed of 86–90 tokens per second, with peak memory consumption of approximately 8.34 GiB under an 8K context length.

Editor:NewsAssistant