Z AI released GLM-5.3-Flash on August 26, 2026, an open-weight reasoning model with 320 billion total parameters and 18 billion active parameters per inference token. The Mixture-of-Experts model accepts text and image inputs, produces text outputs, and supports a 1 million-token context window. Its weights are available on Hugging Face under the MIT license, which permits commercial use. Artificial Analysis gave GLM-5.3-Flash a score of 57 on its Intelligence Index, compared with a median score of 27 for open-weight models of a similar size. The composite benchmark covers reasoning, knowledge, mathematics and coding. The evaluation generated 150 million output tokens, above the comparable-model median of 110 million, indicating relatively verbose outputs. Z AI's API charges $0.15 per million input tokens and $0.50 per million output tokens; Artificial Analysis lists a blended cache-hit/input/output rate of $0.10 per million tokens using a 7:2:1 ratio. The Intelligence Index evaluation cost $138.02. The model produces about 50.2 tokens per second, below the comparable median of 65.8, while its 1.56-second time to first token is faster than the 2.13-second median.
artificialanalysis.ai
4 min
8/26/2026
ZCode is an official development tool optimized for the GLM-5.2 model, enhancing stability and efficiency in programming tasks. It supports over 20 programming tools, offers various subscription plans for different development needs, and enables remote task execution via platforms like WeChat, Feishu, or Telegram.
zcode.z.ai
1 min
7/1/2026
GLM-5.2 (max) is a leading model in intelligence with a score of 51 on the Artificial Analysis Intelligence Index. It offers a 1 million token context window, supports text input and output, is faster than average, but is considered expensive compared to other open weight models of similar size.
artificialanalysis.ai
5 min
6/17/2026
Z AI released GLM-5.3-Flash on August 26, 2026, an open-weight reasoning model with 320 billion total parameters and 18 billion active parameters per inference token. The Mixture-of-Experts model accepts text and image inputs, produces text outputs, and supports a 1 million-token context window. Its weights are available on Hugging Face under the MIT license, which permits commercial use. Artificial Analysis gave GLM-5.3-Flash a score of 57 on its Intelligence Index, compared with a median score of 27 for open-weight models of a similar size. The composite benchmark covers reasoning, knowledge, mathematics and coding. The evaluation generated 150 million output tokens, above the comparable-model median of 110 million, indicating relatively verbose outputs. Z AI's API charges $0.15 per million input tokens and $0.50 per million output tokens; Artificial Analysis lists a blended cache-hit/input/output rate of $0.10 per million tokens using a 7:2:1 ratio. The Intelligence Index evaluation cost $138.02. The model produces about 50.2 tokens per second, below the comparable median of 65.8, while its 1.56-second time to first token is faster than the 2.13-second median.
artificialanalysis.ai
4 min
8/26/2026
GLM-5.2 (max) is a leading model in intelligence with a score of 51 on the Artificial Analysis Intelligence Index. It offers a 1 million token context window, supports text input and output, is faster than average, but is considered expensive compared to other open weight models of similar size.
artificialanalysis.ai
5 min
6/17/2026
ZCode is an official development tool optimized for the GLM-5.2 model, enhancing stability and efficiency in programming tasks. It supports over 20 programming tools, offers various subscription plans for different development needs, and enables remote task execution via platforms like WeChat, Feishu, or Telegram.
zcode.z.ai
1 min
7/1/2026
Z AI released GLM-5.3-Flash on August 26, 2026, an open-weight reasoning model with 320 billion total parameters and 18 billion active parameters per inference token. The Mixture-of-Experts model accepts text and image inputs, produces text outputs, and supports a 1 million-token context window. Its weights are available on Hugging Face under the MIT license, which permits commercial use. Artificial Analysis gave GLM-5.3-Flash a score of 57 on its Intelligence Index, compared with a median score of 27 for open-weight models of a similar size. The composite benchmark covers reasoning, knowledge, mathematics and coding. The evaluation generated 150 million output tokens, above the comparable-model median of 110 million, indicating relatively verbose outputs. Z AI's API charges $0.15 per million input tokens and $0.50 per million output tokens; Artificial Analysis lists a blended cache-hit/input/output rate of $0.10 per million tokens using a 7:2:1 ratio. The Intelligence Index evaluation cost $138.02. The model produces about 50.2 tokens per second, below the comparable median of 65.8, while its 1.56-second time to first token is faster than the 2.13-second median.
artificialanalysis.ai
4 min
8/26/2026
ZCode is an official development tool optimized for the GLM-5.2 model, enhancing stability and efficiency in programming tasks. It supports over 20 programming tools, offers various subscription plans for different development needs, and enables remote task execution via platforms like WeChat, Feishu, or Telegram.
zcode.z.ai
1 min
7/1/2026
GLM-5.2 (max) is a leading model in intelligence with a score of 51 on the Artificial Analysis Intelligence Index. It offers a 1 million token context window, supports text input and output, is faster than average, but is considered expensive compared to other open weight models of similar size.
artificialanalysis.ai
5 min
6/17/2026
No more articles to load