GLM-5.2 Deployment: vLLM unterstützt ab Tag 0, 744B Parameter gesamt / 40B aktiv, 1M Kontext, 128K max Output.
UnslothAI meldet eine 239GB 2-Bit-Variante mit etwa 81% Genauigkeit nach Schrumpfung von 1,51TB.
Unabhängige Schätzungen: $0,15/1M Input, $0,47/1M Output, rund 20× günstiger und ~2× schneller als Qwen3.8 Max.
GLM-5.2 deployment specs: vLLM day-0 support, 744B total / 40B active parameters, 1M context, 128K max output.
UnslothAI claims a 239GB 2-bit variant retains about 81% accuracy after shrinking from 1.51TB.
Independent estimates put pricing at $0.15/1M input and $0.47/1M output, roughly 20× cheaper and ~2× faster than Qwen3.8 Max.