We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Unsloth has a UD-IQ4_XS quant at 157 GB https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF
You must log in or register to comment.
After testing via api for a few days, this model is genuinely impressive for its size/price.
Feels a little like running opus 4.6 on medium think. Quick enough, clean enough. Still makes some minor errors once in a while, but it’s fast enough to be useful and cheap enough that you can use it for everything and still pay very little.


