We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

Unsloth has a UD-IQ4_XS quant at 157 GB https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF

  • ArchAengelus@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    2
    ·
    3 days ago

    After testing via api for a few days, this model is genuinely impressive for its size/price.

    Feels a little like running opus 4.6 on medium think. Quick enough, clean enough. Still makes some minor errors once in a while, but it’s fast enough to be useful and cheap enough that you can use it for everything and still pay very little.