GLM-5.3 Artificial Analysis Benchmarks

(artificialanalysis.ai)

56 points | by apitman 1 hour ago

7 comments

  • scotttrinh 33 minutes ago
    I like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics:

        Model                        Score    Cost / Task    Output Tokens / Task
        -------------------------------------------------------------------------
        GLM-5.3 (max)                 59.5          $0.68                  41,107
        GLM-5.2 (max)                 53.0          $0.56                  32,200
        Claude Opus 5 (high)          61.5          $1.52                  21,353
        GPT-5.6 Sol (max)             60.9          $1.23                  16,879
        Grok 4.6 (high)               60.9          $0.84                  21,735
        Kimi K3 (max)                 59.7          $0.84                  25,474
        GPT-5.6 Sol (xhigh)           59.0          $0.87                  11,098
        Claude Opus 5 (medium)        58.6          $0.98                  12,459
        Qwen3.8 Max                   58.1          $1.13                  38,287
        Qwen3.8 2.4T A95B             57.7          $0.95                  32,472
        Claude Opus 4.8 (max)         57.3          $1.65                  33,557
        GPT-5.6 Sol (high)            57.3          $0.52                   7,545
        Muse Spark 1.2 (xhigh)        56.8          $0.40                  30,430
        GPT-5.6 Terra (max)           56.6          $0.51                  20,838
        GPT-5.5 (xhigh)               56.3          $0.69                  16,893
        Gemini 3.7 Flash (high)       56.0          $0.40                  36,847
    
    Edited for accuracy and more models.
    • teravor 13 minutes ago
      these $/task figures aren't very useful in my experience. it doesn't tell you how well it did the task.

      generally I choose models by their intelligence and then personal preference from direct experience.

    • sourcecodeplz 26 minutes ago
      Muse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.
      • sscaryterry 18 minutes ago
        I found the sweetspot here: GPT-5.6 Sol (high) 57.3 $0.52 7,545
  • BinRoo 32 minutes ago
    Beware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html
    • Onavo 4 minutes ago
      The Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).
  • markasoftware 44 minutes ago
    Very impressive score for the size, though token use is higher than k3 and far higher than proprietary models, and its price to performance isn't all that far ahead of k3 as a result
    • Havoc 35 minutes ago
      >token use is higher than k3 and far higher than proprietary models

      GLM sets effort to max by default historically.

  • Zaheer 35 minutes ago
    Is it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.
    • culi 31 minutes ago
      Use a unified proxy that lets you switch between models seamlessly. We are far from an equilibrium in this market and you will continue to have FOMO no matter who you pick if you go all in on one company
    • colingauvin 32 minutes ago
      At least by API usage, they aren't yet lower cost than subscriptions. Not sure about GLM's subscription plans though.
    • notatoad 27 minutes ago
      no, at subscription prices claude is a better value than GLM.

      They're only a better value if you're paying API rates

  • colingauvin 31 minutes ago
    Tied for #1 by agentic index (with Opus 5).
  • colingauvin 57 minutes ago
    ...do I take out a double mortgage to buy a 4 Spark cluster?
    • nvme0n1p1 36 minutes ago
      No, you use openrouter and spend 10% as much as using a proprietary model.
    • lisplist 50 minutes ago
      $20k is personal loan territory, not a second mortgage lol
  • fenestella 1 minute ago
    [flagged]