Research note · 05 August 2026

Frontier Model
Benchmark Ledger

A compact, source-led comparison of GLM-5.2, DeepSeek-V4-Flash, and GPT-5.6 Luna. Publisher claims and independent measurements are kept visibly separate.

Static reference edition · v1.0

00 / Reading note

Scores are only comparable when the test is.

Benchmark versions, agent harnesses, reasoning effort, token budgets, and tool access can materially change a result. A dash means that no directly comparable figure was reported in the cited first-party release. Higher is better unless noted otherwise.

01 / Model cards

Three efficiency tiers, three different bets

Z.ai · Open weights

GLM-5.2

Long-horizon flagship

Context
1M tokens
Parameters
753B / 40B active
License
MIT

DeepSeek · Open weights

V4 Flash

Fast reasoning model

Context
1M tokens
Parameters
284B / 13B active
License
MIT

OpenAI · Hosted

GPT-5.6 Luna

High-volume efficiency tier

Context
1.05M tokens
Max output
128K tokens
Weights
Proprietary

02 / Publisher-reported evaluations

Reported capability results

Percent unless noted

Evaluation GLM-5.2 DeepSeek V4 Flash GPT-5.6 Luna
GPQA Diamond 91.2publisher setup 88.1Max · Pass@1 92.3publisher setup
SWE-bench Pro 62.1OpenHands 52.6Max · resolved 62.7publisher setup
Humanity's Last Exam 40.5text-only 34.8Max · Pass@1 not reported in release table
Terminal-Bench 81.0v2.1 · Terminus-2 56.9v2.0 · Max 84.7v2.1
BrowseComp not reported 73.2Max · Pass@1 83.3publisher setup
MCP-Atlas public set 76.8thinking 69.0Max · Pass@1 not reported
DeepSWE 46.2mini-swe-agent not reported for Flash 67.2v1.1

Terminal-Bench uses different versions for DeepSeek and the other two models. SWE-bench Pro harnesses also differ. These rows are preserved as reported records, not presented as controlled head-to-head experiments.

03 / Independent snapshot

Artificial Analysis, comparable reasoning tiers

Intelligence Index v4.1

GLM-5.2 max51
GPT-5.6 Luna high46
DeepSeek V4 Flash max40

Observed output speed

179 t/s

GLM-5.2 max, provider median snapshot

161 t/s

GPT-5.6 Luna high

116 t/s

DeepSeek V4 Flash max

Published API price

GLM-5.2$1.40 / $4.40
DeepSeek V4 Flash$0.14 / $0.28
GPT-5.6 Luna$1.00 / $6.00

USD per 1M input / output tokens. GLM figure is the independent provider median; DeepSeek and OpenAI figures are publisher prices.

04 / What the records suggest

There is no single winner—only a clearer workload fit.

GLM-5.2 leads this independent intelligence snapshot and posts strong agentic results while remaining open-weight.

DeepSeek V4 Flash is the price outlier: materially cheaper, smaller per-token activation, and competitive at Max effort.

GPT-5.6 Luna combines strong publisher-reported coding and browsing scores with a hosted, high-volume product profile.

05 / Sources

Primary records and independent methodology

  1. Z.ai · GLM-5.2 official model card
  2. DeepSeek · V4 Flash official model card and technical report
  3. OpenAI · GPT-5.6 launch and evaluation tables
  4. OpenAI · GPT-5.6 Luna model reference
  5. Artificial Analysis · Intelligence, speed, and pricing methodology

Compiled 05 August 2026 · Figures may change as providers update models, harnesses, and pricing.