01

Reasoning

FrontierMath

Western lead

Performance on difficult, research-level mathematical problems.

Western 52.4%

GPT-5.5 Pro pre-release (high) OpenAI

Chinese 39.0%

Kimi k2.6 Moonshot AI

How to read this measure

Higher scores show stronger hard-problem solving; they do not establish broad reliability.

02

Software work

SWE-bench Verified

Western lead

Resolution of verified, real-world software engineering issues.

Western 83.5%

Claude Opus 4.7 (max) Anthropic

Chinese 78.7%

GLM 5.2 (max) Zhipu AI / Z.ai

How to read this measure

Measures bounded repository tasks, not unattended ownership of production systems.

03

Multimodal

Video-MME

Chinese lead

Understanding across short, medium, and long video without subtitle assistance.

Western 75.0%

Gemini 1.5 Pro Google DeepMind

Chinese 79.7%

video-SALMONN 2+ ByteDance

How to read this measure

Captures one form of visual-language understanding, not embodied world modelling.

04

Autonomy

METR task horizon (80%)

No direct comparison

The length of software tasks models complete at an estimated 80% success rate.

Western 3.1 hr

Claude mythos Preview early Anthropic

Chinese No comparable result

Public dataset has no positive scored record.

How to read this measure

Longer horizons matter, but a benchmark task is not the same as safe, persistent agency.

05

Generalization

ARC-AGI-2

Western lead

Adaptation to novel abstract visual reasoning tasks.

Western 92.5%

GPT-5.6 Sol (Max) OpenAI

Chinese 61.4%

DeepSeek V4 Flash 0731 (Max) DeepSeek

How to read this measure

Strong performance can reveal flexible reasoning, while benchmark saturation can weaken the signal.

Optional plain-language interpretation

Understand a measure without changing the evidence.

Optional interpretation

Explain this signal in plain language

The model receives the selected public record, not an open web prompt. It cannot alter the score.