01

Reasoning

FrontierMath

Western lead

Performance on difficult, research-level mathematical problems.

Western 52.4%

GPT-5.5 Pro pre-release (high) OpenAI

Chinese 39.0%

Kimi k2.6 Moonshot AI

How to read this measure

Higher scores show stronger hard-problem solving; they do not establish broad reliability.

02

Software work

SWE-bench Verified

Western lead

Resolution of verified, real-world software engineering issues.

Western 83.5%

Claude Opus 4.7 (max) Anthropic

Chinese 78.7%

GLM 5.2 (max) Zhipu AI / Z.ai

How to read this measure

Measures bounded repository tasks, not unattended ownership of production systems.

03

Multimodal

Video-MME

Chinese lead

Understanding across short, medium, and long video without subtitle assistance.

Western 75.0%

Gemini 1.5 Pro Google DeepMind

Chinese 79.7%

video-SALMONN 2+ ByteDance

How to read this measure

Captures one form of visual-language understanding, not embodied world modelling.

04

Autonomy

METR task horizon (80%)

No direct comparison

The length of software tasks models complete at an estimated 80% success rate.

Western 1.5 hr

Gemini 3.1 Pro Preview Google DeepMind

Chinese No comparable result

Public dataset has no positive scored record.

How to read this measure

Longer horizons matter, but a benchmark task is not the same as safe, persistent agency.

05

Generalization

ARC-AGI-2

Western lead

Adaptation to novel abstract visual reasoning tasks.

Western 92.5%

GPT-5.6 Sol (Max) OpenAI

Chinese 22.8%

GLM-5.2 Zhipu AI / Z.ai

How to read this measure

Strong performance can reveal flexible reasoning, while benchmark saturation can weaken the signal.

Optional Groq interpretation

Understand a measure without changing the evidence.

Optional interpretation

Ask GPT-OSS to explain this signal

The model receives the selected public record, not an open web prompt. It cannot alter the score.