Precision caliper, engineering drawings, and a glowing measurement path

PE Precision Bench · Frontier update

Newer isn't always better.

Grok 4.6 and Qwen3.8 Max joined our engineering benchmark. Neither took the lead—and the reason was precision, not raw problem solving.

97.1%

Kimi K3 · third overall

95.7%

Grok 4.6 · sixth overall

94.6%

Qwen3.8 Max · eighth overall

The headline is simple: model release order did not predict engineering precision. The same sealed questions and scoring rules produced a stable leader—and a revealing gap between selecting an answer and supporting it correctly.

The updated leaderboard

RankModelOverallAnswersStandards
1DeepSeek V4 Pro97.8%99.4%91.7%
2GPT-5.6 Sol97.5%99.7%88.9%
3Kimi K397.1%99.4%88.0%
4Grok 4.596.7%99.4%86.1%
5Claude Sonnet 596.2%98.8%86.1%
6Grok 4.695.7%98.1%86.1%
7Gemini 3.5 Flash95.1%97.8%84.3%
8Qwen3.8 Max94.6%98.1%80.6%

Grok 4.6 did not beat Grok 4.5

Grok 4.6 finished at 95.7%, one point behind Grok 4.5 at 96.7%. Their standards citation scores were identical at 86.1%; the difference came from answer accuracy, where Grok 4.6 scored 98.1% against Grok 4.5's 99.4%.

That is exactly why independent, task-specific evaluation matters. A newer model may be stronger in broad capability tests and still regress on a narrow professional protocol.

Kimi K3 was the strongest of the new generation

Kimi K3 placed third overall at 97.1%. It matched the field's best answer accuracy tier at 99.4% and delivered an 88.0% standards score. Its 97.6% standards-task score made it the most convincing new entrant when the question required both an engineering answer and the governing code or standard.

Qwen3.8 Max exposed the reliability tax

Qwen3.8 Max selected answers well—98.1% were correct—but its standards citation F1 fell to 80.6%. 4 of its 320 responses were still unusable after the benchmark's retry policy. We retained those failures as zeroes, because silently removing them would overstate deployment reliability.

The real separator was standards discipline

The new models all remained excellent on engineering analysis. The larger spread appeared when they had to name the correct governing standard without adding unsupported citations. This is a useful warning for real practice: polished reasoning and a correct multiple-choice selection do not guarantee dependable code grounding.

What this benchmark does not prove

PE Precision Bench is a one-run, closed-book, zero-temperature multiple-choice protocol. It does not establish engineering competence and it does not replace current standards, project documents, or review by a licensed engineer. It measures a narrower question: under identical constraints, which models most precisely combine answer selection with standards support?

Inspect the evidence

Explore every model and discipline

The interactive benchmark includes all eight models, discipline-level rankings, scoring details, and public development examples.

Open the benchmark

Model references