A model in the suite · Anthropic

Claude Opus 5

Anthropic · Claude · Opus 5 / xhigh effort (Anthropic-recommended for agentic coding) / Suite 2.0 API harness · 2026-07-25

81/100
Strict suite averageNo legacy score · 4 benchmarks

Opus 5 / xhigh effort (Anthropic-recommended for agentic coding) / Suite 2.0 API harness

Copies Claude Opus 5's full data pack — paste it into ChatGPT, Claude, or any AI to talk it through.

How Claude Opus 5 handled each benchmark

Score, capability radar, and the honest read on what it nailed and where it slipped. Hit Overlay to drop other models onto the same axes.

Dingo & Co. Knowledge Work

A 23-deliverable consulting brief: research, financial reconciliation, regulatory analysis, decks and spreadsheets. Tests whether a model can run an entire knowledge-work engagement end to end.

86
Excellent

An excellent, unusually thoughtful work package that handles the benchmark's legal, ethical, quantitative, and strategic traps with mature judgment. Its canonical fact set and gated GTM architecture are standout work. It falls short of near mastery because research citations need a real verification pass, several cross-document arithmetic and scope inconsistencies reach public-facing copy, and minor dashboard polish remains. A VP could use this internally immediately, but should not release the external assets without final research, legal, and numerical QA.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegrityResearch GroundingSemantic JudgmentQuant. Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeedSpatial Reas.
1GPT-5.6 Sol93
2Kimi K392
3Grok 4.591
4GPT-5.6 Luna90
5GPT-5.6 Terra88
6GLM 5.2 (OpenRouter)88
7Claude Opus 586
8Claude Fable 581
9Claude Sonnet 5 (xhigh)81
10Claude Opus 4.880
11GPT-5.578
12Gemini 3.5 Flash (High) Fast62
13Opus 4.754
14Sonnet 4.652
15Gemini 3.1 Pro38

What it nailed

  • Complete, native-format deliverable set with substantive Word, PowerPoint, Excel, PDF, and interactive HTML artifacts.
  • Exceptional judgment around exotic-animal legality, welfare, support liability, company-created demand, and the Alaska/Australia mismatch.
  • Strong quantitative reconciliation of recognized revenue, channel totals, import economics, attach attribution, paid-media quality, and launch-budget conflicts.
  • A genuinely strategic GTM plan with falsifiable gates, tranche releases, channel rules, stop conditions, and a board-level subsidiary decision.
  • Distinct, credible copy and unusually specific personas and lifecycle email sequences.

Where it slipped

  • Several competitor and pricing citations are weak or mismatched, especially the Litter-Robot 5 Pro and Litter-Robot 4 rows.
  • The public NPS breakdown is arithmetically inconsistent: six promoters and nine detractors across 19 responses do not produce -21.
  • Strict-TAM wording, sensing-channel counts, and launch-gate numbering drift across documents.
  • Public positioning for broader primitive-canid households conflicts with the package's stated prohibition on species endorsement beyond dingoes and hybrids.
  • The dashboard has minor visual polish issues, and full-format validation beyond expected-file checks was not independently supplied.
Wall clock 58m 20s

From the run

Car Wash Operations

A filthy operational dataset — ghost records, orphaned orders, typo'd customers, raw enum variants. Tests judgment under messy real-world data: what gets fixed, quarantined, or wrongly promoted.

75
Strong

A strong and highly auditable migration scaffold with an excellent schema, provenance layer, documentation set, and reviewer interface. It falls short of the excellent band because image-only evidence is not extracted, fuzzy typo-order resolution is not convincingly implemented, several semantic checks are questionable, and key determinism and security claims remain self-reported. The package is useful for expert review but is not safe for unsupervised production migration.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegritySemantic JudgmentQuant. Reas.UX ReviewabilityProd. ReadinessSpeed
1Claude Fable 588
2Claude Opus 4.886
3Claude Opus 575
4Kimi K365
5Claude Sonnet 5 (xhigh)64
6GPT-5.6 Sol55
7GPT-5.6 Terra55
8GPT-5.555
9GPT-5.6 Luna55
10Grok 4.555
11GLM 5.2 (OpenRouter)55
12Gemini 3.5 Flash (High) Fast51
13GPT-5.451
14Opus 4.748

What it nailed

  • Openable, deeply normalized SQLite package with unusually strong source-record and provenance modeling.
  • Comprehensive source inventory and broad extraction across structured and text formats.
  • Clear separation of raw evidence, canonical business data, conflicts, rejects, reconciliation, and human review tasks.
  • Strong handling of duplicate files and the SVC-007 service-code collision.
  • Reviewer-facing UI is coherent, feature-rich, and visually polished on desktop.
  • Documentation is candid about uncertainty, image limitations, and the need for human review.

Where it slipped

  • No OCR or manual transcription is provided for the 14 image receipts, notes, and whiteboards, leaving substantive planted records and conflicts unextracted.
  • The documented hard-evidence resolver does not establish that the 13 typo-order surnames were fuzzily matched; the 200-customer result suggests over-splitting.
  • The price-era canary reports zero 2025 jobs using prior-era prices despite describing that condition as present.
  • The date validator accepts February 29 in non-leap years, affecting the planted 2025-02-29 case.
  • Matching all 1,504 payments to jobs is insufficiently substantiated given the source's unmatched-payment pattern.
  • Byte-identical rerun and zero-secret-leak claims were not independently verified.
  • The mobile horizontal navigation visibly clips the next tab label.
Wall clock 44m 32s

From the run

Brick — The AI LEGO Build

Four buildable LEGO models from prompt to part list to runnable browser guide. Tests spatial reasoning, physical plausibility, and whether large builds hold together or collapse into repetition.

72
Strong

Equal-weight mean of four isolated Brick case scores: 100-piece-lunar-rover=83, 250-piece-rescue-helicopter=83.5, 500-piece-cyberpunk-food-stall=55, 1000-piece-airship-research-station=67 -> 72.12.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegritySemantic JudgmentQuant. Reas.Spatial Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeed
1Claude Fable 588
2Claude Opus 4.882
3Claude Sonnet 5 (xhigh)78
4Claude Opus 572
5Kimi K368
6GPT-5.6 Sol59
7Gemini 3.5 Flash (High) Fast56
8Grok 4.555
9GPT-5.6 Terra54
10GPT-5.6 Luna50
11GLM 5.2 (OpenRouter)50

What it nailed

  • [100-piece-lunar-rover] Exact 100-piece specification with a 98-piece grounded main component.
  • [100-piece-lunar-rover] A genuine structured kit specification drives the rendered pieces, instruction interface, and manifest rather than a separate decorative model.
  • [100-piece-lunar-rover] Clear 32-step, six-chapter assembly narrative with concrete part-ID mappings.
  • [100-piece-lunar-rover] Recognizable and polished lunar-rover presentation with a complete control set and responsive layouts.
  • [100-piece-lunar-rover] Required isolated-case artifact is present and renders successfully.
  • [250-piece-rescue-helicopter] Authoritative connectivity is strong, with 95.6% of pieces in the grounded main component.
  • [250-piece-rescue-helicopter] Exactly 250 structured parts are carried through the specification, manifest, instructions, and viewer.
  • [250-piece-rescue-helicopter] The 68-step, 14-chapter sequence is unusually concrete and reviewable for this benchmark.
  • [250-piece-rescue-helicopter] The finished model is a recognizable, detailed rescue helicopter with functional-looking rotor, door, winch, and helipad features.
  • [250-piece-rescue-helicopter] Desktop and mobile renders succeed, and the control surface is comprehensive and clearly labeled.
  • [500-piece-cyberpunk-food-stall] Delivers the required runnable index.html plus a structured 500-part kit specification, manifest, instructions, build scripts, and local Three.js dependency.
  • [500-piece-cyberpunk-food-stall] Maintains a large 446-piece connected core despite the strict connectivity failures.
  • [500-piece-cyberpunk-food-stall] Uses a strong, recognizable cyberpunk food-stall concept with neon lighting, signage, cooking details, accessories, removable roof, and scooter.
  • [500-piece-cyberpunk-food-stall] Desktop rendering and information design are polished and visually distinctive.
  • [500-piece-cyberpunk-food-stall] The instruction set is extensively chaptered and maps steps to concrete part IDs.
  • [1000-piece-airship-research-station] A substantial 997-part, 239-step artifact with a recognizable and visually polished airship-research-station concept.
  • [1000-piece-airship-research-station] The viewer, manifest, and instructions are structurally generated from the same kit specification rather than from a separate decorative scene.
  • [1000-piece-airship-research-station] Strong assembly-guide functionality, including highlighting, scrubbing, chapter navigation, subassembly movement, removable sections, animated propellers, and a final hero orbit.
  • [1000-piece-airship-research-station] Detailed, recording-friendly desktop presentation with useful alternate camera and preview states.

Where it slipped

  • [100-piece-lunar-rover] Strict validation finds two floating pieces, including the prominent windshield.
  • [100-piece-lunar-rover] The 432 near misses are very high relative to 100 parts and indicate imprecise spatial arithmetic.
  • [100-piece-lunar-rover] The model-authored claim of zero floating pieces is contradicted by independent measurement.
  • [100-piece-lunar-rover] The earliest chassis steps rely on loose coplanar plate adjacency before later reinforcement.
  • [100-piece-lunar-rover] The mobile completed view has minor badge/card overlap and leaves limited room to inspect the model.
  • [100-piece-lunar-rover] Offline local operation is not established because three.js is loaded from a CDN.
  • [250-piece-rescue-helicopter] Eleven pieces are strictly floating despite the artifact's successful-support and rotor-joint claims.
  • [250-piece-rescue-helicopter] The 925 measured near misses indicate substantial approximate placement rather than consistently precise stud engagement.
  • [250-piece-rescue-helicopter] The completed-state title overlaps the helicopter and becomes difficult to read.
  • [250-piece-rescue-helicopter] Mobile completion views leave limited unobstructed space for inspecting the final model.
  • [250-piece-rescue-helicopter] Not every advertised control has independent interaction evidence, and the viewer depends on externally hosted three.js.
  • [500-piece-cyberpunk-food-stall] Strict measurement finds 54 floating pieces and 45 connected components.
  • [500-piece-cyberpunk-food-stall] The 2,305 near misses indicate extensive approximate placement rather than precise stud engagement.
  • [500-piece-cyberpunk-food-stall] The artifact's own no-floating-parts claim is contradicted by independent connectivity measurement.
  • [500-piece-cyberpunk-food-stall] Several written steps state quantities that disagree with their listed part IDs.
  • [500-piece-cyberpunk-food-stall] Mobile presentation has major clipping, cropping, and overlay-obstruction defects.
  • [500-piece-cyberpunk-food-stall] The run was slow, taking about 3,479 seconds across 60 API calls.
  • [1000-piece-airship-research-station] Authoritative validation finds 244 mechanically floating parts and 213 components, despite the artifact claiming zero unsupported pieces.
  • [1000-piece-airship-research-station] The 3,948 near misses indicate extensive approximate placement rather than precise stud-grid engagement.
  • [1000-piece-airship-research-station] The envelope depends heavily on idealized custom shell and former elements that are not demonstrated as ordinary buildable brick parts.
  • [1000-piece-airship-research-station] Mobile completed views are crowded and crop much of the model behind instruction, chapter, and control panels.
  • [1000-piece-airship-research-station] The visualizer depends on external three.js CDNs and is not fully offline-ready.
Major UI Readability DefectOperator Major Primary ArtifactOperator Multiple Major Defects
Wall clock 3h 31m 54s

From the run

Artemis II Mission Visualization

A fact sheet plus an interactive 3D visualization of the Artemis II mission. Tests factual grounding, source integrity, and the ability to dramatize the hard beats — launch, staging, re-entry, recovery.

89
Excellent

An excellent and unusually complete Artemis II package with rigorous provenance labeling, broad mission storytelling, and a real, polished interactive 3D implementation. It falls short of near mastery because factual source support was not independently audited, the displayed continuous path is modelled rather than flown ephemeris, and the otherwise strong responsive design clips narrative text and crowds controls on mobile.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegrityResearch GroundingSemantic JudgmentQuant. Reas.Spatial Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeed
1Claude Opus 589
2GPT-5.6 Sol89
3GPT-5.6 Luna87
4Claude Fable 586
5GPT-5.6 Terra86
6Kimi K383
7GPT-5.579
8Grok 4.579
9Claude Opus 4.876
10Claude Sonnet 5 (xhigh)71
11Opus 4.760
12GLM 5.2 (OpenRouter)58
13Gemini 3.5 Flash (High) Fast54

What it nailed

  • Exceptional separation of sourced facts from modelled trajectory state, including explicit treatment of conflicting published figures.
  • A genuinely mission-specific 3D application with synchronized telemetry, timeline, cameras, event cards, staging, flyby, entry, parachute, and recovery states.
  • Comprehensive storytelling coverage extending beyond the headline flyby to checkout, proximity operations, anomalies, engineering decisions, and post-flight assessment.
  • Strong local packaging with vendored three.js, bundled imagery, build scripts, structured JSON/CSV, preset recording beats, and screenshot evidence.
  • Desktop visual design is cohesive and polished enough for publication screenshots and video capture.

Where it slipped

  • The mobile event description is visibly clipped, and the mobile timeline and control labels are crowded and small.
  • The continuous trajectory is a fitted, time-warped model despite an available NASA ephemeris, limiting its use as a flown-path source of truth.
  • No independent report verifies that the cited URLs resolve and substantively support every current or post-flight claim; several citations use generic index or homepage URLs.
  • The sphere-of-influence narration overstates the concept by saying lunar gravity became the dominant pull.
  • The run took about 50 minutes across 76 API calls, making speed a notable weakness.
Wall clock 50m 17s

From the run