A model in the suite · OpenAI

GPT-5.5

OpenAI · GPT · historical runs from 2026-04-23 through 2026-06-01 staging

71/100
Strict suite averageLegacy 81 · 3 benchmarks

Strongest historical non-image packet in this backfill set, led by Dingo and Artemis. The Car Wash result keeps the strict average grounded because operational canary misses remain substantial; the Brick rover is preserved only as a single-prompt reference.

No transcript compaction

This run kept the entire transcript in context on every turn, with nothing evicted, so usable context shrank as the run progressed. Models with smaller context windows are disadvantaged under this policy. Compaction was added to the harness on 2026-08-16; results marked “compaction enabled” are not directly comparable on this dimension.

Copies GPT-5.5's full data pack — paste it into ChatGPT, Claude, or any AI to talk it through.

How GPT-5.5 handled each benchmark

Score, capability radar, and the honest read on what it nailed and where it slipped. Hit Overlay to drop other models onto the same axes.

Dingo & Co. Knowledge Work

A 23-deliverable consulting brief: research, financial reconciliation, regulatory analysis, decks and spreadsheets. Tests whether a model can run an entire knowledge-work engagement end to end.

78legacy 87
Strong

This is the cleanest historical Dingo run: all 23 deliverables exist as real files, source integrity passed, regulatory/import ambiguity was handled unusually well, and the strategy work is coherent. It stays below excellent strict territory because the board deck has a real PPTX XML/rendering defect, one visible NPS inconsistency crosses artifacts, pricing research includes stale or imprecise claims, and there is no raw transcript evidence.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegrityResearch GroundingSemantic JudgmentQuant. Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeedSpatial Reas.
1GPT-5.6 Sol93
2Grok 4.693
3Kimi K392
4Ox Alpha92
5Grok 4.591
6GPT-5.6 Luna90
7Qwen3.8 27B89
8GPT-5.6 Terra88
9GLM 5.2 (OpenRouter)88
10Claude Opus 586
11Claude Fable 581
12Claude Sonnet 5 (xhigh)81
13Claude Opus 4.880
14GPT-5.578
15Gemini 3.5 Flash (High) Fast62
16Opus 4.754
17Sonnet 4.652
18Gemini 3.1 Pro38

What it nailed

  • Completed all 23 required deliverables as real files with valid types and preserved source integrity.
  • Handled dingo ownership, import-created demand, legal uncertainty, ethics, and Alaska/Australia mismatch as central operating constraints.
  • Used provided source imagery heavily in the deck, sales one-pager, and dashboard.
  • Delivered coherent GTM, board, pricing, risk, and investor-facing strategy with staged decisions and guardrails.

Where it slipped

  • Board deck contains invalid PPTX metadata XML because the Company value uses an unescaped ampersand, blocking Quick Look rendering.
  • Board deck slide 5 reports average NPS as 6.6 while source math and other artifacts use about 6.2.
  • Some pricing research was stale or imprecise, especially Halo membership pricing and PetSafe blended pricing.
  • Raw model output/transcript evidence is absent from the evaluation package.
Pptx Metadata Xml Rendering DefectCross Document Number DriftStale Or Imprecise Pricing ClaimsEmpty Raw Model Output Evidence Confidence Cap

Car Wash Operations

A filthy operational dataset — ghost records, orphaned orders, typo'd customers, raw enum variants. Tests judgment under messy real-world data: what gets fixed, quarantined, or wrongly promoted.

55legacy 74
Interesting but Unreliable

This is the strongest inspected audit scaffold: complete artifacts, full source discovery, a working frontend, strong provenance, fake/test rejection, and an empirically verified idempotent rebuild. It still fails too many operational canaries for a high strict score: Terrence Blackwood became a canonical customer, SVC-007 was missed, department codes were dropped, status/payment enums stayed raw, canonical jobs were overcounted, and several name variants stayed split. The run reaches the cap for multiple primary canary misses but not above it.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegritySemantic JudgmentQuant. Reas.UX ReviewabilityProd. ReadinessSpeed
1Claude Fable 588
2Claude Opus 4.886
3Grok 4.684
4Qwen3.8 27B77
5Claude Opus 575
6Kimi K365
7Claude Sonnet 5 (xhigh)64
8GPT-5.6 Sol55
9GPT-5.6 Terra55
10GPT-5.555
11GPT-5.6 Luna55
12Grok 4.555
13GLM 5.2 (OpenRouter)55
14Ox Alpha54
15Gemini 3.5 Flash (High) Fast51
16GPT-5.451
17Opus 4.748

What it nailed

  • Produced every expected artifact, including screenshots.
  • Discovered 465 of 465 source files and processed or partially processed almost all business-relevant files.
  • Rejected planted ghost/test records and preserved a large source-record provenance layer.
  • Passed an isolated idempotency rerun with identical counts.

Where it slipped

  • Created Terrence Blackwood as a canonical customer instead of an orphan review case.
  • Missed the DeShawn SVC-007 conflict and lacked a service-code column.
  • Dropped department/role code normalization and left raw status/payment values.
  • Overcounted jobs and retained several duplicate/nickname customer splits.
Misses Three Or More Primary CanariesPromotes Orphan Order

Artemis II Mission Visualization

79legacy 79
Strong

The run produced a complete, runnable React/Vite/Three package with a separately maintained missionData.js source of truth, a detailed fact sheet, NASA-heavy citations, screenshots, desktop/mobile verification images, and mission-specific visual beats for launch, ascent, staging, TLI, lunar flyby, max distance, re-entry, splashdown, and recovery. It stays below excellent because several values need strict current-source cleanup, including closest lunar approach finalization, actual ascent milestone timings, official total miles after NASA's May 7 update, and a non-primary Orion helium-leak claim. The re-entry and recovery visuals are informative but still more app-like than cinematic.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegrityResearch GroundingSemantic JudgmentQuant. Reas.Spatial Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeed
1Claude Opus 589
2GPT-5.6 Sol89
3GPT-5.6 Luna87
4Claude Fable 586
5GPT-5.6 Terra86
6Kimi K383
7GPT-5.579
8Grok 4.579
9Claude Opus 4.876
10Claude Sonnet 5 (xhigh)71
11Ox Alpha60
12Opus 4.760
13GLM 5.2 (OpenRouter)58
14Qwen3.8 27B55
15Grok 4.655
16Gemini 3.5 Flash (High) Fast54

What it nailed

  • Complete fact sheet plus runnable React/Vite/Three visualization.
  • Separates mission facts, crew, vehicle facts, component details, events, and telemetry into missionData.js.
  • Uses many NASA source links directly in both fact sheet and visualization.
  • Covers the hard mission sequence rather than staying in generic orbit-only mode.
  • Includes 10 staged screenshots and desktop/mobile verification images.

Where it slipped

  • No formal historical scorecard or raw model output was found.
  • Some public-facing numbers need current-source reconciliation before publication.
  • Re-entry, splashdown, and recovery scenes are visually abstract and less video-useful than the stronger Opus visual treatment.
  • At least one anomaly/issue claim relies on non-primary reporting.

From the run