A model in the suite · Alibaba

Qwen3.8 27B

Alibaba · Qwen3.8 · Qwen3.8 27B via OpenRouter / xhigh reasoning (model default, vendor-recommended) / streaming transport + harness Perplexity web-search accommodation / Suite 2.0 API harness · 2026-08-17

3 of 4 benchmarks — Brick could not complete

Brick (The AI LEGO Build) did not complete and is not scored. Qwen3.8 27B produced zero artifacts on it across four attempts, each failing on the very first reasoning turn of the first isolated case: 900s non-streaming; a malformed JSON response; 900s streaming with the harness Perplexity search tool; and 1800s streaming with compaction available. This model has exactly one upstream provider on OpenRouter (AkashML) with no fallback route and no published throughput telemetry, and it never returned that first long response. The other three benchmarks completed normally under the same configuration, so this reads as a route limitation rather than a measured model capability. The suite average shown is over the 3 benchmarks that scored; it is not a 4-benchmark result and should not be compared as one.

74/100
Strict suite averageNo legacy score · 3 benchmarks

Qwen3.8 27B via OpenRouter / xhigh reasoning (model default, vendor-recommended) / streaming transport + harness Perplexity web-search accommodation / Suite 2.0 API harness

No transcript compaction

This run kept the entire transcript in context on every turn, with nothing evicted, so usable context shrank as the run progressed. Models with smaller context windows are disadvantaged under this policy. Compaction was added to the harness on 2026-08-16; results marked “compaction enabled” are not directly comparable on this dimension.

Copies Qwen3.8 27B's full data pack — paste it into ChatGPT, Claude, or any AI to talk it through.

How Qwen3.8 27B handled each benchmark

Score, capability radar, and the honest read on what it nailed and where it slipped. Hit Overlay to drop other models onto the same axes.

Dingo & Co. Knowledge Work

A 23-deliverable consulting brief: research, financial reconciliation, regulatory analysis, decks and spreadsheets. Tests whether a model can run an entire knowledge-work engagement end to end.

89
Excellent

Excellent Dingo & Co. run. It completed the full package in real formats and showed unusually strong judgment on the benchmark's legal, ethical, regulatory, import, TAM, quote-permission, and tone traps. The work is strategically useful and close to board-ready as a draft. It falls short of near-mastery because a human QA pass is still clearly needed: the dashboard has a misleading chart axis, one workbook row carries a material TAM typo, and a few visible copy/spec errors remain.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegrityResearch GroundingSemantic JudgmentQuant. Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeedSpatial Reas.
1GPT-5.6 Sol93
2Grok 4.693
3Kimi K392
4Grok 4.591
5GPT-5.6 Luna90
6Qwen3.8 27B89
7GPT-5.6 Terra88
8GLM 5.2 (OpenRouter)88
9Claude Opus 586
10Claude Fable 581
11Claude Sonnet 5 (xhigh)81
12Claude Opus 4.880
13GPT-5.578
14Gemini 3.5 Flash (High) Fast62
15Opus 4.754
16Sonnet 4.652
17Gemini 3.1 Pro38

What it nailed

  • Complete artifact set with real DOCX, PPTX, XLSX, PDF, HTML, markdown, JSON, images, screenshots, and manifest outputs.
  • Excellent handling of the benchmark's central absurdities: dingo/litter-box fit, Alaska/Australia mismatch, import-market creation, legal ambiguity, ethics, reputation, and support-language liability.
  • Strong citation posture with official/legal sources separated from commercial sources and fictional competitors labeled as fictional/unverified.
  • Impressive strategic package: gated budget, GTM plan, investor FAQ, risk matrix, pricing rationale, KPIs, and channel-specific messaging.
  • Copy and personas are tailored to the actual niche rather than generic pet-tech filler.

Where it slipped

  • Dashboard revenue/CAC chart has misleading left-axis tick labels, a minor visual/data-trust defect confirmed by operator visual review.
  • Competitive workbook contains an isolated but material TAM inconsistency: '~$8M/yr hardware ceiling' where the package's own math and other documents use ~$840K/year.
  • A few QA issues would embarrass a polished board package, including 'dino-hybrid' in the executive summary and a likely inconsistent Pro dimension row in an email table.
  • Some financial modeling is simplified and appears partly calibrated to target outputs rather than fully independently derived.
  • Full source-integrity hash verification and exhaustive independent citation verification were not provided in the validation evidence.
Wall clock 1h 52m 31s

From the run

Car Wash Operations

A filthy operational dataset — ghost records, orphaned orders, typo'd customers, raw enum variants. Tests judgment under messy real-world data: what gets fixed, quarantined, or wrongly promoted.

77
Strong

A strong, reviewer-ready migration scaffold with broad digital extraction, a real SQLite database, extensive provenance/review infrastructure, and a polished static UI. It is not production-clean: handwritten image records are mostly not extracted, some planted multimodal facts are missed or mislabeled, department semantics are weak, and referential integrity is not enforced with database constraints. Overall, it is useful for human audit and repair, but not trustworthy as an autonomous final migration.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegritySemantic JudgmentQuant. Reas.UX ReviewabilityProd. ReadinessSpeed
1Claude Fable 588
2Claude Opus 4.886
3Grok 4.684
4Qwen3.8 27B77
5Claude Opus 575
6Kimi K365
7Claude Sonnet 5 (xhigh)64
8GPT-5.6 Sol55
9GPT-5.6 Terra55
10GPT-5.555
11GPT-5.6 Luna55
12Grok 4.555
13GLM 5.2 (OpenRouter)55
14Gemini 3.5 Flash (High) Fast51
15GPT-5.451
16Opus 4.748

What it nailed

  • Complete required artifact set with an openable SQLite database, migration script, reports, static UI, and screenshots.
  • Strong audit/review architecture: source inventory, conflicts, rejected records, review queue, canaries, and provenance are first-class.
  • Good handling of many digital normalization problems: services, statuses, payments, dates, price eras, corrupted JSON, and duplicate image byte detection.
  • Reviewer UI renders cleanly on desktop and mobile with no operator-reported visual defects.
  • Sensitive credential-like file is inventoried without exposing contents.

Where it slipped

  • Handwritten/image receipt extraction is materially weak: the package reports no usable OCR recovery for 14 images and misses several planted image-derived business facts.
  • Some image context appears hard-coded and partially inaccurate, which reduces trust in multimodal grounding.
  • No explicit SQLite foreign-key constraints are declared, despite referential-integrity goals.
  • Department/role-code normalization is not clearly represented in the canonical schema.
  • Evidence for all planted typo-order merges is incomplete, and the final customer/job counts suggest unresolved duplicate or transaction-only entities remain.
Wall clock 37m 16s

From the run

Artemis II Mission Visualization

55
Interesting but Unreliable

This is an ambitious and mostly complete Artemis II visualization package with strong desktop WebGL presentation, broad mission-event coverage, and substantial source-backed factual material. However, the primary visualization has multiple operator-documented major mobile responsive/readability defects, including overlapped title/HUD panels, clipping, inaccessible phases, and controls that dominate the viewport. Under v2 hard caps, those primary-artifact defects cap the strict score at 55 despite the stronger desktop and research work.

OverlayDownload radar
Instr. FollowingArtifact ValiditySource IntegrityResearch GroundingSemantic JudgmentQuant. Reas.Spatial Reas.Visual StorytellingUX ReviewabilityProd. ReadinessSpeed
1Claude Opus 589
2GPT-5.6 Sol89
3GPT-5.6 Luna87
4Claude Fable 586
5GPT-5.6 Terra86
6Kimi K383
7GPT-5.579
8Grok 4.579
9Claude Opus 4.876
10Claude Sonnet 5 (xhigh)71
11Opus 4.760
12GLM 5.2 (OpenRouter)58
13Qwen3.8 27B55
14Grok 4.655
15Gemini 3.5 Flash (High) Fast54

What it nailed

  • Complete required artifact set with fact sheet, source list, visualization, documentation, app code, research archive, and screenshot evidence paths.
  • Desktop visualization is nonblank and coherent, with Earth/Moon framing, trajectory, HUD, mission phase list, event markers, and playback controls.
  • Substantial mission coverage across launch, ascent, staging, TLI, coast, lunar flyby, max-distance record, return, re-entry, splashdown, and recovery.
  • Research package is mostly grounded in official NASA/CSA/ESA-style material and includes local archived source pages.
  • Visualization data model includes many source-attributed events and a synchronized mission timeline.

Where it slipped

  • Primary visualization has operator-documented major mobile layout defects: overlapping panels, clipped/covered title content, horizontally clipped scene context, inaccessible phase structure, and oversized controls.
  • Visual/UX production readiness is not publication-grade across responsive viewports even though the desktop render is attractive.
  • Fact sheet contains some public-facing precision/accuracy issues, including the 'Apollo 2 astronaut Charlie Duke' typo and minor quantitative/timeline inconsistencies.
  • Citation discipline is good at the package level but not fully independently link-validated in the provided evidence.
  • Packaging has minor inconsistencies, including the Visualization README referring to 24 EV events while the artifact otherwise describes 26.
Operator Major Primary ArtifactOperator Multiple Major DefectsMajor UI Readability Defect
Wall clock 30m 32s

From the run