v0.16

Can AI build consultant-grade decks?

I ran 13 model configurations across 4 consulting test cases and 53 published runs. Within each test case, every tool received the same prompt, source files, and constraints; each valid deck was scored on six dimensions: visual design, factual accuracy, narrative quality, clarity, editability, and data visualization.

Primary benchmarkCase 4 · AI and workforce transitionMaster-following test

13
model configurations
4
test cases
53
published runs
  • Benchmark deck by Claude Code / Sonnet 5
    Claude Code / Sonnet 5Case 4 · AI and workforce transition81/ 100
1 / 4Current deck preview, 1 of 4: Claude Code / Sonnet 5, Case 4 · AI and workforce transition, score 81 out of 100.

01 — LEADERBOARD

Quality first. Then time and cost.

Case 4 is the primary test: can an AI agent build an editable deck while following a supplied PowerPoint master and slide library? Its score combines task quality, factuality, narrative, clarity, visual design, and editability.

Model companyAnthropicDeepSeekGoogleMiniMaxMoonshot AIOpenAIQwen / AlibabaxAIZ.AI

Best quality

Case 4 overall quality · Higher is better

9 configurations · Case 4 only

Fastest

Average generation time in minutes · Lower is better

8 configurations · Case 4 only

Cheapest

Reported cost per complete deck · Lower is better

4 configurations · Case 4 only

Reported cost captured for 4 of 9 case 4 configurations. Missing values are not ranked.

Cost not captured: Opus 5 · Sonnet 5 · Sonnet 5 · GPT-5.6 Sol · DeepSeek V4 Flash Latest

Supporting evidence

General presentation capability

Cases 1–3 · no PowerPoint master supplied

Analysis, storytelling, charting, and source handling across the unconstrained test cases. This score is intentionally separate from Case 4 master adherence.

01Sonnet 5Anthropic · Claude Code67/ 1003 of 3 cases
02Opus 5Anthropic · Claude Code61/ 1003 of 3 cases

Limited evidence: Fable 5 (1 of 3 cases) · GPT-5.6 Sol (3 of 3 cases) · GPT-5.6 Terra (2 of 3 cases) · DeepSeek V4 Pro (1 of 3 cases) · Gemini 3.7 Flash (2 of 3 cases) · Grok 4.6 (1 of 3 cases) · Kimi K3 (1 of 3 cases) · MiniMax M3 (2 of 3 cases) · Qwen 3.8 Max (1 of 3 cases). These configurations are shown for coverage, not ranked.

Permanent evidence

Every Test Case 4 result stays visible.

Legacy publication data · legacy release view

The prompt is unchanged across releases. Each result keeps its generation version, review status, compliance findings, reported cost, API-equivalent estimate, and elapsed time, even when it is not part of the current ranking.

Test Case 4 generation and evaluation history
ResultConfigurationVersionQualityComplianceReported costEstimateTimeEvidence
1RecordedSonnet 5Claude Code / Sonnet 5Historical rubriclocal-v16local-v781 / 100Overall scoreUnverifiedChecks not publishedNot recordedProvider reportedNot recordedAPI-equivalent9m 43sGenerationView deck preview for Sonnet 5
View ledger

Score ledger

Clarity
8.75 / 10
Data Visualization
10 / 10
Editability
8.125 / 10
Factuality
8.75 / 10
Narrative
10 / 10
Visual Design
5 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

2RecordedOpus 5Claude Code / Opus 5Historical rubriclocal-v16local-v769 / 100Overall scoreUnverifiedChecks not publishedNot recordedProvider reportedNot recordedAPI-equivalent11m 41sGenerationView deck preview for Opus 5
View ledger

Score ledger

Clarity
7.5 / 10
Data Visualization
9.375 / 10
Editability
8.438 / 10
Factuality
10 / 10
Narrative
10 / 10
Visual Design
4.583 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

3RecordedGLM 5.3 FlashOpenCode + OpenRouter / GLM 5.3 FlashHistorical rubriclocal-v16local-v968 / 100Overall scoreUnverifiedChecks not published$0.05Provider reportedNot recordedAPI-equivalent42m 00sGenerationView deck preview for GLM 5.3 Flash
View ledger

Score ledger

Clarity
6.667 / 10
Data Visualization
8.333 / 10
Editability
7.917 / 10
Factuality
9.167 / 10
Narrative
9.167 / 10
Visual Design
3.889 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

4RecordedGPT-5.6 SolCodex / GPT-5.6 Sol / MaxHistorical rubriclocal-v16local-v767 / 100Overall scoreUnverifiedChecks not publishedNot recordedProvider reportedNot recordedAPI-equivalent10m 01sGeneration
View ledger

Score ledger

Clarity
7.083 / 10
Data Visualization
8.333 / 10
Editability
9.167 / 10
Factuality
10 / 10
Narrative
10 / 10
Visual Design
4.444 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

5RecordedGrok 4.6OpenCode + OpenRouter / Grok 4.6Historical rubriclocal-v16local-v1665 / 100Overall scoreUnverifiedChecks not published$0.77Provider reportedNot recordedAPI-equivalent8m 45sGeneration
View ledger

Score ledger

Clarity
6.875 / 10
Data Visualization
8.125 / 10
Editability
8.438 / 10
Factuality
9.375 / 10
Narrative
9.062 / 10
Visual Design
5 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

6RecordedSonnet 5Claude Code / Sonnet 5Historical rubriclocal-v16local-v1660 / 100Overall scoreUnverifiedChecks not publishedNot recordedProvider reportedNot recordedAPI-equivalent17m 51sGenerationView deck preview for Sonnet 5
View ledger

Score ledger

Clarity
10 / 10
Data Visualization
8.75 / 10
Editability
8.75 / 10
Factuality
10 / 10
Narrative
10 / 10
Visual Design
9.167 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

7RecordedDeepSeek V4 ProOpenCode + OpenRouter / DeepSeek V4 ProHistorical rubriclocal-v16local-v1660 / 100Overall scoreUnverifiedChecks not published$0.24Provider reportedNot recordedAPI-equivalent13m 56sGeneration
View ledger

Score ledger

Clarity
8.75 / 10
Data Visualization
7.5 / 10
Editability
9.375 / 10
Factuality
10 / 10
Narrative
7.5 / 10
Visual Design
5.833 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

8RecordedQwen 3.8 MaxOpenCode + OpenRouter / Qwen 3.8 MaxHistorical rubriclocal-v16local-v760 / 100Overall scoreUnverifiedChecks not published$0.81Provider reportedNot recordedAPI-equivalent16m 35sGeneration
View ledger

Score ledger

Clarity
8.75 / 10
Data Visualization
8.75 / 10
Editability
8.75 / 10
Factuality
10 / 10
Narrative
10 / 10
Visual Design
3.333 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

9RecordedDeepSeek V4 Flash LatestOpenCode + OpenRouter / DeepSeek V4 Flash LatestHistorical rubriclocal-v16local-v751 / 100Overall scoreUnverifiedChecks not publishedNot recordedProvider reportedNot recordedAPI-equivalentNot recordedGeneration
View ledger

Score ledger

Clarity
6.25 / 10
Data Visualization
7.5 / 10
Editability
7.5 / 10
Factuality
10 / 10
Narrative
10 / 10
Visual Design
0.833 / 10

PowerPoint compliance

Compliance checks are not published for this historical result.

Evidence view

Evidence view

02 — COMPARE

What do the decks look like?

Expanded published deck comparison

Claude Code / Sonnet 581/ 100
Claude Code / Sonnet 5, AI and workforce transition, slide 1
Claude Code / Opus 569/ 100
Claude Code / Opus 5, AI and workforce transition, slide 1

03 — HOW I TESTED

How I tested each deck.

What each AI tool was asked to do

  1. Structured

    Test case 1 · Freight modal split · Storyline supplied
  2. Interpretive

    Test case 2 · Road freight activity · Find the story
  3. Multi-source

    Test case 3 · Shipment operations · Reconcile evidence
  4. Strategic

    Test case 4 · AI and workforce transition · Build the argument

Every submission follows the same six-stage pipeline.

  1. Lock the task

    Freeze the exact prompt, redistributable sources, requirements, and reference boundary.

  2. Generate one deck

    Start a fresh local agent session and submit the task exactly once.

  3. Check validity and editability

    Inspect OOXML integrity, required content, source notes, and editable structure.

  4. Render in PowerPoint

    Export locally through Microsoft PowerPoint and preserve raw visual evidence.

  5. Blind-judge quality

    Send a blinded, sanitized bundle to the release-pinned judge models.

  6. Rank eligible results

    Rank every configuration with at least one eligible deck and incorporate all eligible repetitions.

Only Test Case 4 used Meridian Signal’s corporate template.

Test Case 4 was the only test case supplied with Meridian Signal’s corporate PowerPoint system. The reference library is split into the master layouts that define structure and the slide templates that show how to use them. Cases 1–3 had no corporate template, so their visual-design scores are not directly comparable with Case 4.