Skip to content

Tested prompt · Data analysis

Correlation between ad spend and sign-ups: every AI model's reply, tested

We sent this hard data analysis prompt to all 16 models in llmwise, the same way the app sends a message, and checked every reply the same way. Here's each one as it came, with whether it passed, what it cost and how long it took.

Based on 16 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .

Short answer

12 of 16 models passed this data analysis prompt's check (final answer); Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.8 Flash, and GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.00037; the fastest, Claude Sonnet 5.5's in 4.2 s. The dearest reply, Claude Fable 5.1's, cost 179 times as much ($0.0663).

The prompt, as sent, and its check

Checked by final answer, the same way for every model.

Correlation between ad spend and sign-ups (hard)

Here's weekly ad spend and sign-ups as CSV. What's the Pearson correlation coefficient between ad_spend and signups, to 2 decimal places?

week,ad_spend,signups
1,500,42
2,800,55
3,650,49
4,1200,71
5,900,60
6,300,35
7,1100,64
8,700,58

Show your working briefly, then end with a line that says "Final answer: " followed by the answer alone.

The final answer must be 0.97.

Exactly what this prompt's replies are checked against, with every other prompt of our test runs.

Every model's result

All 16 models on this prompt, in catalog order.

Every model's reply to “Correlation between ad spend and sign-ups”
ModelResultCostTimeReply
Claude Fable 5.1AnthropicPassed: Final answer 0.97: right.$0.066314.2 s202 tokens
Claude Opus 5.5AnthropicPassed: Final answer 0.97: right.$0.026111.5 s209 tokens
Claude Sonnet 5.5AnthropicPassed: Final answer 0.97: right.$0.00704.2 s547 tokens
Claude Sonnet 5AnthropicFailed: Final answer 0.99; expected 0.97.$0.016513.8 s176 tokens
Claude Haiku 4.5AnthropicFailed: Final answer 0.99; expected 0.97.$0.00375.0 s642 tokens
GPT-6 AstraOpenAIPassed: Final answer 0.97: right.$0.02948.8 s159 tokens
GPT-6 SolOpenAIPassed: Final answer 0.97: right.$0.00696.6 s100 tokens
GPT-6 LunaOpenAIPassed: Final answer 0.97: right.$0.000376.7 s99 tokens
Gemini 3.1 Pro (preview)GooglePassed: Final answer 0.97: right.$0.031016.5 s588 tokens
Gemini 3.8 FlashGoogleFailed: Final answer 0.91; expected 0.97.$0.00375.4 s887 tokens
DeepSeek V4.1 FlashDeepSeekPassed: Final answer 0.97: right.$0.00134.8 s213 tokens
DeepSeek V4 ProDeepSeekPassed: Final answer 0.97: right.$0.002141.2 s121 tokens
Grok 4.7xAIPassed: Final answer 0.97: right.$0.022245.4 s177 tokens
Kimi K3MoonshotPassed: Final answer 0.97: right.$0.013120.2 s454 tokens
GLM 5.3Z.aiPassed: Final answer 0.97: right.$0.00625.6 s140 tokens
GLM 5.3 FlashZ.aiFailed: Final answer 0.99; expected 0.97.$0.0005419.4 s684 tokens

Cost: what OpenRouter charged us for the reply. Time: from sending to the whole reply. Reply: its length in tokens, thinking not counted. In llmwise you pay per message, not per token: each of these replies counted as one message on Pro.

Every reply

The 4 replies that failed first, then the ones that passed, from the cheapest up.

  1. GLM 5.3 Flash

    Failed: Final answer 0.99; expected 0.97.

    Working:
    - Mean ad_spend = 6150/8 = 768.75; mean signups = 434/8 = 54.25
    - Σ(dx·dy) = 24428.125
    - Σdx² = 624687.5; Σdy² = 971.5
    - r = 24428.125 / √(624687.5 × 971.5) = 24428.125 / √606,871,906.25 ≈ 24428.125 / 24634.8 ≈ 0.9916
    
    Final answer: 0.99

    468 tokens in, 1,022 out (338 of them reasoning) · 19.4 s · $0.0005 · 1 message on Pro · answered by z-ai/glm-5.3-flash via Relace ·

  2. Claude Haiku 4.5

    Failed: Final answer 0.99; expected 0.97.

    I'll calculate the Pearson correlation coefficient between ad_spend and signups.
    
    **Data:**
    - ad_spend: 500, 800, 650, 1200, 900, 300, 1100, 700
    - signups: 42, 55, 49, 71, 60, 35, 64, 58
    
    **Calculations:**
    
    Mean of ad_spend: (500+800+650+1200+900+300+1100+700)/8 = 6150/8 = 768.75
    
    Mean of signups: (42+55+49+71+60+35+64+58)/8 = 434/8 = 54.25
    
    **Deviations and products:**
    | Week | ad_spend - 768.75 | signups - 54.25 | Product |
    |------|------------------|-----------------|---------|
    | 1 | -268.75 | -12.25 | 3291.19 |
    | 2 | 31.25 | 0.75 | 23.44 |
    | 3 | -118.75 | -5.25 | 623.44 |
    | 4 | 431.25 | 16.75 | 7223.44 |
    | 5 | 131.25 | 5.75 | 755.19 |
    | 6 | -468.75 | -19.25 | 9023.44 |
    | 7 | 331.25 | 9.75 | 3229.69 |
    | 8 | -68.75 | 3.75 | -257.81 |
    
    Sum of products: 23,891.98
    
    **Variance calculations:**
    - Variance of ad_spend: (72031.25 + 976.56 + 14101.56 + 185975.56 + 17226.56 + 219726.56 + 109726.56 + 4726.56) / 8 = 107995.31
    - Variance of signups: (150.06 + 0.56 + 27.56 + 280.56 + 33.06 + 370.56 + 95.06 + 14.06) / 8 = 121.63
    
    **Correlation:**
    r = 23,891.98 / √(107,995.31 × 121.63) = 23,891.98 / √13,128.46 = 23,891.98 / 114.57 = 0.9875 ≈ **0.99**
    
    Final answer: 0.99

    482 tokens in, 642 out · 5.0 s · $0.0037 · 1 message on Pro · answered by anthropic/claude-haiku-4.5 via Amazon Bedrock ·

  3. Gemini 3.8 Flash

    Failed: Final answer 0.91; expected 0.97.

    Let $x$ represent `ad_spend` and $y$ represent `signups`. The number of data points is $n = 8$.
    
    **Data:**
    - $x = [500, 800, 650, 1200, 900, 300, 1100, 700]$
    - $y = [42, 55, 49, 71, 60, 35, 64, 58]$
    
    **Summary Calculations:**
    - $\sum x = 500 + 800 + 650 + 1200 + 900 + 300 + 1100 + 700 = 6150$
    - $\sum y = 42 + 55 + 49 + 71 + 60 + 35 + 64 + 58 = 434$
    - $\sum x^2 = 500^2 + 800^2 + 650^2 + 1200^2 + 900^2 + 300^2 + 1100^2 + 700^2 = 5,342,500$
    - $\sum y^2 = 42^2 + 55^2 + 49^2 + 71^2 + 60^2 + 35^2 + 64^2 + 58^2 = 24,676$
    - $\sum xy = (500)(42) + (800)(55) + (650)(49) + (1200)(71) + (900)(60) + (300)(35) + (1100)(64) + (700)(58)$
      $= 21,000 + 44,000 + 31,850 + 85,200 + 54,000 + 10,500 + 70,400 + 40,600 = 357,550$
    
    **Pearson Correlation Formula:**
    $$r = \frac{n\sum xy - (\sum x)(\sum y)}{\sqrt{[n\sum x^2 - (\sum x)^2][n\sum y^2 - (\sum y)^2]}}$$
    
    - **Numerator:**
      $8(357,550) - (6150)(434) = 2,860,400 - 2,669,100 = 191,300$
    
    - **Denominator:**
      $n\sum x^2 - (\sum x)^2 = 8(5,342,500) - 6150^2 = 42,740,000 - 37,822,500 = 4,917,500$
      $n\sum y^2 - (\sum y)^2 = 8(24,676) - 434^2 = 197,408 - 188,356 = 9,052$
      $\sqrt{4,917,500 \times 9,052} = \sqrt{44,513,210,000} \approx 210,981.54$
    
    - **Correlation ($r$):**
      $r = \frac{191,300}{210,981.54} \approx 0.9067 \approx 0.91$
    
    Final answer: 0.91

    505 tokens in, 887 out · 5.4 s · $0.0037 · 1 message on Pro · answered by google/gemini-3.8-flash via Google ·

  4. Claude Sonnet 5

    Failed: Final answer 0.99; expected 0.97.

    **Working:**
    
    n=8, mean(x)=768.75, mean(y)=54.25
    
    Computed deviations and cross-products:
    - Σ(dx·dy) = 24428.125
    - Σ(dx²) = 624687.5
    - Σ(dy²) = 971.5
    
    r = Σ(dx·dy) / √(Σdx² · Σdy²) = 24428.125 / √(624687.5 × 971.5) = 24428.125 / 24635.01 ≈ 0.9916
    
    **Final answer: 0.99**

    605 tokens in, 1,524 out (1,348 of them reasoning) · 13.8 s · $0.0165 · 1 message on Pro · answered by anthropic/claude-sonnet-5 via Claude Platform on AWS ·

  5. GPT-6 Luna

    Passed: Final answer 0.97: right.

    Using \(r=\frac{\sum (x-\bar{x})(y-\bar{y})}{\sqrt{\sum(x-\bar{x})^2\sum(y-\bar{y})^2}}\), the centered sums are \(23{,}912.5\), \(624{,}687.5\), and \(971.5\). Thus \(r \approx 0.9707\).
    
    Final answer: 0.97

    451 tokens in, 651 out (552 of them reasoning) · 6.7 s · $0.0004 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·

  6. DeepSeek V4.1 Flash

    Passed: Final answer 0.97: right.

    **Working**
    
    n = 8, Σx = 6150, Σy = 434, Σxy = 357,550, Σx² = 5,352,500, Σy² = 24,516
    
    Numerator: n·Σxy − Σx·Σy = 8(357,550) − 6150(434) = 2,860,400 − 2,669,100 = 191,300
    
    Denominator:
    - n·Σx² − (Σx)² = 42,820,000 − 37,822,500 = 4,997,500
    - n·Σy² − (Σy)² = 196,128 − 188,356 = 7,772
    - √(4,997,500 × 7,772) ≈ 197,080.1
    
    r = 191,300 / 197,080.1 ≈ 0.9707
    
    Final answer: 0.97

    492 tokens in, 1,515 out (1,302 of them reasoning) · 4.8 s · $0.0013 · 1 message on Pro · answered by deepseek/deepseek-v4.1-flash via Novita ·

  7. DeepSeek V4 Pro

    Passed: Final answer 0.97: right.

    Means: ad_spend = 768.75, signups = 54.25
    
    - \(\sum (x-\bar{x})(y-\bar{y}) = 23912.5\)
    - \(\sum (x-\bar{x})^2 = 624687.5\)
    - \(\sum (y-\bar{y})^2 = 971.5\)
    
    \[
    r = \frac{23912.5}{\sqrt{624687.5 \times 971.5}} \approx 0.9707
    \]
    
    Final answer: 0.97

    464 tokens in, 2,841 out (2,720 of them reasoning) · 41.2 s · $0.0021 · 1 message on Pro · answered by deepseek/deepseek-v4-pro-0813 via Baidu ·

  8. GLM 5.3

    Passed: Final answer 0.97: right.

    **Working:**
    - n = 8; mean ad_spend = 768.75, mean signups = 54.25
    - Σ(x−x̄)(y−ȳ) = 23,912.5
    - Σ(x−x̄)² = 624,687.5; Σ(y−ȳ)² = 971.5
    - r = 23,912.5 / √(624,687.5 × 971.5) = 23,912.5 / 24,635 ≈ 0.9707
    
    Final answer: 0.97

    491 tokens in, 1,336 out (1,196 of them reasoning) · 5.6 s · $0.0062 · 1 message on Pro · answered by z-ai/glm-5.3 via Wafer ·

  9. GPT-6 Sol

    Passed: Final answer 0.97: right.

    Using Pearson’s formula, \(r=\frac{\sum(x-\bar x)(y-\bar y)}{\sqrt{\sum(x-\bar x)^2\sum(y-\bar y)^2}}\). Here, the three sums are \(23{,}912.5\), \(624{,}687.5\), and \(971.5\), respectively, giving \(r\approx0.9707\).
    
    Final answer: 0.97

    451 tokens in, 602 out (502 of them reasoning) · 6.6 s · $0.0069 · 1 message on Pro · answered by openai/gpt-6-sol via OpenAI ·

  10. Claude Sonnet 5.5

    Passed: Final answer 0.97: right.

    Data: x = ad_spend, y = signups (n = 8)
    
    Sums:
    - Σx = 500+800+650+1200+900+300+1100+700 = 6150 → mean x = 768.75
    - Σy = 42+55+49+71+60+35+64+58 = 434 → mean y = 54.25
    
    Deviations (dx, dy), products:
    - 500: dx=-268.75, dy=-12.25 → dxdy=3292.19; dx²=72226.56; dy²=150.06
    - 800: dx=31.25, dy=0.75 → 23.44; 976.56; 0.56
    - 650: dx=-118.75, dy=-5.25 → 623.44; 14101.56; 27.56
    - 1200: dx=431.25, dy=16.75 → 7223.44; 185976.56; 280.56
    - 900: dx=131.25, dy=5.75 → 754.69; 17226.56; 33.06
    - 300: dx=-468.75, dy=-19.25 → 9023.44; 219726.56; 370.56
    - 1100: dx=331.25, dy=9.75 → 3229.69; 109726.56; 95.06
    - 700: dx=-68.75, dy=3.75 → -257.81; 4726.56; 14.06
    
    Totals:
    - Sxy ≈ 23,912.5
    - Sxx ≈ 624,687.5
    - Syy ≈ 971.5
    
    r = 23,912.5 / √(624,687.5 × 971.5) = 23,912.5 / √606,884,000 ≈ 23,912.5 / 24,635.4 ≈ 0.97
    
    Final answer: 0.97

    609 tokens in, 547 out · 4.2 s · $0.0070 · 1 message on Pro · answered by anthropic/claude-sonnet-5.5 via Anthropic ·

  11. Kimi K3

    Passed: Final answer 0.97: right.

    ## Working
    
    **Means:** x̄ = 6150/8 = 768.75, ȳ = 434/8 = 54.25
    
    **Deviations and sums:**
    
    | week | x−x̄ | y−ȳ | (x−x̄)(y−ȳ) | (x−x̄)² | (y−ȳ)² |
    |---|---|---|---|---|---|
    | 1 | −268.75 | −12.25 | 3292.19 | 72226.56 | 150.06 |
    | 2 | 31.25 | 0.75 | 23.44 | 976.56 | 0.56 |
    | 3 | −118.75 | −5.25 | 623.44 | 14101.56 | 27.56 |
    | 4 | 431.25 | 16.75 | 7223.44 | 185976.56 | 280.56 |
    | 5 | 131.25 | 5.75 | 754.69 | 17226.56 | 33.06 |
    | 6 | −468.75 | −19.25 | 9023.44 | 219726.56 | 370.56 |
    | 7 | 331.25 | 9.75 | 3229.69 | 109726.56 | 95.06 |
    | 8 | −68.75 | 3.75 | −257.81 | 4726.56 | 14.06 |
    | **Σ** | | | **23912.50** | **624687.50** | **971.50** |
    
    **Compute r:**
    
    r = 23912.5 / √(624687.5 × 971.5) = 23912.5 / √(606,883,906.25) = 23912.5 / 24635.0 ≈ 0.9707
    
    Rounded to 2 decimal places: 0.97 (a very strong positive correlation).
    
    Final answer: 0.97

    537 tokens in, 1,418 out (964 of them reasoning) · 20.2 s · $0.0131 · 1 message on Pro · answered by moonshotai/kimi-k3 via Wafer ·

  12. Grok 4.7

    Passed: Final answer 0.97: right.

    **Pearson r (computational formula)**
    
    n = 8, x = ad_spend, y = signups
    
    - Σx = 6150, Σy = 434, Σxy = 357550, Σx² = 5352500, Σy² = 24516
    
    Numerator: nΣxy − ΣxΣy = 8×357550 − 6150×434 = 191300
    
    Denominator parts: nΣx² − (Σx)² = 4997500, nΣy² − (Σy)² = 7772
    
    r = 191300 / √(4997500 × 7772) = 191300 / √38840570000 ≈ 191300 / 197080.11 ≈ 0.9707
    
    **Final answer: 0.97**

    1,686 tokens in, 4,422 out (4,245 of them reasoning) · 45.4 s · $0.0222 · 1 message on Pro · answered by x-ai/grok-4.7 via xAI ·

  13. Claude Opus 5.5

    Passed: Final answer 0.97: right.

    **Means**
    - Mean ad_spend: 6150 / 8 = 768.75
    - Mean signups: 434 / 8 = 54.25
    
    **Sums of deviations**
    - Σ(dx·dy) = 23,912.5
    - Σ(dx²) = 624,687.5
    - Σ(dy²) = 971.5
    
    **Pearson r**
    
    r = 23,912.5 / √(624,687.5 × 971.5)
    = 23,912.5 / √606,883,906.25
    = 23,912.5 / 24,635.01
    ≈ 0.9707
    
    This is a strong positive correlation.
    
    Final answer: 0.97

    607 tokens in, 1,155 out (946 of them reasoning) · 11.5 s · $0.0261 · 1 message on Pro · answered by anthropic/claude-opus-5.5 via Claude Platform on AWS ·

  14. GPT-6 Astra

    Passed: Final answer 0.97: right.

    Pearson’s correlation coefficient is
    \[
    r=\frac{n\sum xy-(\sum x)(\sum y)}
    {\sqrt{[n\sum x^2-(\sum x)^2][n\sum y^2-(\sum y)^2]}}.
    \]
    
    For these 8 weeks:
    - \(\sum x=6150,\quad \sum y=434\)
    - \(\sum xy=357550\)
    - \(\sum x^2=5352500,\quad \sum y^2=24516\)
    
    Thus,
    \[
    r=\frac{191300}{\sqrt{4997500\times7772}}\approx0.9707.
    \]
    
    Final answer: 0.97

    451 tokens in, 497 out (338 of them reasoning) · 8.8 s · $0.0294 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·

  15. Gemini 3.1 Pro

    Passed: Final answer 0.97: right.

    To calculate the Pearson correlation coefficient ($r$), we use the formula:
    $r = \frac{n(\sum xy) - (\sum x)(\sum y)}{\sqrt{[n\sum x^2 - (\sum x)^2][n\sum y^2 - (\sum y)^2]}}$
    
    Given the data for ad_spend ($x$) and signups ($y$):
    - $n = 8$
    - $\sum x = 500 + 800 + 650 + 1200 + 900 + 300 + 1100 + 700 = 6,150$
    - $\sum y = 42 + 55 + 49 + 71 + 60 + 35 + 64 + 58 = 434$
    - $\sum xy = (500 \times 42) + \dots + (700 \times 58) = 357,550$
    - $\sum x^2 = (500^2) + \dots + (700^2) = 5,352,500$
    - $\sum y^2 = (42^2) + \dots + (58^2) = 24,516$
    
    Plugging these into the formula:
    Numerator: $8(357,550) - (6,150)(434) = 2,860,400 - 2,669,100 = 191,300$
    Denominator $x$: $8(5,352,500) - (6,150)^2 = 42,820,000 - 37,822,500 = 4,997,500$
    Denominator $y$: $8(24,516) - (434)^2 = 196,128 - 188,356 = 7,772$
    
    $r = \frac{191,300}{\sqrt{4,997,500 \times 7,772}}$
    $r = \frac{191,300}{\sqrt{38,840,570,000}}$
    $r = \frac{191,300}{197,080.11} \approx 0.97067$
    
    Final answer: 0.97

    505 tokens in, 2,496 out (1,908 of them reasoning) · 16.5 s · $0.0310 · 1 message on Pro · answered by google/gemini-3.1-pro-preview via Google ·

  16. Claude Fable 5.1

    Passed: Final answer 0.97: right.

    **Working**
    
    - Means: x̄ (ad_spend) = 6150/8 = 768.75; ȳ (signups) = 434/8 = 54.25
    - Σ(x−x̄)(y−ȳ) = 23,912.5
    - Σ(x−x̄)² = 624,687.5
    - Σ(y−ȳ)² = 971.5
    - r = 23,912.5 / √(624,687.5 × 971.5) = 23,912.5 / 24,635.0 ≈ 0.9707
    
    Strong positive correlation between ad spend and sign-ups.
    
    Final answer: 0.97

    607 tokens in, 1,177 out (975 of them reasoning) · 14.2 s · $0.0663 · 1 message on Pro · answered by anthropic/claude-fable-5.1 via Anthropic ·

More data analysis prompts

The other data analysis prompts, each with every model's reply.

Questions

Which AI does best on “Correlation between ad spend and sign-ups”?

12 of 16 models passed this data analysis prompt's check (final answer); Claude Sonnet 5, Claude Haiku 4.5, Gemini 3.8 Flash, and GLM 5.3 Flash didn't. The cheapest reply that passed was GPT-6 Luna's, at $0.00037; the fastest, Claude Sonnet 5.5's in 4.2 s. The dearest reply, Claude Fable 5.1's, cost 179 times as much ($0.0663).

What does a reply to “Correlation between ad spend and sign-ups” cost?

Through the models' APIs, what OpenRouter charged us ran from $0.00037 (GPT-6 Luna) to $0.0663 (Claude Fable 5.1) for this prompt. In llmwise you don't pay by the token: a reply like these counts as one message on Pro, whichever model answers.

Claude, GPT, Gemini, DeepSeek, Grok, Kimi, and GLM, in one chat.

See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.