GPT · Summarization
GPT-6 for summarization
llmwise has 4 of OpenAI's GPT models, from GPT-6 Luna (up to 60 messages a day on Pro) to GPT-6 Astra (up to 31 messages a month). We ran the same summarization prompts on every one and published every reply: which GPT model to use, from the results, what each costs per message, and how to get more out of it.
Based on 20 of our test runs on , through OpenRouter with the app's own prompt and settings. Updated .
Short answer
In our test runs on September 29, 2026, 2 models passed 5 of 5 summarization prompts, so these prompts don't pick one for hard problems. For value, GPT-6.1 Sol (5 of 5), 125 a month on Pro; for everyday summarization, GPT-6 Luna (4 of 5), from the daily count.
Our picks for summarization
Hard problems
Shared by 2 models
2 models passed 5 of 5, both hard ones: GPT-6 Astra and GPT-6.1 Sol. These prompts don't tell them apart, so they share the pick.
Best value
Passed 5 of 5 summarization prompts, with 125 a month on Pro.
Everyday
Passed 4 of 5 summarization prompts; an everyday model, so its messages come from the daily count (60 a day on Pro), not the monthly allowance.
These picks aren't our opinion: they're what the results below give, by these rules, among the GPT models in llmwise. They change when the results do.
- Hard problems: the model that passed the most prompts and, of those, the most hard ones. Models level on both share the pick: the prompts don't tell them apart, so we don't break the tie by price or by name.
- Best value: among the models that draw on the monthly allowance, the one with the most messages on Pro that passed no more than one prompt fewer than the top model. Ties go to the one that passed more, then to the lower cost per reply.
- Everyday: among the cheapest models on the page (the everyday models, which come from the daily count, when the page has any), the one that passed the most. Ties go to the one that passed more of the hard prompts, then to the lower cost per reply.
Our summarization test runs, model by model
How each model did on our 5 summarization prompts, what each reply counted as on Pro, and what it cost to run.
| Model | Passed | Hard ones | Messages used on Pro | Cost per reply | Time per reply |
|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 5 of 5 | 2 of 2 | 1 each, of 31 a month on Pro | $0.0109 | 3.0 s |
| GPT-6.1 SolOpenAI | 5 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0010 | 2.0 s |
| GPT-6 SolOpenAI | 4 of 5 | 2 of 2 | 1 each, of 125 a month on Pro | $0.0022 | 2.4 s |
| GPT-6 LunaOpenAI | 4 of 5 | 2 of 2 | 1 each, of 60 a day on Pro | $0.0001 | 1.4 s |
One prompt, side by side
Each model's reply to “A quarterly memo for the CEO”, as it came, with whether it passed and what it cost.
The prompt: A quarterly memo for the CEO (hard)
Summarize this memo for the CEO in at most 4 sentences, keeping every number exact. Q3 at Fernway: revenue was $2.4M, up 18% on Q2, driven by the new team plan, which now makes up 41% of new sign-ups. Churn fell from 3.1% to 2.6% a month. Gross margin slipped from 71% to 68% because of higher hosting costs, which the platform team expects to bring back down by the end of Q4 after moving storage providers. We ended the quarter with 14 months of runway. Hiring is frozen until Q2 next year, except for two support roles already approved. The biggest risk is our largest customer, 9% of revenue, whose contract is up for renewal in November.
GPT-6 Astra
Passed: Graded 4.7 of 5 on average (lowest 4).
Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan, which accounts for 41% of new sign-ups; monthly churn fell from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to higher hosting costs, which the platform team expects to reduce by the end of Q4 after switching storage providers. Quarter-end runway was 14 months, and hiring is frozen until Q2 next year except for two approved support roles. The biggest risk is the November renewal of our largest customer, which represents 9% of revenue.
507 tokens in, 134 out · 2.9 s · $0.0118 · 1 message on Pro · answered by openai/gpt-6-astra via OpenAI ·
GPT-6.1 Sol
Passed: Graded 4.7 of 5 on average (lowest 4).
Fernway’s Q3 revenue was $2.4M, up 18% on Q2, driven by the new team plan, now 41% of new sign-ups; monthly churn fell from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to higher hosting costs, which the platform team expects to reduce by the end of Q4 after switching storage providers. The company ended Q3 with 14 months of runway, and hiring is frozen until Q2 next year except for two approved support roles. The biggest risk is the November renewal of the largest customer, which accounts for 9% of revenue.
509 tokens in, 140 out · 2.4 s · $0.0012 · 1 message on Pro · answered by openai/gpt-6.1-sol via OpenAI ·
GPT-6 Luna
Passed: Graded 4.0 of 5 on average (lowest 4).
Q3 revenue was $2.4M, up 18% on Q2, with the new team plan driving growth and accounting for 41% of new sign-ups; monthly churn fell from 3.1% to 2.6%. Gross margin declined from 71% to 68% due to higher hosting costs, which the platform team expects to reduce by the end of Q4 after a storage-provider move. Fernway ended the quarter with 14 months of runway, and hiring is frozen until Q2 next year except for two approved support roles. The main risk is the largest customer, representing 9% of revenue, whose contract renews in November.
507 tokens in, 141 out · 1.5 s · $0.0001 · 1 message on Pro · answered by openai/gpt-6-luna via OpenAI ·
The summarization prompts, and how they're scored
Rubric (graded). The grader model scores the reply from 1 to 5 on each published criterion. It passes with an average of 4 or more and no criterion under 3, and only if it also meets the prompt's automatic rules (length, words it must or mustn't use).
The grader is Claude Opus 5.5 at low reasoning effort; its own replies are graded by GPT-6 Astra, so no model grades itself. Its prompt and every rubric are on the methods page.
Each prompt was sent the way llmwise sends a message in a side-by-side comparison, which offers no tools: the app's own system prompt, the model's own settings, and Pro's reply size limit (8,000 tokens). Read all five summarization prompts and how every reply was scored.
Prompts like these to try yourself
GPT in llmwise
| Model | On Pro | On Free | Context window | Images | PDFs | Reasoning | API price per 1M, in / out |
|---|---|---|---|---|---|---|---|
| GPT-6 AstraOpenAI | 31/mo on Pro | No | 1.05M tokens | Yes | Whole file | Yes | $10.00 / $50.00 |
| GPT-6.1 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $2.00 / $10.00 |
| GPT-6 SolOpenAI | 125/mo on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $2.00 / $10.00 |
| GPT-6 LunaOpenAI | 60/day on Pro | Yes | 1.05M tokens | Yes | Whole file | Yes | $0.10 / $0.50 |
Every limit is published. Paid plans also have a monthly fair-use limit on AI cost: Pro $7.50, Max $20, Ultra $42, Studio $85. Using every message on your plan at typical sizes stays under it; very large messages and heavy research use it faster. Every limit, explained.
What matters for summarization
How much it can read
A model can only summarize what fits in its context window, listed for each model below.
Staying faithful
A summary is only useful if it's faithful. Ask for key numbers and quotes so you can check them.
The right length
Say who the summary is for and how long it should be; a one-line TL;DR and a one-page brief are different jobs.
GPT for summarization in llmwise
Attach your documents
Attach PDFs (up to 100 pages), images, Word, Excel, and PowerPoint files, and text or code files (plain text, Markdown, and CSV, JSON and more), up to 10 files a message.
Summarize a link
Paste a link and the model reads the page. It counts as one web search.
Keep the summary
Ask for the summary as a document to edit it, keep its versions and download it as PDF, Word or Markdown.
Long chats, stated up front
Past 64k tokens each reply counts as 2, past 128k as 4, and a chat stops at 200k; the composer says so before you send. One document per chat keeps it short.
In llmwise, a message to GPT goes to its maker, OpenAI, or through OpenRouter when llmwise can't reach the maker directly. See the Privacy Policy.
Tips
Say who the summary is for and how long it should be.
Ask for the key numbers and quotes, with page references.
Ask what the summary leaves out.
Use one chat per long document.
Questions
Is GPT good for summarization?
In our test runs on September 29, 2026, GPT-6 Astra passed 5 of 5, GPT-6.1 Sol passed 5 of 5, GPT-6 Sol passed 4 of 5, GPT-6 Luna passed 4 of 5 of our summarization prompts, against a best result of 5 of 5 among all 19 models. Every reply is published on this page and the methods page, so you can judge them yourself.
Which GPT model should I use for summarization?
Start with GPT-6 Luna for emails, short articles and quick TL;DRs, and move up to GPT-6 Astra for dense or high-stakes material: contracts, research papers, anything you'll act on.
Can I use GPT for summarization for free?
Yes: GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Luna are on the Free plan (a one-time trial of 5 messages). GPT-6 Astra needs a paid plan.
How long a document can I summarize?
PDFs up to 100 pages and files up to 20 MB, within the model's context window (listed below).
Claude, GPT, Gemini, DeepSeek, Grok, Kimi, GLM, and Mistral, in one chat.
See what a message costs before you send it. Free is 5 messages to try; sign in with an email link, no password or card.