Method and data · measured 2026-10-10 to 2026-10-11
How we measured
Everything behind the calculator and the report: the exact prompts, where the data came from, how the emails were judged, where every price comes from, what is assumed, what this test cannot tell you, and the raw files to check or rerun it yourself.
Models
Run 2026-10-11 · official APIs
Both tested models ran through their vendors' official APIs at medium effort with identical prompts. The judges ran at medium effort. Request overhead is the number of tokens each API adds for message structure, measured with the token-counting endpoints.
| Model | API id | Role | Effort | Max input tokens | Request overhead |
|---|---|---|---|---|---|
| Claude Haiku 5.5 | claude-haiku-5-5 | tested | medium | 1,000,000 | 9 tokens |
| GPT-6 Luna | gpt-6-luna | tested | medium | 922,000 | 6 tokens |
| Claude Opus 5.5 | claude-opus-5-5 | email judge | medium | 1,000,000 | not measured |
| GPT-6.1 Sol | gpt-6.1-sol | email judge | medium | 922,000 | not measured |
Prompts
These are the exact strings the scripts send, exported from the scripts themselves. None of the prompts tells the model today's date, except the date-fix sentence used in the rerun.
System prompt for ticket classification classify_system · 829 characters · from eval/run_classify.py
You classify customer support tickets for a B2B SaaS product.
Reply with exactly one label and nothing else: billing, bug, feature_request, account, cancellation.
- billing: invoices, prices, payment methods, refunds (not stopping renewal or cancelling)
- bug: something in the product is broken, errors, wrong display
- feature_request: asking for a new or improved capability
- account: login, password, 2FA, email change, adding/removing members or permissions (including deleting one member's account)
- cancellation: cancel the contract, stop auto-renewal, don't convert trial, leave the service / close the whole organization account
Tie-breaks: asking for a capability that doesn't exist yet is feature_request even if it is about login (e.g. SSO).
A user who cannot log in is account even if a reset email doesn't arrive.Sentence appended to the email system prompt in the date-fix rerun date_suffix · 20 characters · from eval/check_date_fix.py
今日は2026年10月10日(土)です。System prompt for Japanese business email email_system · 92 characters · from eval/run_tasks.py
あなたは日本企業で働く人のメール作成を手伝うアシスタントです。依頼内容をもとに、そのまま送れるビジネスメールを書いてください。件名と本文だけを出力し、説明や前置きは書かないでください。JSON schema sent with the extraction request (structured output) extract_schema · 553 characters · from eval/run_tasks.py
{
"type": "object",
"properties": {
"company": {
"type": "string"
},
"invoice_no": {
"type": "string"
},
"issue_date": {
"type": "string"
},
"due_date": {
"type": "string"
},
"amount": {
"type": "integer"
},
"currency": {
"type": "string",
"enum": [
"JPY",
"USD",
"EUR"
]
}
},
"required": [
"company",
"invoice_no",
"issue_date",
"due_date",
"amount",
"currency"
],
"additionalProperties": false
}System prompt for invoice extraction extract_system · 551 characters · from eval/run_tasks.py
Extract the invoice that the reader is being asked to pay from the message. Return JSON with:
- company: the issuing (billing) company, exactly as written, including its legal suffix
- invoice_no: the invoice number exactly as written
- issue_date, due_date: ISO 8601 (YYYY-MM-DD); convert Japanese era years and terms like "net 30"
- amount: total payable including tax, as an integer in whole currency units; if only a subtotal and tax are given, add them; if the message corrects an earlier amount, use the corrected one
- currency: JPY, USD or EURSystem prompt for both email judges judge_system · 1,248 characters · from eval/judge_emails.py
You are an experienced reviewer of Japanese business email at a Japanese company.
You will see a request an office worker typed to an AI assistant, and two candidate emails (A and B) written from it.
Judge each email on whether it could be sent as-is. Score 1-5 (5 best):
- accuracy: every fact in the request (people, companies, dates, amounts, quantities, conditions, sender) is present and correct
- keigo: honorifics and business manners fit the relationship (external partner, first contact, customer, boss, colleague)
- natural: reads like a native Japanese business writer; no translationese or awkward phrasing
- concise: length fits the situation; no padding or content the request didn't need
- sendable: 5 = send with no edits, 3 = needs a few edits, 1 = needs a rewrite
List concrete problems only:
- missing_or_wrong: facts from the request that are missing or stated incorrectly
- invented: specifics or commitments not in the request that could cause trouble if sent (new dates, amounts, promises, policies)
winner: the email a careful Japanese manager would rather send ("A", "B" or "tie"). Do not prefer an email for being longer or shorter.
Write reason and the problem lists in Simplified Chinese; reason is one or two sentences.User message template for the judges ({brief}, {email_a}, {email_b} are filled in) judge_user_template · 55 characters · from eval/judge_emails.py
## 依頼内容
{brief}
## メール A
{email_a}
## メール B
{email_b}Datasets
Checked 2026-10-11
All task data was written for this test by AI agents: 50 support tickets (25 English, 25 Japanese), 25 English and 25 Japanese invoice messages with gold field values, and 25 short Japanese email requests. No real customer data was used. The English, Japanese and Chinese tokenizer samples were also written by AI agents; four email addresses in them that looked like real domains were changed to .example. The code and JSON samples are excerpts of real open-source files (see corpus licenses).
A script checks every dataset mechanically before the site is built. These are the only data-quality claims we make. For the invoices, each gold field records the source text it was taken from; the script checks that this text appears verbatim in the document and that the gold values are well-formed. It does not check that a normalized gold value (an amount as a plain integer, a date in ISO format) means the same as its source text; that was set when the data was written.
| File | Items | Check | Passed |
|---|---|---|---|
| eval/classify.jsonl | 50 | label is one of the allowed labels | 50 of 50 |
| ids are unique | 50 of 50 | ||
| ticket text is not empty | 50 of 50 | ||
| eval/email_ja.jsonl | 25 | ids are unique | 25 of 25 |
| request text is not empty | 25 of 25 | ||
| requests used in the report are present | 3 of 3 | ||
| eval/extract_en.jsonl | 25 | the source text recorded for each gold field appears verbatim in its document | 155 of 155 |
| gold values are complete and well-formed: valid dates, due date not before issue date, whole-number amount, known currency | 25 of 25 | ||
| ids are unique | 25 of 25 | ||
| eval/extract_ja.jsonl | 25 | the source text recorded for each gold field appears verbatim in its document | 153 of 153 |
| gold values are complete and well-formed: valid dates, due date not before issue date, whole-number amount, known currency | 25 of 25 | ||
| ids are unique | 25 of 25 |
How date errors are counted
eval/audit_dates.py reads every email and checks each date written with a weekday or a year against the 2026 and 2027 calendars. It recognizes four ways of writing a date (after normalizing full-width digits and brackets): 「10月16日(金)」 or 「2026年10月16日(金)」, 「2026年10月16日」, 「金曜日(10月16日)」, and 「10/16(金)」. A year is only checked when the full year, month and day are written, and it must be 2026 or 2027 (「2025年度版」 is not a date). A weekday without a year counts as correct if it matches either year. Claude Haiku 5.5: 7 of 25 emails with at least one error; GPT-6 Luna: 0 of 25 emails with at least one error. The same script checked the 14 outputs of the date-fix rerun.
How placeholders are counted
An email counts once if it contains any of these:
- an empty fill-in bracket: a label and colon followed only by spaces, such as 「(連絡先: )」, an empty bracket, or a bracket that tells the writer to fill something in (「ここに」「記入」「入力」「挿入」), such as 「(URLをここに記載)」;
- placeholder characters: runs of 〇 or ○, or of X or x, including inside names, phone numbers and email addresses;
- a 【…】 or [...] item that mentions a URL or asks for input, or that stands alone on a line with nothing under it.
A heading such as 【送付資料】 followed by real content does not count. The rules and real examples are part of the export script's self-test.
Labels and gold values
The tickets and their labels were written together, 10 per label across 5 labels (billing, bug, feature_request, account, cancellation). The invoice gold values were written with each document, including 8 Japanese and 8 English hard cases, such as totals given only as subtotal plus tax and amounts corrected later in the same thread. We did not run an independent human relabeling; the checks above are mechanical. Both models matched every label and every gold value, so these two tasks do not separate them on accuracy.
Judging
Judged 2026-10-10
Each email pair was judged by Claude Opus 5.5 (Anthropic) and GPT-6.1 Sol (OpenAI). A judge saw the request and the two emails as A and B, without model names, and judged every pair twice with A and B swapped. For each email it scored five criteria from 1 to 5 (accuracy, keigo, naturalness, concision, ready to send as-is), listed specifics or commitments not in the request and facts that were missing or wrong, and picked A, B or a tie.
The two verdicts for a pair are combined by adding them: a win counts +1 for that model and a tie 0. A positive total is a win, a negative total a loss, and zero a tie, so one win plus one tie still counts as a win.
Before judging, both emails go through the same cleanup (review/build.py clean()): trailing spaces are removed from every line, leading and trailing blank lines are dropped, and three or more line breaks are collapsed to one blank line. This keeps formatting habits from giving away which model wrote which email; the text itself is not changed.
| Judge | Luna / tie / Haiku | Same verdict in both orders |
|---|---|---|
| Claude Opus 5.5 | 22 / 2 / 1 | 23 of 25 |
| GPT-6.1 Sol | 25 / 0 / 0 | 25 of 25 |
The two judges agreed on 22 of 25 pairs. Judges are AI models, not native-speaker reviewers.
Prices and sources
Every price, tier rule, batch discount and context limit comes from a vendor page, opened and checked on the date shown. The export fails if any price field is neither covered by a source nor listed as an assumption. All costs on the site are recomputed from the billed tokens at these prices.
| Model | Base: in / out / cache read | Higher tier | Batch | Max input |
|---|---|---|---|---|
| Claude Haiku 5.5 | $0.10 / $0.50 / $0.01 | above 100K: $0.50 / $2.50 / $0.05, whole request | × 0.5 | 1,000,000 |
| GPT-6 Luna | $0.10 / $0.50 / $0.01 | above 272K: $0.20 / $0.75 / $0.02, whole request | × 0.5 | 922,000 |
| Claude Opus 5.5 | $4.00 / $20.00 / $0.20 | none | × 0.5 | 1,000,000 |
| GPT-6.1 Sol | $2.00 / $10.00 / $0.10 | none | × 0.5 | 922,000 |
- Claude Haiku 5.5: platform.claude.com/docs/en/about-claude/pricing, checked 2026-10-11. Covers base.in, base.out, base.cache_read, tier.threshold_tokens, tier.in, tier.out, tier.cache_read, tier.applies_to, batch_multiplier.
Claude Haiku 5.5 (for prompts up to 100,000 tokens) $0.10 / $0.01 (cache hits) / $0.50; (for prompts over 100,000 tokens) $0.50 / $0.05 / $2.50; batch $0.05/$0.25 and $0.25/$1.25. "a request whose prompt is over 100,000 tokens pays higher prices"
- Claude Haiku 5.5: platform.claude.com/docs/en/about-claude/models/overview, checked 2026-10-11. Covers max_input_tokens.
Context window: 1M tokens (Models API max_input_tokens = 1000000 on 2026-10-11)
- GPT-6 Luna: developers.openai.com/api/docs/pricing, checked 2026-10-11. Covers base.in, base.out, base.cache_read, tier.in, tier.out, tier.cache_read, batch_multiplier, batch_multiplier.cache_read.
gpt-6-luna | $0.10 | $0.01 | $0.125 | $0.50 | $0.20 | $0.02 | $0.25 | $0.75 (short / long context); Batch and Flex $0.05 | $0.005 | ... | $0.25 | $0.10 | $0.01 | ... | $0.375
- GPT-6 Luna: developers.openai.com/api/docs/models/gpt-6-luna, checked 2026-10-11. Covers max_input_tokens, tier.threshold_tokens, tier.applies_to.
Maximum input tokens: 922,000. "Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request."
- Claude Opus 5.5: platform.claude.com/docs/en/about-claude/pricing, checked 2026-10-11. Covers base.in, base.out, base.cache_read, batch_multiplier.
Claude Opus 5.5 | $4 / MTok | ... | $0.20 / MTok | $20 / MTok; batch $2 / $10
- Claude Opus 5.5: platform.claude.com/docs/en/about-claude/models/overview, checked 2026-10-11. Covers max_input_tokens.
Context window: 1M tokens (Models API max_input_tokens = 1000000 on 2026-10-11)
- GPT-6.1 Sol: developers.openai.com/api/docs/pricing, checked 2026-10-11. Covers base.in, base.out, base.cache_read, batch_multiplier.
gpt-6.1-sol | $2.00 | $0.10 | $2.50 | $10.00 | ...; Batch $1.00 | $0.05 | ... | $5.00
- GPT-6.1 Sol: developers.openai.com/api/docs/models/gpt-6.1-sol, checked 2026-10-11. Covers max_input_tokens.
Maximum input tokens: 922,000
Assumptions
Values we could not confirm from a source are labeled assumption wherever they appear.
- assumption Claude Haiku 5.5, batch_multiplier.cache_read: The batch table lists only input and output prices; we assume the 50% batch discount also applies to cache reads ("Batch API requests are 50% off").
- assumption Calculator scenario “Ticket classification (Japanese)”: Inputs come from the measured run, but the calculator estimates prompt tokens from character counts by language. This task's short, keyword-dense English system prompt makes that estimate run more than 15% above Haiku's measured cost, so we label the scenario an assumption. The measured table shows the real bill.
- assumption Calculator scenario “Code review”: Output length and reasoning tokens are assumptions, not measurements. The reasoning numbers reuse the email task's averages.
- assumption Calculator scenario “English long document → Japanese summary”: Output length and reasoning tokens are assumptions, not measurements. The reasoning numbers reuse the email task's averages.
- assumption Haiku's reasoning tokens: the Anthropic API bills reasoning and visible output together, so we estimate Haiku's reasoning tokens as billed output tokens minus the token count of the visible answer. The measured task costs use the API's billed totals, not this estimate; the calculator's scenarios do use it as Haiku's reasoning tokens. Luna reports its reasoning tokens directly.
- assumption Characters to tokens: the calculator converts characters to tokens with the average ratio for each content type (table on the calculator page). Real text varies; the Japanese genres alone range from ×1.15 to ×1.34 (3 samples each).
Limits
- Small samples: 50 tickets, 50 invoices and 25 email requests per model, 1 run each. Results can shift with another run.
- All task data was written by AI agents, not taken from real workloads. Real tickets, invoices and requests are messier.
- The email preference comes from two AI judges, not from native Japanese speakers.
- Haiku's reasoning tokens are estimated: billed output tokens minus the counted visible answer. Luna reports its reasoning tokens directly.
- Latency was measured from one client location at one time of day.
- Prices and model behavior change. Every number carries the date it was measured or checked.
Cost of this test
About $1.30 in API fees, including judging and superseded runs. Sum of every billed call saved in eval/results, at current list prices. A few ad-hoc calls made during exploration were not saved and are not included. Token-counting calls are free and count as $0.
Cost by result file
| File | USD | Note |
|---|---|---|
| eval/results/classify-haiku-20261009-200854.jsonl | $0 | |
| eval/results/classify-haiku-20261009-233558.jsonl | $0.00016 | |
| eval/results/classify-haiku-20261009-233602.jsonl | $0.0016 | |
| eval/results/classify-haiku-20261011-181146.jsonl | $0.0016 | |
| eval/results/classify-luna-20261009-200854.jsonl | $0.00015 | |
| eval/results/classify-luna-20261009-233602.jsonl | $0.0017 | |
| eval/results/classify-luna-20261011-181146.jsonl | $0.0017 | |
| eval/results/classify-summary-20261009-200854.json | $0 | summary file; its calls are counted in the row files |
| eval/results/classify-summary-20261009-233558.json | $0 | summary file; its calls are counted in the row files |
| eval/results/classify-summary-20261009-233602.json | $0 | summary file; its calls are counted in the row files |
| eval/results/classify-summary-20261011-181146.json | $0 | summary file; its calls are counted in the row files |
| eval/results/coeff-20261010-011346.json | $0 | token counting or local check; not billed |
| eval/results/coeff-20261011-181620.json | $0 | token counting or local check; not billed |
| eval/results/coeff-20261011-182428.json | $0 | token counting or local check; not billed |
| eval/results/coeff-rows-20261010-011346.jsonl | $0 | token counting or local check; not billed |
| eval/results/coeff-rows-20261011-181620.jsonl | $0 | token counting or local check; not billed |
| eval/results/coeff-rows-20261011-182428.jsonl | $0 | token counting or local check; not billed |
| eval/results/dataset-check-20261011-181205.json | $0 | token counting or local check; not billed |
| eval/results/date-audit-20261011-181215.json | $0 | token counting or local check; not billed |
| eval/results/date-audit-20261011-181315.json | $0 | token counting or local check; not billed |
| eval/results/date-fix-20261011-181315.json | $0.0039 | |
| eval/results/email-haiku-20261010-020535.jsonl | $0.00068 | |
| eval/results/email-haiku-20261010-020957.jsonl | $0.0059 | |
| eval/results/email-luna-20261010-020535.jsonl | $0.00051 | |
| eval/results/email-luna-20261010-020957.jsonl | $0.004 | |
| eval/results/extract-haiku-20261010-020957.jsonl | $0.0073 | |
| eval/results/extract-luna-20261010-020957.jsonl | $0.0045 | |
| eval/results/judge-20261010-040157.jsonl | $0.097 | |
| eval/results/judge-20261010-040432.jsonl | $1.17 | |
| eval/results/judge-summary-20261010-040157.json | $0 | summary file; its calls are counted in the row files |
| eval/results/judge-summary-20261010-040432.json | $0 | summary file; its calls are counted in the row files |
| eval/results/schema-overhead-20261011-181144.json | $0 | token counting or local check; not billed |
| eval/results/tasks-summary-20261010-020535.json | $0 | summary file; its calls are counted in the row files |
| eval/results/tasks-summary-20261010-020957.json | $0 | summary file; its calls are counted in the row files |
| eval/results/tokens-20261009-234710.json | $0 | token counting or local check; not billed |
Corpus sources and licenses
The tokenizer corpus has 24 English, 24 Japanese, 24 Chinese, 24 Python, 24 JavaScript/TypeScript, 15 JSON samples. en/ja/zh samples are original texts written for this test by AI agents (eval/prose.json). Python samples are excerpts of the Python 3.12.13 standard library; JavaScript/TypeScript and JSON samples are excerpts of npm packages, taken only from packages under MIT, ISC, BSD-2-Clause, BSD-3-Clause, Apache-2.0, CC0-1.0. Copyright lines in the samples are kept exactly as written, and each package's license text ships in the download under LICENSES/.
Replaced on 2026-10-11 (code_js_ts, json, prose emails): Earlier JS/TS and JSON samples came from bundled front-end files inside wheel packages in the uv cache, whose licenses could not be verified one by one; they were re-sampled from site/node_modules with each license recorded. Four email addresses in the prose samples that looked like real domains were changed to .example. Standard-library samples were re-drawn from a pinned Python version.
Every code and JSON sample
| Sample | Source file | Package | License | License text |
|---|---|---|---|---|
| code_js_ts:01 | @astrojs/compiler-binding/index.js | @astrojs/[email protected] | MIT | LICENSES/@astrojs/[email protected]/LICENSE |
| code_js_ts:02 | aria-query/lib/rolesMap.js | [email protected] | Apache-2.0 | LICENSES/[email protected]/LICENSE |
| code_js_ts:03 | css-tree/lib/lexer/error.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:04 | csso/lib/clean/Rule.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:05 | domutils/lib/esm/querying.js | [email protected] | BSD-2-Clause | LICENSES/[email protected]/LICENSE |
| code_js_ts:06 | postcss/lib/map-generator.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:07 | prismjs/components/prism-dax.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:08 | prismjs/components/prism-mermaid.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:09 | prismjs/components/prism-smarty.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:10 | prismjs/plugins/match-braces/prism-match-braces.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:11 | source-map-js/lib/util.js | [email protected] | BSD-3-Clause | LICENSES/[email protected]/LICENSE |
| code_js_ts:12 | svgo/plugins/moveElemsAttrsToGroup.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:13 | undici/lib/core/socks5-client.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:14 | undici/lib/interceptor/decompress.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:15 | undici/lib/web/websocket/events.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:16 | zod/src/v3/tests/map.test.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:17 | zod/src/v4/classic/tests/cyclic-data.test.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:18 | zod/src/v4/classic/tests/recursive-types.test.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:19 | zod/src/v4/core/tests/locales/el.test.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:20 | zod/src/v4/locales/de.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:21 | zod/src/v4/locales/mk.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:22 | zod/src/v4/locales/uz.ts | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:23 | zod/v4/locales/bn.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_js_ts:24 | zod/v4/locales/ka.js | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| code_python:01 | __future__.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:02 | _markupbase.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:03 | _pyio.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:04 | _threading_local.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:05 | base64.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:06 | cgi.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:07 | codeop.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:08 | copy.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:09 | decimal.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:10 | fileinput.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:11 | genericpath.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:12 | gzip.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:13 | ipaddress.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:14 | mailbox.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:15 | nntplib.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:16 | operator.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:17 | pickletools.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:18 | profile.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:19 | queue.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:20 | runpy.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:21 | shlex.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:22 | socket.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:23 | string.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| code_python:24 | sysconfig.py | [email protected] | PSF-2.0 | LICENSES/[email protected]/LICENSE.txt |
| json:01 | ci-info/vendors.json | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| json:02 | css-tree/data/patch.json | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| json:03 | csso/node_modules/mdn-data/api/inheritance.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:04 | csso/node_modules/mdn-data/css/at-rules.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:05 | csso/node_modules/mdn-data/css/selectors.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:06 | csso/node_modules/mdn-data/css/syntaxes.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:07 | csso/node_modules/mdn-data/css/types.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:08 | csso/node_modules/mdn-data/css/units.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:09 | csso/node_modules/mdn-data/l10n/css.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:10 | extend/.jscs.json | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| json:11 | mdn-data/css/definitions.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:12 | mdn-data/css/functions.json | [email protected] | CC0-1.0 | LICENSES/[email protected]/LICENSE |
| json:13 | prismjs/components.json | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| json:14 | undici/docs/docs/site.json | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
| json:15 | undici/docs/docs/type-map.json | [email protected] | MIT | LICENSES/[email protected]/LICENSE |
Downloads
The download mirrors the repository layout, so the scripts run from it as-is. It contains the datasets, every raw model output and judge output used on the site, the summaries, the scripts, this project's MIT license, and the license texts of the corpus samples. It contains no API keys or personal data; the email addresses in it are fictional (.example), placeholders written by the models, or copyright lines that a license requires us to keep.
Download everything (llm-bill-data-20261011.zip) manifest.json (paths, sizes, SHA-256)
Every file in the download
| File | What it is | Size |
|---|---|---|
| LICENSE | MIT license for this project's code | 1.3 kB |
| LICENSES/@astrojs/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 10.2 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.3 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 6.6 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 6.6 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE.txt | Third-party license text | 13.9 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.5 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| LICENSES/[email protected]/LICENSE | Third-party license text | 1.1 kB |
| corpus_sources.json | Corpus sources and licenses | 50.4 kB |
| eval/audit_dates.py | Evaluation script | 5.4 kB |
| eval/build_corpus.py | Evaluation script | 12.6 kB |
| eval/check_date_fix.py | Evaluation script | 2.7 kB |
| eval/classify.jsonl | Classification dataset | 7.0 kB |
| eval/corpus.jsonl | Tokenizer corpus | 347.3 kB |
| eval/corpus_licenses.json | Tokenizer corpus | 25.1 kB |
| eval/count_tokens.py | Evaluation script | 6.5 kB |
| eval/email_allowlist.json | Allowed email addresses | 2.7 kB |
| eval/email_ja.jsonl | Email briefs | 11.2 kB |
| eval/excerpts.json | Excerpt selection rules | 5.3 kB |
| eval/export_site_data.py | Evaluation script | 48.5 kB |
| eval/extract_en.jsonl | Extraction dataset | 30.7 kB |
| eval/extract_ja.jsonl | Extraction dataset | 37.7 kB |
| eval/judge_emails.py | Evaluation script | 9.1 kB |
| eval/measure_coeff.py | Evaluation script | 4.1 kB |
| eval/measure_schema_overhead.py | Evaluation script | 4.0 kB |
| eval/presets.json | Calculator presets | 2.7 kB |
| eval/prices.json | Prices with sources | 4.6 kB |
| eval/prose.json | Original prose samples (AI-written) | 130.2 kB |
| eval/results/classify-haiku-20261011-181146.jsonl | Ticket classification raw outputs / summary | 14.0 kB |
| eval/results/classify-luna-20261011-181146.jsonl | Ticket classification raw outputs / summary | 14.3 kB |
| eval/results/classify-summary-20261011-181146.json | Ticket classification raw outputs / summary | 1.1 kB |
| eval/results/coeff-20261011-182428.json | Tokenizer coefficients | 8.3 kB |
| eval/results/coeff-rows-20261011-182428.jsonl | Tokenizer coefficients | 14.4 kB |
| eval/results/dataset-check-20261011-181205.json | Dataset checks | 2.4 kB |
| eval/results/date-audit-20261011-181315.json | Date audit | 1.9 kB |
| eval/results/date-fix-20261011-181315.json | Date-fix rerun outputs | 18.5 kB |
| eval/results/email-haiku-20261010-020957.jsonl | Japanese email raw outputs | 36.7 kB |
| eval/results/email-luna-20261010-020957.jsonl | Japanese email raw outputs | 26.3 kB |
| eval/results/extract-haiku-20261010-020957.jsonl | Invoice extraction raw outputs | 36.2 kB |
| eval/results/extract-luna-20261010-020957.jsonl | Invoice extraction raw outputs | 36.1 kB |
| eval/results/judge-20261010-040432.jsonl | Blind judging outputs / summary | 85.6 kB |
| eval/results/judge-summary-20261010-040432.json | Blind judging outputs / summary | 2.4 kB |
| eval/results/schema-overhead-20261011-181144.json | Structured-output overhead | 0.6 kB |
| eval/results/tasks-summary-20261010-020957.json | Extraction and email summary | 2.1 kB |
| eval/run_classify.py | Evaluation script | 8.9 kB |
| eval/run_tasks.py | Evaluation script | 9.8 kB |
| eval/validate_datasets.py | Evaluation script | 5.6 kB |
| review/build.py | Shared cleaning helper (used by judge_emails.py) | 1.7 kB |
Rerun it
You need uv (it installs the pinned Python and dependencies from each script's header) and, for the billed steps, your own API keys. Unzip the download and run from its top folder.
# 1. Self-tests and dataset checks: free, no API calls for s in eval/*.py; do uv run -q "$s" --selftest; done uv run -q eval/validate_datasets.py # 2. API keys for the next steps. Paste each key at the prompt: the input is # hidden and never typed on a command line, so it stays out of shell history. printf 'Anthropic API key: '; read -rs ANTHROPIC_API_KEY; echo; export ANTHROPIC_API_KEY printf 'OpenAI API key: '; read -rs OPENAI_API_KEY; echo; export OPENAI_API_KEY # Only if your Anthropic key needs a workspace (an id, not a secret): # export ANTHROPIC_WORKSPACE_ID=<workspace id> # 3. Tokenizer ratios and schema overhead: token-counting endpoints, free uv run -q eval/measure_coeff.py uv run -q eval/measure_schema_overhead.py # 4. The three tasks, then the judges: billed uv run -q eval/run_classify.py uv run -q eval/run_tasks.py uv run -q eval/judge_emails.py # 5. Date check and the date-fix rerun uv run -q eval/audit_dates.py uv run -q eval/check_date_fix.py uv run -q eval/audit_dates.py
Each script writes a new timestamped file to eval/results/ and never overwrites earlier runs. Our complete measurement, including superseded runs, cost $1.30; see the cost table above for the per-file split.
How the numbers are checked
Pages never contain hand-typed figures. Measured numbers come from the data files above; calculator numbers (default bills, ratios, tier points) are computed when the site is built by the same code the calculator runs in your browser.
After every build, a check reads the visible text of each page. Each calculator number is marked and recomputed from the data and must match exactly. Every other number must appear in the data files (allowing common formatting such as rounding, percentages, K/M and thousands separators) or in a short list of model version numbers, years and method constants. Any mismatch fails the build.
Limits of the check: single-digit numbers almost always match something in the data, so the check says little about them; numbers written as words or kanji numerals are not checked; dates are skipped.