llm-bill

Method and data · measured 2026-10-10 to 2026-10-11

How we measured

Everything behind the calculator and the report: the exact prompts, where the data came from, how the emails were judged, where every price comes from, what is assumed, what this test cannot tell you, and the raw files to check or rerun it yourself.

Models

Run 2026-10-11 · official APIs

Both tested models ran through their vendors' official APIs at medium effort with identical prompts. The judges ran at medium effort. Request overhead is the number of tokens each API adds for message structure, measured with the token-counting endpoints.

ModelAPI idRoleEffortMax input tokensRequest overhead
Claude Haiku 5.5claude-haiku-5-5testedmedium1,000,0009 tokens
GPT-6 Lunagpt-6-lunatestedmedium922,0006 tokens
Claude Opus 5.5claude-opus-5-5email judgemedium1,000,000not measured
GPT-6.1 Solgpt-6.1-solemail judgemedium922,000not measured

Prompts

These are the exact strings the scripts send, exported from the scripts themselves. None of the prompts tells the model today's date, except the date-fix sentence used in the rerun.

System prompt for ticket classification classify_system · 829 characters · from eval/run_classify.py
You classify customer support tickets for a B2B SaaS product.
Reply with exactly one label and nothing else: billing, bug, feature_request, account, cancellation.
- billing: invoices, prices, payment methods, refunds (not stopping renewal or cancelling)
- bug: something in the product is broken, errors, wrong display
- feature_request: asking for a new or improved capability
- account: login, password, 2FA, email change, adding/removing members or permissions (including deleting one member's account)
- cancellation: cancel the contract, stop auto-renewal, don't convert trial, leave the service / close the whole organization account
Tie-breaks: asking for a capability that doesn't exist yet is feature_request even if it is about login (e.g. SSO).
A user who cannot log in is account even if a reset email doesn't arrive.
Sentence appended to the email system prompt in the date-fix rerun date_suffix · 20 characters · from eval/check_date_fix.py
今日は2026年10月10日(土)です。
System prompt for Japanese business email email_system · 92 characters · from eval/run_tasks.py
あなたは日本企業で働く人のメール作成を手伝うアシスタントです。依頼内容をもとに、そのまま送れるビジネスメールを書いてください。件名と本文だけを出力し、説明や前置きは書かないでください。
JSON schema sent with the extraction request (structured output) extract_schema · 553 characters · from eval/run_tasks.py
{
  "type": "object",
  "properties": {
    "company": {
      "type": "string"
    },
    "invoice_no": {
      "type": "string"
    },
    "issue_date": {
      "type": "string"
    },
    "due_date": {
      "type": "string"
    },
    "amount": {
      "type": "integer"
    },
    "currency": {
      "type": "string",
      "enum": [
        "JPY",
        "USD",
        "EUR"
      ]
    }
  },
  "required": [
    "company",
    "invoice_no",
    "issue_date",
    "due_date",
    "amount",
    "currency"
  ],
  "additionalProperties": false
}
System prompt for invoice extraction extract_system · 551 characters · from eval/run_tasks.py
Extract the invoice that the reader is being asked to pay from the message. Return JSON with:
- company: the issuing (billing) company, exactly as written, including its legal suffix
- invoice_no: the invoice number exactly as written
- issue_date, due_date: ISO 8601 (YYYY-MM-DD); convert Japanese era years and terms like "net 30"
- amount: total payable including tax, as an integer in whole currency units; if only a subtotal and tax are given, add them; if the message corrects an earlier amount, use the corrected one
- currency: JPY, USD or EUR
System prompt for both email judges judge_system · 1,248 characters · from eval/judge_emails.py
You are an experienced reviewer of Japanese business email at a Japanese company.
You will see a request an office worker typed to an AI assistant, and two candidate emails (A and B) written from it.
Judge each email on whether it could be sent as-is. Score 1-5 (5 best):
- accuracy: every fact in the request (people, companies, dates, amounts, quantities, conditions, sender) is present and correct
- keigo: honorifics and business manners fit the relationship (external partner, first contact, customer, boss, colleague)
- natural: reads like a native Japanese business writer; no translationese or awkward phrasing
- concise: length fits the situation; no padding or content the request didn't need
- sendable: 5 = send with no edits, 3 = needs a few edits, 1 = needs a rewrite
List concrete problems only:
- missing_or_wrong: facts from the request that are missing or stated incorrectly
- invented: specifics or commitments not in the request that could cause trouble if sent (new dates, amounts, promises, policies)
winner: the email a careful Japanese manager would rather send ("A", "B" or "tie"). Do not prefer an email for being longer or shorter.
Write reason and the problem lists in Simplified Chinese; reason is one or two sentences.
User message template for the judges ({brief}, {email_a}, {email_b} are filled in) judge_user_template · 55 characters · from eval/judge_emails.py
## 依頼内容
{brief}

## メール A
{email_a}

## メール B
{email_b}

Datasets

Checked 2026-10-11

All task data was written for this test by AI agents: 50 support tickets (25 English, 25 Japanese), 25 English and 25 Japanese invoice messages with gold field values, and 25 short Japanese email requests. No real customer data was used. The English, Japanese and Chinese tokenizer samples were also written by AI agents; four email addresses in them that looked like real domains were changed to .example. The code and JSON samples are excerpts of real open-source files (see corpus licenses).

A script checks every dataset mechanically before the site is built. These are the only data-quality claims we make. For the invoices, each gold field records the source text it was taken from; the script checks that this text appears verbatim in the document and that the gold values are well-formed. It does not check that a normalized gold value (an amount as a plain integer, a date in ISO format) means the same as its source text; that was set when the data was written.

FileItemsCheckPassed
eval/classify.jsonl50label is one of the allowed labels50 of 50
ids are unique50 of 50
ticket text is not empty50 of 50
eval/email_ja.jsonl25ids are unique25 of 25
request text is not empty25 of 25
requests used in the report are present3 of 3
eval/extract_en.jsonl25the source text recorded for each gold field appears verbatim in its document155 of 155
gold values are complete and well-formed: valid dates, due date not before issue date, whole-number amount, known currency25 of 25
ids are unique25 of 25
eval/extract_ja.jsonl25the source text recorded for each gold field appears verbatim in its document153 of 153
gold values are complete and well-formed: valid dates, due date not before issue date, whole-number amount, known currency25 of 25
ids are unique25 of 25

How date errors are counted

eval/audit_dates.py reads every email and checks each date written with a weekday or a year against the 2026 and 2027 calendars. It recognizes four ways of writing a date (after normalizing full-width digits and brackets): 「10月16日(金)」 or 「2026年10月16日(金)」, 「2026年10月16日」, 「金曜日(10月16日)」, and 「10/16(金)」. A year is only checked when the full year, month and day are written, and it must be 2026 or 2027 (「2025年度版」 is not a date). A weekday without a year counts as correct if it matches either year. Claude Haiku 5.5: 7 of 25 emails with at least one error; GPT-6 Luna: 0 of 25 emails with at least one error. The same script checked the 14 outputs of the date-fix rerun.

How placeholders are counted

An email counts once if it contains any of these:

  1. an empty fill-in bracket: a label and colon followed only by spaces, such as 「(連絡先: )」, an empty bracket, or a bracket that tells the writer to fill something in (「ここに」「記入」「入力」「挿入」), such as 「(URLをここに記載)」;
  2. placeholder characters: runs of 〇 or ○, or of X or x, including inside names, phone numbers and email addresses;
  3. a 【…】 or [...] item that mentions a URL or asks for input, or that stands alone on a line with nothing under it.

A heading such as 【送付資料】 followed by real content does not count. The rules and real examples are part of the export script's self-test.

Labels and gold values

The tickets and their labels were written together, 10 per label across 5 labels (billing, bug, feature_request, account, cancellation). The invoice gold values were written with each document, including 8 Japanese and 8 English hard cases, such as totals given only as subtotal plus tax and amounts corrected later in the same thread. We did not run an independent human relabeling; the checks above are mechanical. Both models matched every label and every gold value, so these two tasks do not separate them on accuracy.

Judging

Judged 2026-10-10

Each email pair was judged by Claude Opus 5.5 (Anthropic) and GPT-6.1 Sol (OpenAI). A judge saw the request and the two emails as A and B, without model names, and judged every pair twice with A and B swapped. For each email it scored five criteria from 1 to 5 (accuracy, keigo, naturalness, concision, ready to send as-is), listed specifics or commitments not in the request and facts that were missing or wrong, and picked A, B or a tie.

The two verdicts for a pair are combined by adding them: a win counts +1 for that model and a tie 0. A positive total is a win, a negative total a loss, and zero a tie, so one win plus one tie still counts as a win.

Before judging, both emails go through the same cleanup (review/build.py clean()): trailing spaces are removed from every line, leading and trailing blank lines are dropped, and three or more line breaks are collapsed to one blank line. This keeps formatting habits from giving away which model wrote which email; the text itself is not changed.

JudgeLuna / tie / HaikuSame verdict in both orders
Claude Opus 5.522 / 2 / 123 of 25
GPT-6.1 Sol25 / 0 / 025 of 25

The two judges agreed on 22 of 25 pairs. Judges are AI models, not native-speaker reviewers.

Prices and sources

Every price, tier rule, batch discount and context limit comes from a vendor page, opened and checked on the date shown. The export fails if any price field is neither covered by a source nor listed as an assumption. All costs on the site are recomputed from the billed tokens at these prices.

USD per million tokens
ModelBase: in / out / cache readHigher tierBatchMax input
Claude Haiku 5.5$0.10 / $0.50 / $0.01above 100K: $0.50 / $2.50 / $0.05, whole request× 0.51,000,000
GPT-6 Luna$0.10 / $0.50 / $0.01above 272K: $0.20 / $0.75 / $0.02, whole request× 0.5922,000
Claude Opus 5.5$4.00 / $20.00 / $0.20none× 0.51,000,000
GPT-6.1 Sol$2.00 / $10.00 / $0.10none× 0.5922,000

Assumptions

Values we could not confirm from a source are labeled assumption wherever they appear.

  • assumption Claude Haiku 5.5, batch_multiplier.cache_read: The batch table lists only input and output prices; we assume the 50% batch discount also applies to cache reads ("Batch API requests are 50% off").
  • assumption Calculator scenario “Ticket classification (Japanese)”: Inputs come from the measured run, but the calculator estimates prompt tokens from character counts by language. This task's short, keyword-dense English system prompt makes that estimate run more than 15% above Haiku's measured cost, so we label the scenario an assumption. The measured table shows the real bill.
  • assumption Calculator scenario “Code review”: Output length and reasoning tokens are assumptions, not measurements. The reasoning numbers reuse the email task's averages.
  • assumption Calculator scenario “English long document → Japanese summary”: Output length and reasoning tokens are assumptions, not measurements. The reasoning numbers reuse the email task's averages.
  • assumption Haiku's reasoning tokens: the Anthropic API bills reasoning and visible output together, so we estimate Haiku's reasoning tokens as billed output tokens minus the token count of the visible answer. The measured task costs use the API's billed totals, not this estimate; the calculator's scenarios do use it as Haiku's reasoning tokens. Luna reports its reasoning tokens directly.
  • assumption Characters to tokens: the calculator converts characters to tokens with the average ratio for each content type (table on the calculator page). Real text varies; the Japanese genres alone range from ×1.15 to ×1.34 (3 samples each).

Limits

  • Small samples: 50 tickets, 50 invoices and 25 email requests per model, 1 run each. Results can shift with another run.
  • All task data was written by AI agents, not taken from real workloads. Real tickets, invoices and requests are messier.
  • The email preference comes from two AI judges, not from native Japanese speakers.
  • Haiku's reasoning tokens are estimated: billed output tokens minus the counted visible answer. Luna reports its reasoning tokens directly.
  • Latency was measured from one client location at one time of day.
  • Prices and model behavior change. Every number carries the date it was measured or checked.

Cost of this test

About $1.30 in API fees, including judging and superseded runs. Sum of every billed call saved in eval/results, at current list prices. A few ad-hoc calls made during exploration were not saved and are not included. Token-counting calls are free and count as $0.

Cost by result file
FileUSDNote
eval/results/classify-haiku-20261009-200854.jsonl$0
eval/results/classify-haiku-20261009-233558.jsonl$0.00016
eval/results/classify-haiku-20261009-233602.jsonl$0.0016
eval/results/classify-haiku-20261011-181146.jsonl$0.0016
eval/results/classify-luna-20261009-200854.jsonl$0.00015
eval/results/classify-luna-20261009-233602.jsonl$0.0017
eval/results/classify-luna-20261011-181146.jsonl$0.0017
eval/results/classify-summary-20261009-200854.json$0summary file; its calls are counted in the row files
eval/results/classify-summary-20261009-233558.json$0summary file; its calls are counted in the row files
eval/results/classify-summary-20261009-233602.json$0summary file; its calls are counted in the row files
eval/results/classify-summary-20261011-181146.json$0summary file; its calls are counted in the row files
eval/results/coeff-20261010-011346.json$0token counting or local check; not billed
eval/results/coeff-20261011-181620.json$0token counting or local check; not billed
eval/results/coeff-20261011-182428.json$0token counting or local check; not billed
eval/results/coeff-rows-20261010-011346.jsonl$0token counting or local check; not billed
eval/results/coeff-rows-20261011-181620.jsonl$0token counting or local check; not billed
eval/results/coeff-rows-20261011-182428.jsonl$0token counting or local check; not billed
eval/results/dataset-check-20261011-181205.json$0token counting or local check; not billed
eval/results/date-audit-20261011-181215.json$0token counting or local check; not billed
eval/results/date-audit-20261011-181315.json$0token counting or local check; not billed
eval/results/date-fix-20261011-181315.json$0.0039
eval/results/email-haiku-20261010-020535.jsonl$0.00068
eval/results/email-haiku-20261010-020957.jsonl$0.0059
eval/results/email-luna-20261010-020535.jsonl$0.00051
eval/results/email-luna-20261010-020957.jsonl$0.004
eval/results/extract-haiku-20261010-020957.jsonl$0.0073
eval/results/extract-luna-20261010-020957.jsonl$0.0045
eval/results/judge-20261010-040157.jsonl$0.097
eval/results/judge-20261010-040432.jsonl$1.17
eval/results/judge-summary-20261010-040157.json$0summary file; its calls are counted in the row files
eval/results/judge-summary-20261010-040432.json$0summary file; its calls are counted in the row files
eval/results/schema-overhead-20261011-181144.json$0token counting or local check; not billed
eval/results/tasks-summary-20261010-020535.json$0summary file; its calls are counted in the row files
eval/results/tasks-summary-20261010-020957.json$0summary file; its calls are counted in the row files
eval/results/tokens-20261009-234710.json$0token counting or local check; not billed

Corpus sources and licenses

The tokenizer corpus has 24 English, 24 Japanese, 24 Chinese, 24 Python, 24 JavaScript/TypeScript, 15 JSON samples. en/ja/zh samples are original texts written for this test by AI agents (eval/prose.json). Python samples are excerpts of the Python 3.12.13 standard library; JavaScript/TypeScript and JSON samples are excerpts of npm packages, taken only from packages under MIT, ISC, BSD-2-Clause, BSD-3-Clause, Apache-2.0, CC0-1.0. Copyright lines in the samples are kept exactly as written, and each package's license text ships in the download under LICENSES/.

Replaced on 2026-10-11 (code_js_ts, json, prose emails): Earlier JS/TS and JSON samples came from bundled front-end files inside wheel packages in the uv cache, whose licenses could not be verified one by one; they were re-sampled from site/node_modules with each license recorded. Four email addresses in the prose samples that looked like real domains were changed to .example. Standard-library samples were re-drawn from a pinned Python version.

Every code and JSON sample
SampleSource filePackageLicenseLicense text
code_js_ts:01@astrojs/compiler-binding/index.js@astrojs/[email protected]MITLICENSES/@astrojs/[email protected]/LICENSE
code_js_ts:02aria-query/lib/rolesMap.js[email protected]Apache-2.0LICENSES/[email protected]/LICENSE
code_js_ts:03css-tree/lib/lexer/error.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:04csso/lib/clean/Rule.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:05domutils/lib/esm/querying.js[email protected]BSD-2-ClauseLICENSES/[email protected]/LICENSE
code_js_ts:06postcss/lib/map-generator.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:07prismjs/components/prism-dax.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:08prismjs/components/prism-mermaid.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:09prismjs/components/prism-smarty.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:10prismjs/plugins/match-braces/prism-match-braces.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:11source-map-js/lib/util.js[email protected]BSD-3-ClauseLICENSES/[email protected]/LICENSE
code_js_ts:12svgo/plugins/moveElemsAttrsToGroup.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:13undici/lib/core/socks5-client.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:14undici/lib/interceptor/decompress.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:15undici/lib/web/websocket/events.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:16zod/src/v3/tests/map.test.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:17zod/src/v4/classic/tests/cyclic-data.test.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:18zod/src/v4/classic/tests/recursive-types.test.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:19zod/src/v4/core/tests/locales/el.test.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:20zod/src/v4/locales/de.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:21zod/src/v4/locales/mk.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:22zod/src/v4/locales/uz.ts[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:23zod/v4/locales/bn.js[email protected]MITLICENSES/[email protected]/LICENSE
code_js_ts:24zod/v4/locales/ka.js[email protected]MITLICENSES/[email protected]/LICENSE
code_python:01__future__.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:02_markupbase.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:03_pyio.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:04_threading_local.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:05base64.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:06cgi.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:07codeop.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:08copy.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:09decimal.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:10fileinput.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:11genericpath.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:12gzip.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:13ipaddress.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:14mailbox.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:15nntplib.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:16operator.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:17pickletools.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:18profile.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:19queue.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:20runpy.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:21shlex.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:22socket.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:23string.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
code_python:24sysconfig.py[email protected]PSF-2.0LICENSES/[email protected]/LICENSE.txt
json:01ci-info/vendors.json[email protected]MITLICENSES/[email protected]/LICENSE
json:02css-tree/data/patch.json[email protected]MITLICENSES/[email protected]/LICENSE
json:03csso/node_modules/mdn-data/api/inheritance.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:04csso/node_modules/mdn-data/css/at-rules.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:05csso/node_modules/mdn-data/css/selectors.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:06csso/node_modules/mdn-data/css/syntaxes.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:07csso/node_modules/mdn-data/css/types.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:08csso/node_modules/mdn-data/css/units.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:09csso/node_modules/mdn-data/l10n/css.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:10extend/.jscs.json[email protected]MITLICENSES/[email protected]/LICENSE
json:11mdn-data/css/definitions.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:12mdn-data/css/functions.json[email protected]CC0-1.0LICENSES/[email protected]/LICENSE
json:13prismjs/components.json[email protected]MITLICENSES/[email protected]/LICENSE
json:14undici/docs/docs/site.json[email protected]MITLICENSES/[email protected]/LICENSE
json:15undici/docs/docs/type-map.json[email protected]MITLICENSES/[email protected]/LICENSE

Downloads

The download mirrors the repository layout, so the scripts run from it as-is. It contains the datasets, every raw model output and judge output used on the site, the summaries, the scripts, this project's MIT license, and the license texts of the corpus samples. It contains no API keys or personal data; the email addresses in it are fictional (.example), placeholders written by the models, or copyright lines that a license requires us to keep.

Download everything (llm-bill-data-20261011.zip) manifest.json (paths, sizes, SHA-256)

Every file in the download
FileWhat it isSize
LICENSEMIT license for this project's code1.3 kB
LICENSES/@astrojs/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text10.2 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text1.3 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text6.6 kB
LICENSES/[email protected]/LICENSEThird-party license text6.6 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSE.txtThird-party license text13.9 kB
LICENSES/[email protected]/LICENSEThird-party license text1.5 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
LICENSES/[email protected]/LICENSEThird-party license text1.1 kB
corpus_sources.jsonCorpus sources and licenses50.4 kB
eval/audit_dates.pyEvaluation script5.4 kB
eval/build_corpus.pyEvaluation script12.6 kB
eval/check_date_fix.pyEvaluation script2.7 kB
eval/classify.jsonlClassification dataset7.0 kB
eval/corpus.jsonlTokenizer corpus347.3 kB
eval/corpus_licenses.jsonTokenizer corpus25.1 kB
eval/count_tokens.pyEvaluation script6.5 kB
eval/email_allowlist.jsonAllowed email addresses2.7 kB
eval/email_ja.jsonlEmail briefs11.2 kB
eval/excerpts.jsonExcerpt selection rules5.3 kB
eval/export_site_data.pyEvaluation script48.5 kB
eval/extract_en.jsonlExtraction dataset30.7 kB
eval/extract_ja.jsonlExtraction dataset37.7 kB
eval/judge_emails.pyEvaluation script9.1 kB
eval/measure_coeff.pyEvaluation script4.1 kB
eval/measure_schema_overhead.pyEvaluation script4.0 kB
eval/presets.jsonCalculator presets2.7 kB
eval/prices.jsonPrices with sources4.6 kB
eval/prose.jsonOriginal prose samples (AI-written)130.2 kB
eval/results/classify-haiku-20261011-181146.jsonlTicket classification raw outputs / summary14.0 kB
eval/results/classify-luna-20261011-181146.jsonlTicket classification raw outputs / summary14.3 kB
eval/results/classify-summary-20261011-181146.jsonTicket classification raw outputs / summary1.1 kB
eval/results/coeff-20261011-182428.jsonTokenizer coefficients8.3 kB
eval/results/coeff-rows-20261011-182428.jsonlTokenizer coefficients14.4 kB
eval/results/dataset-check-20261011-181205.jsonDataset checks2.4 kB
eval/results/date-audit-20261011-181315.jsonDate audit1.9 kB
eval/results/date-fix-20261011-181315.jsonDate-fix rerun outputs18.5 kB
eval/results/email-haiku-20261010-020957.jsonlJapanese email raw outputs36.7 kB
eval/results/email-luna-20261010-020957.jsonlJapanese email raw outputs26.3 kB
eval/results/extract-haiku-20261010-020957.jsonlInvoice extraction raw outputs36.2 kB
eval/results/extract-luna-20261010-020957.jsonlInvoice extraction raw outputs36.1 kB
eval/results/judge-20261010-040432.jsonlBlind judging outputs / summary85.6 kB
eval/results/judge-summary-20261010-040432.jsonBlind judging outputs / summary2.4 kB
eval/results/schema-overhead-20261011-181144.jsonStructured-output overhead0.6 kB
eval/results/tasks-summary-20261010-020957.jsonExtraction and email summary2.1 kB
eval/run_classify.pyEvaluation script8.9 kB
eval/run_tasks.pyEvaluation script9.8 kB
eval/validate_datasets.pyEvaluation script5.6 kB
review/build.pyShared cleaning helper (used by judge_emails.py)1.7 kB

Rerun it

You need uv (it installs the pinned Python and dependencies from each script's header) and, for the billed steps, your own API keys. Unzip the download and run from its top folder.

# 1. Self-tests and dataset checks: free, no API calls
for s in eval/*.py; do uv run -q "$s" --selftest; done
uv run -q eval/validate_datasets.py

# 2. API keys for the next steps. Paste each key at the prompt: the input is
#    hidden and never typed on a command line, so it stays out of shell history.
printf 'Anthropic API key: '; read -rs ANTHROPIC_API_KEY; echo; export ANTHROPIC_API_KEY
printf 'OpenAI API key: '; read -rs OPENAI_API_KEY; echo; export OPENAI_API_KEY
# Only if your Anthropic key needs a workspace (an id, not a secret):
# export ANTHROPIC_WORKSPACE_ID=<workspace id>

# 3. Tokenizer ratios and schema overhead: token-counting endpoints, free
uv run -q eval/measure_coeff.py
uv run -q eval/measure_schema_overhead.py

# 4. The three tasks, then the judges: billed
uv run -q eval/run_classify.py
uv run -q eval/run_tasks.py
uv run -q eval/judge_emails.py

# 5. Date check and the date-fix rerun
uv run -q eval/audit_dates.py
uv run -q eval/check_date_fix.py
uv run -q eval/audit_dates.py

Each script writes a new timestamped file to eval/results/ and never overwrites earlier runs. Our complete measurement, including superseded runs, cost $1.30; see the cost table above for the per-file split.

How the numbers are checked

Pages never contain hand-typed figures. Measured numbers come from the data files above; calculator numbers (default bills, ratios, tier points) are computed when the site is built by the same code the calculator runs in your browser.

After every build, a check reads the visible text of each page. Each calculator number is marked and recomputed from the data and must match exactly. Every other number must appear in the data files (allowing common formatting such as rounding, percentages, K/M and thousands separators) or in a short list of model version numbers, years and method constants. Any mismatch fails the build.

Limits of the check: single-digit numbers almost always match something in the data, so the check says little about them; numbers written as words or kanji numerals are not checked; dates are skipped.