The generation benchmark
“Leading LLM programming language” is an empirical claim: a model receives a task, writes POLYTONE, and ptc test judges — deterministically, no human in the loop. Two numbers per task, both first-class: pass@1 (the first candidate passes) and pass@2e (after a failure the model sees the actual error and gets one repair attempt — POLYTONE’s thesis is that its errors teach).
Methodology
- The judge is ptc test: candidate + hidden tests → one file → pass/fail. The judge never trusts the model — tests are appended after generation.
- Capability-clean: effects appear only through mocks (mock_fs, fixed_rng, fixed_clock, mock_env, mock_http) — a run touches no network, disk, clock, or entropy.
- A CI gate keeps every reference solution green against its own hidden tests — the corpus cannot rot. CI never calls a model; runs are local and BYO-key.
- The report always shows every task × every attempt — no cherry-picking.
- The system prompt is the frozen IDE language card. Since Sprint 252 the harness sends the EVALUATED card — byte-identical with the IDE and the Pro CLI, pinned by a gate; runs 1–7 sent the raw source form (doubled backslashes). Tokens-to-green stays directly comparable only between runs of the same card form.
Results
pass@1 34/36 · pass@2e 36/36 · tokens-to-green median 62827 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 63220 |
audio_probe | ✓ | ✓ | 63857 |
bit_parity | ✓ | ✓ | 62081 |
clock_iso | ✓ | ✓ | 62282 |
config_port | ✓ | ✓ | 65270 |
content_tag | ✓ | ✓ | 62631 |
csv_totals | ✓ | ✓ | 63061 |
dice_walk | ✓ | ✓ | 62228 |
env_mode | ✓ | ✓ | 62063 |
grade_book | ✓ | ✓ | 63908 |
hex_dump | ✓ | ✓ | 62190 |
histogram | ✓ | ✓ | 61990 |
image_probe | ✓ | ✓ | 62940 |
json_pluck | ✓ | ✓ | 64443 |
log_scan | ✓ | ✓ | 65030 |
mesh_probe | ✓ | ✓ | 63003 |
money_order | ✓ | ✓ | 63404 |
parse_point | ✓ | ✓ | 61959 |
rider_film | ✓ | ✓ | 65927 |
row_sums | ✓ | ✓ | 62012 |
run_length | ✓ | ✓ | 62627 |
save_report | ✗ | ✓ | 125537 |
season_label | ✓ | ✓ | 62479 |
set_overlap | ✓ | ✓ | 61873 |
shape_area | ✓ | ✓ | 62239 |
stereo_field | ✓ | ✓ | 63284 |
swap_pairs | ✓ | ✓ | 62142 |
task_batch | ✓ | ✓ | 62198 |
top_scorer | ✓ | ✓ | 63275 |
total_of | ✓ | ✓ | 62348 |
uniques | ✓ | ✓ | 62357 |
video_probe | ✓ | ✓ | 62295 |
web_toc | ✓ | ✓ | 64580 |
week_later | ✓ | ✓ | 64442 |
wipe_reveal | ✗ | ✓ | 127583 |
word_stats | ✓ | ✓ | 62827 |
pass@1 33/33 · pass@2e 33/33 · tokens-to-green median 56514 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 56619 |
audio_probe | ✓ | ✓ | 56530 |
bit_parity | ✓ | ✓ | 56499 |
clock_iso | ✓ | ✓ | 56460 |
config_port | ✓ | ✓ | 56582 |
content_tag | ✓ | ✓ | 56516 |
csv_totals | ✓ | ✓ | 56600 |
dice_walk | ✓ | ✓ | 56534 |
env_mode | ✓ | ✓ | 56451 |
grade_book | ✓ | ✓ | 56524 |
hex_dump | ✓ | ✓ | 56416 |
histogram | ✓ | ✓ | 56418 |
image_probe | ✓ | ✓ | 56520 |
json_pluck | ✓ | ✓ | 56554 |
log_scan | ✓ | ✓ | 56521 |
mesh_probe | ✓ | ✓ | 56491 |
money_order | ✓ | ✓ | 56589 |
parse_point | ✓ | ✓ | 56529 |
row_sums | ✓ | ✓ | 56432 |
run_length | ✓ | ✓ | 56514 |
save_report | ✓ | ✓ | 56541 |
season_label | ✓ | ✓ | 56524 |
set_overlap | ✓ | ✓ | 56443 |
shape_area | ✓ | ✓ | 56510 |
swap_pairs | ✓ | ✓ | 56474 |
task_batch | ✓ | ✓ | 56496 |
top_scorer | ✓ | ✓ | 56504 |
total_of | ✓ | ✓ | 56514 |
uniques | ✓ | ✓ | 56444 |
video_probe | ✓ | ✓ | 56458 |
web_toc | ✓ | ✓ | 56493 |
week_later | ✓ | ✓ | 56547 |
word_stats | ✓ | ✓ | 56519 |
pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 56471 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 56521 |
audio_probe | ✓ | ✓ | 56471 |
bit_parity | ✓ | ✓ | 56435 |
clock_iso | ✓ | ✓ | 56429 |
config_port | ✓ | ✓ | 56531 |
content_tag | ✓ | ✓ | 56456 |
csv_totals | ✓ | ✓ | 56583 |
dice_walk | ✓ | ✓ | 56480 |
env_mode | ✓ | ✓ | 56431 |
grade_book | ✓ | ✓ | 56488 |
hex_dump | ✓ | ✓ | 56404 |
histogram | ✗ | ✓ | 113469 |
image_probe | ✓ | ✓ | 56482 |
json_pluck | ✗ | ✓ | 114016 |
log_scan | ✓ | ✓ | 56471 |
mesh_probe | ✓ | ✓ | 56443 |
money_order | ✓ | ✓ | 56542 |
parse_point | ✓ | ✓ | 56485 |
row_sums | ✓ | ✓ | 56378 |
run_length | ✓ | ✓ | 56476 |
save_report | ✓ | ✓ | 56485 |
season_label | ✓ | ✓ | 56426 |
set_overlap | ✓ | ✓ | 56410 |
shape_area | ✓ | ✓ | 56477 |
swap_pairs | ✓ | ✓ | 56435 |
task_batch | ✓ | ✓ | 56467 |
top_scorer | ✓ | ✓ | 56445 |
total_of | ✓ | ✓ | 56485 |
uniques | ✓ | ✓ | 56389 |
video_probe | ✓ | ✓ | 56437 |
web_toc | ✓ | ✓ | 56436 |
week_later | ✓ | ✓ | 56487 |
word_stats | ✗ | ✓ | 113864 |
pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 56431 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 56462 |
audio_probe | ✓ | ✓ | 56446 |
bit_parity | ✓ | ✓ | 56384 |
clock_iso | ✓ | ✓ | 56381 |
config_port | ✓ | ✓ | 56513 |
content_tag | ✓ | ✓ | 56411 |
csv_totals | ✓ | ✓ | 56526 |
dice_walk | ✓ | ✓ | 56460 |
env_mode | ✓ | ✓ | 56359 |
grade_book | ✓ | ✓ | 56416 |
hex_dump | ✓ | ✓ | 56351 |
histogram | ✗ | ✓ | 113382 |
image_probe | ✓ | ✓ | 56428 |
json_pluck | ✗ | ✓ | 113895 |
log_scan | ✓ | ✓ | 56431 |
mesh_probe | ✓ | ✓ | 56405 |
money_order | ✓ | ✓ | 56503 |
parse_point | ✗ | ✓ | 113770 |
row_sums | ✓ | ✓ | 56342 |
run_length | ✓ | ✓ | 56432 |
save_report | ✓ | ✓ | 56457 |
season_label | ✓ | ✓ | 56378 |
set_overlap | ✓ | ✓ | 56434 |
shape_area | ✓ | ✓ | 56434 |
swap_pairs | ✓ | ✓ | 56392 |
task_batch | ✓ | ✓ | 56416 |
top_scorer | ✓ | ✓ | 56405 |
total_of | ✓ | ✓ | 56435 |
uniques | ✓ | ✓ | 56353 |
video_probe | ✓ | ✓ | 56387 |
web_toc | ✓ | ✓ | 56404 |
week_later | ✓ | ✓ | 56446 |
word_stats | ✓ | ✓ | 56439 |
pass@1 30/33 · pass@2e 32/33 · tokens-to-green median 50669 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 50695 |
audio_probe | ✓ | ✓ | 50681 |
bit_parity | ✓ | ✓ | 50645 |
clock_iso | ✓ | ✓ | 50633 |
config_port | ✓ | ✓ | 50746 |
content_tag | ✓ | ✓ | 50649 |
csv_totals | ✓ | ✓ | 50765 |
dice_walk | ✓ | ✓ | 50689 |
env_mode | ✗ | ✓ | 101841 |
grade_book | ✓ | ✓ | 50682 |
hex_dump | ✓ | ✓ | 50580 |
histogram | ✗ | ✓ | 101799 |
image_probe | ✓ | ✓ | 50668 |
json_pluck | ✗ | ✗ | — |
log_scan | ✓ | ✓ | 50685 |
mesh_probe | ✓ | ✓ | 50636 |
money_order | ✓ | ✓ | 50746 |
parse_point | ✓ | ✓ | 50669 |
row_sums | ✓ | ✓ | 50607 |
run_length | ✓ | ✓ | 50662 |
save_report | ✓ | ✓ | 50678 |
season_label | ✓ | ✓ | 50629 |
set_overlap | ✓ | ✓ | 50598 |
shape_area | ✓ | ✓ | 50700 |
swap_pairs | ✓ | ✓ | 50618 |
task_batch | ✓ | ✓ | 50660 |
top_scorer | ✓ | ✓ | 50664 |
total_of | ✓ | ✓ | 50686 |
uniques | ✓ | ✓ | 50607 |
video_probe | ✓ | ✓ | 50624 |
web_toc | ✓ | ✓ | 50641 |
week_later | ✓ | ✓ | 50706 |
word_stats | ✓ | ✓ | 50684 |
pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 50655 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 50687 |
audio_probe | ✓ | ✓ | 50666 |
bit_parity | ✓ | ✓ | 50635 |
clock_iso | ✓ | ✓ | 50616 |
config_port | ✓ | ✓ | 50727 |
content_tag | ✓ | ✓ | 50640 |
csv_totals | ✓ | ✓ | 50757 |
dice_walk | ✓ | ✓ | 50683 |
env_mode | ✓ | ✓ | 50599 |
grade_book | ✓ | ✓ | 50667 |
hex_dump | ✓ | ✓ | 50572 |
histogram | ✗ | ✓ | 101773 |
image_probe | ✗ | ✓ | 101976 |
json_pluck | ✗ | ✓ | 102320 |
log_scan | ✓ | ✓ | 50677 |
mesh_probe | ✓ | ✓ | 50632 |
money_order | ✓ | ✓ | 50731 |
parse_point | ✓ | ✓ | 50661 |
row_sums | ✓ | ✓ | 50601 |
run_length | ✓ | ✓ | 50655 |
save_report | ✓ | ✓ | 50677 |
season_label | ✓ | ✓ | 50638 |
set_overlap | ✓ | ✓ | 50598 |
shape_area | ✓ | ✓ | 50652 |
swap_pairs | ✓ | ✓ | 50626 |
task_batch | ✓ | ✓ | 50647 |
top_scorer | ✓ | ✓ | 50626 |
total_of | ✓ | ✓ | 50665 |
uniques | ✓ | ✓ | 50587 |
video_probe | ✓ | ✓ | 50625 |
web_toc | ✓ | ✓ | 50633 |
week_later | ✓ | ✓ | 50671 |
word_stats | ✓ | ✓ | 50671 |
pass@1 21/33 · pass@2e 30/33 · tokens-to-green median 50372 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)
| task | pass@1 | pass@2e | tok→green |
|---|---|---|---|
api_status | ✓ | ✓ | 50401 |
audio_probe | ✓ | ✓ | 50368 |
bit_parity | ✓ | ✓ | 50324 |
clock_iso | ✓ | ✓ | 50297 |
config_port | ✗ | ✓ | 101628 |
content_tag | ✓ | ✓ | 50335 |
csv_totals | ✓ | ✓ | 50450 |
dice_walk | ✓ | ✓ | 50376 |
env_mode | ✓ | ✓ | 50297 |
grade_book | ✓ | ✓ | 50356 |
hex_dump | ✓ | ✓ | 50281 |
histogram | ✗ | ✓ | 101160 |
image_probe | ✗ | ✓ | 101370 |
json_pluck | ✓ | ✓ | 50385 |
log_scan | ✗ | ✓ | 101442 |
mesh_probe | ✗ | ✗ | — |
money_order | ✓ | ✓ | 50422 |
parse_point | ✗ | ✓ | 101544 |
row_sums | ✗ | ✓ | 101220 |
run_length | ✗ | ✓ | 101400 |
save_report | ✓ | ✓ | 50375 |
season_label | ✓ | ✓ | 50314 |
set_overlap | ✓ | ✓ | 50285 |
shape_area | ✓ | ✓ | 50351 |
swap_pairs | ✗ | ✗ | — |
task_batch | ✗ | ✓ | 101356 |
top_scorer | ✓ | ✓ | 50334 |
total_of | ✗ | ✓ | 101495 |
uniques | ✓ | ✓ | 50282 |
video_probe | ✓ | ✓ | 50308 |
web_toc | ✗ | ✗ | — |
week_later | ✓ | ✓ | 50369 |
word_stats | ✓ | ✓ | 50369 |
The first delta proof (replays)
The phase's first failure dataset is the corpus-authoring session itself (Sprint 133) — an LLM writing POLYTONE cold, its first attempts recorded verbatim. Sprint 138 replays them against the hardened toolchain, CI-gated (gen_replays.rs):
content_tag: failed on the missingText.slice— now passes verbatim (the language grew to meet the model, Sprint 137).json_pluck: the error read like nonsense — now teaches the qualified form (the repair signal pass@2e depends on).log_scan: still fails (a semantics miss no compiler error can prevent) — but the assertion diff shows the actual values, and Card v6 teaches the rule.
Fresh pass@1/pass@2e runs against frontier models stay local and BYO-key — this section shows only what has actually run.
The corpus (36)
17 × tier S · 19 × tier M — each task: a prompt, a hidden judge, a reference solution.
| task | tier | area |
|---|---|---|
api_status | M | Capabilities (mocked) |
audio_probe | M | Media codecs |
bit_parity | M | Bitwise & encoding |
clock_iso | S | Capabilities (mocked) |
config_port | M | Capabilities (mocked) |
content_tag | S | Bitwise & encoding |
csv_totals | M | Data formats |
dice_walk | S | Capabilities (mocked) |
env_mode | S | Capabilities (mocked) |
grade_book | S | Collections |
hex_dump | S | Bitwise & encoding |
histogram | S | Collections |
image_probe | M | Media codecs |
json_pluck | M | Data formats |
log_scan | S | Data formats |
mesh_probe | M | Media codecs |
money_order | M | Methods & traits |
parse_point | S | Errors & Result |
rider_film | M | Media codecs |
row_sums | S | Collections |
run_length | M | Prelude & text |
save_report | M | Capabilities (mocked) |
season_label | S | Methods & traits |
set_overlap | S | Collections |
shape_area | S | Methods & traits |
stereo_field | M | Media codecs |
swap_pairs | S | Generics |
task_batch | M | Async & Tasks |
top_scorer | S | Collections |
total_of | S | Errors & Result |
uniques | M | Generics |
video_probe | M | Media codecs |
web_toc | M | Media codecs |
week_later | M | Data formats |
wipe_reveal | M | Media codecs |
word_stats | S | Prelude & text |