POLYTONE — the AI-native programming language

The generation benchmark

“Leading LLM programming language” is an empirical claim: a model receives a task, writes POLYTONE, and ptc test judges — deterministically, no human in the loop. Two numbers per task, both first-class: pass@1 (the first candidate passes) and pass@2e (after a failure the model sees the actual error and gets one repair attempt — POLYTONE’s thesis is that its errors teach).

Methodology

Results

claude-fable-52026-08-14 (run 7, toolchain 0.35.239 + card v14, corpus 36) · workflow-harness

pass@1 34/36 · pass@2e 36/36 · tokens-to-green median 62827 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status63220
audio_probe63857
bit_parity62081
clock_iso62282
config_port65270
content_tag62631
csv_totals63061
dice_walk62228
env_mode62063
grade_book63908
hex_dump62190
histogram61990
image_probe62940
json_pluck64443
log_scan65030
mesh_probe63003
money_order63404
parse_point61959
rider_film65927
row_sums62012
run_length62627
save_report125537
season_label62479
set_overlap61873
shape_area62239
stereo_field63284
swap_pairs62142
task_batch62198
top_scorer63275
total_of62348
uniques62357
video_probe62295
web_toc64580
week_later64442
wipe_reveal127583
word_stats62827
claude-fable-52026-08-11 (run 6, toolchain 0.35.232 + card v14) · workflow-harness

pass@1 33/33 · pass@2e 33/33 · tokens-to-green median 56514 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status56619
audio_probe56530
bit_parity56499
clock_iso56460
config_port56582
content_tag56516
csv_totals56600
dice_walk56534
env_mode56451
grade_book56524
hex_dump56416
histogram56418
image_probe56520
json_pluck56554
log_scan56521
mesh_probe56491
money_order56589
parse_point56529
row_sums56432
run_length56514
save_report56541
season_label56524
set_overlap56443
shape_area56510
swap_pairs56474
task_batch56496
top_scorer56504
total_of56514
uniques56444
video_probe56458
web_toc56493
week_later56547
word_stats56519
claude-fable-52026-08-11 (run 5, toolchain 0.35.230 + card v13) · workflow-harness

pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 56471 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status56521
audio_probe56471
bit_parity56435
clock_iso56429
config_port56531
content_tag56456
csv_totals56583
dice_walk56480
env_mode56431
grade_book56488
hex_dump56404
histogram113469
image_probe56482
json_pluck114016
log_scan56471
mesh_probe56443
money_order56542
parse_point56485
row_sums56378
run_length56476
save_report56485
season_label56426
set_overlap56410
shape_area56477
swap_pairs56435
task_batch56467
top_scorer56445
total_of56485
uniques56389
video_probe56437
web_toc56436
week_later56487
word_stats113864
claude-fable-52026-08-11 (run 4, toolchain 0.35.226 + card v12) · workflow-harness

pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 56431 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status56462
audio_probe56446
bit_parity56384
clock_iso56381
config_port56513
content_tag56411
csv_totals56526
dice_walk56460
env_mode56359
grade_book56416
hex_dump56351
histogram113382
image_probe56428
json_pluck113895
log_scan56431
mesh_probe56405
money_order56503
parse_point113770
row_sums56342
run_length56432
save_report56457
season_label56378
set_overlap56434
shape_area56434
swap_pairs56392
task_batch56416
top_scorer56405
total_of56435
uniques56353
video_probe56387
web_toc56404
week_later56446
word_stats56439
claude-fable-52026-08-08 (run 3, toolchain 0.34.210 + card v12) · workflow-harness

pass@1 30/33 · pass@2e 32/33 · tokens-to-green median 50669 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status50695
audio_probe50681
bit_parity50645
clock_iso50633
config_port50746
content_tag50649
csv_totals50765
dice_walk50689
env_mode101841
grade_book50682
hex_dump50580
histogram101799
image_probe50668
json_pluck
log_scan50685
mesh_probe50636
money_order50746
parse_point50669
row_sums50607
run_length50662
save_report50678
season_label50629
set_overlap50598
shape_area50700
swap_pairs50618
task_batch50660
top_scorer50664
total_of50686
uniques50607
video_probe50624
web_toc50641
week_later50706
word_stats50684
claude-fable-52026-08-08 (run 2, toolchain 0.34.203 + card v10) · workflow-harness

pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 50655 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status50687
audio_probe50666
bit_parity50635
clock_iso50616
config_port50727
content_tag50640
csv_totals50757
dice_walk50683
env_mode50599
grade_book50667
hex_dump50572
histogram101773
image_probe101976
json_pluck102320
log_scan50677
mesh_probe50632
money_order50731
parse_point50661
row_sums50601
run_length50655
save_report50677
season_label50638
set_overlap50598
shape_area50652
swap_pairs50626
task_batch50647
top_scorer50626
total_of50665
uniques50587
video_probe50625
web_toc50633
week_later50671
word_stats50671
claude-fable-52026-08-08 · workflow-harness

pass@1 21/33 · pass@2e 30/33 · tokens-to-green median 50372 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status50401
audio_probe50368
bit_parity50324
clock_iso50297
config_port101628
content_tag50335
csv_totals50450
dice_walk50376
env_mode50297
grade_book50356
hex_dump50281
histogram101160
image_probe101370
json_pluck50385
log_scan101442
mesh_probe
money_order50422
parse_point101544
row_sums101220
run_length101400
save_report50375
season_label50314
set_overlap50285
shape_area50351
swap_pairs
task_batch101356
top_scorer50334
total_of101495
uniques50282
video_probe50308
web_toc
week_later50369
word_stats50369

The first delta proof (replays)

The phase's first failure dataset is the corpus-authoring session itself (Sprint 133) — an LLM writing POLYTONE cold, its first attempts recorded verbatim. Sprint 138 replays them against the hardened toolchain, CI-gated (gen_replays.rs):

Fresh pass@1/pass@2e runs against frontier models stay local and BYO-key — this section shows only what has actually run.

The corpus (36)

17 × tier S · 19 × tier M — each task: a prompt, a hidden judge, a reference solution.

tasktierarea
api_statusMCapabilities (mocked)
audio_probeMMedia codecs
bit_parityMBitwise & encoding
clock_isoSCapabilities (mocked)
config_portMCapabilities (mocked)
content_tagSBitwise & encoding
csv_totalsMData formats
dice_walkSCapabilities (mocked)
env_modeSCapabilities (mocked)
grade_bookSCollections
hex_dumpSBitwise & encoding
histogramSCollections
image_probeMMedia codecs
json_pluckMData formats
log_scanSData formats
mesh_probeMMedia codecs
money_orderMMethods & traits
parse_pointSErrors & Result
rider_filmMMedia codecs
row_sumsSCollections
run_lengthMPrelude & text
save_reportMCapabilities (mocked)
season_labelSMethods & traits
set_overlapSCollections
shape_areaSMethods & traits
stereo_fieldMMedia codecs
swap_pairsSGenerics
task_batchMAsync & Tasks
top_scorerSCollections
total_ofSErrors & Result
uniquesMGenerics
video_probeMMedia codecs
web_tocMMedia codecs
week_laterMData formats
wipe_revealMMedia codecs
word_statsSPrelude & text