Evaluations
A question's answer counts only where it has been measured. This page reports the text each question is trained and measured on, the thresholds and error bounds it qualified at, and what its decisions would have done on held-out text.
Phrase banks
Each row is written by one author and labelled again by a second labeller who never sees the author's label; only agreements are kept. Training and validation rows come from one group of writers. Calibration rows come from a second group the model never trains on, and holdout rows from a third, each with other regions, vocabularies, carriers and systems.
| Question | train | validation | calibration | holdout |
|---|---|---|---|---|
asn_unit | 618 | 82 | 300 | 600 |
carrier_outcome | 628 | 72 | 300 | 600 |
cost_usable | 627 | 73 | 300 | 599 |
destination_outcome | 620 | 80 | 300 | 600 |
discovered_stock | 632 | 67 | 300 | 600 |
handover_by_cutoff | 619 | 81 | 300 | 600 |
movement_during_count | 634 | 66 | 300 | 600 |
order_change_kind | 626 | 74 | 300 | 600 |
same_order | 634 | 66 | 300 | 600 |
scan_is_serial | 646 | 54 | 300 | 600 |
weight_measured | 637 | 63 | 300 | 599 |
Writers and second labellers are agents of one model family, so their agreement shows that rows read unambiguously to that family; it does not show the labels are right. An auditor from another model labelled a seeded sample of 60 rows from each holdout without seeing any label and differed from the bank on 3 of 660 rows: at most 1.2% at 95% confidence.
Qualification
Measured for model 47ed8d6e03938bfc (float32-cpu), the dtype the hosted service runs, at an error budget of 3% with 95% confidence.
| Question | Holdout rows | Coverage |
|---|---|---|
asn_unit | 600 | not qualified |
carrier_outcome | 600 | 21.7% |
cost_usable | 599 | 18.2% |
destination_outcome | 600 | not qualified |
discovered_stock | 600 | 33.5% |
handover_by_cutoff | 600 | not qualified |
movement_during_count | 600 | 30.2% |
order_change_kind | 600 | 60.2% |
same_order | 600 | 27.2% |
scan_is_serial | 600 | not qualified |
weight_measured | 599 | 30.9% |
Certified options
An option decides only when its probability clears the threshold. Decided is the number of holdout rows in that region; the bound is the 95% upper limit on its error rate.
| Question | Option | Threshold | Decided | Errors | Bound |
|---|---|---|---|---|---|
carrier_outcome | not_performed | 0.938726 | 130 | 0 | 2.3% |
cost_usable | no | 0.992683 | 109 | 0 | 2.7% |
discovered_stock | yes | 0.993178 | 201 | 1 | 2.3% |
movement_during_count | yes | 0.988847 | 181 | 1 | 2.6% |
order_change_kind | additional_units | 0.995227 | 127 | 0 | 2.3% |
order_change_kind | cancellation | 0.991816 | 101 | 0 | 2.9% |
order_change_kind | revised_quantity | 0.909274 | 133 | 0 | 2.2% |
same_order | no | 0.999187 | 163 | 1 | 2.9% |
weight_measured | yes | 0.995324 | 185 | 1 | 2.5% |
Verdicts on held-out rows
Each holdout row becomes its decision's request. Settled counts verdicts that need no person or lookup and match the verdict the agreed label gives, first with every question abstaining and then with the model. Unsafe counts the model's allow or replay verdicts where the label gives something else.
| Question | Decision | Rows | Settled by rules | Settled with model | Unsafe |
|---|---|---|---|---|---|
asn_unit | receive | 600 | 0 | 0 | 0 |
carrier_outcome | carrier | 600 | 0 | 130 | 0 |
cost_usable | variance | 599 | 0 | 109 | 0 |
destination_outcome | replay | 600 | 0 | 0 | 0 |
discovered_stock | variance | 600 | 0 | 0 | 0 |
handover_by_cutoff | handover | 600 | 0 | 0 | 0 |
movement_during_count | count | 600 | 0 | 0 | 0 |
order_change_kind | reserve | 600 | 0 | 361 | 0 |
same_order | wave | 600 | 0 | 162 | 1 |
scan_is_serial | serial | 600 | 115 | 115 | 0 |
weight_measured | pack | 599 | 0 | 185 | 0 |
The method: five thresholds per option are proposed on calibration rows and tested on holdout rows from the narrowest to the widest, each with a one-sided 95% Clopper-Pearson bound on the error rate inside the region it decides. Testing stops at the first threshold whose bound misses the error budget, so the claim holds at 95% for the whole sequence, and the widest threshold that passed is used. An option with no passing threshold never decides.