Evaluations

A question's answer counts only where it has been measured. This page reports the text each question is trained and measured on, the thresholds and error bounds it qualified at, and what its decisions would have done on held-out text.

Phrase banks

Each row is written by one author and labelled again by a second labeller who never sees the author's label; only agreements are kept. Training and validation rows come from one group of writers. Calibration rows come from a second group the model never trains on, and holdout rows from a third, each with other regions, vocabularies, carriers and systems.

Questiontrainvalidationcalibrationholdout
asn_unit61882300600
carrier_outcome62872300600
cost_usable62773300599
destination_outcome62080300600
discovered_stock63267300600
handover_by_cutoff61981300600
movement_during_count63466300600
order_change_kind62674300600
same_order63466300600
scan_is_serial64654300600
weight_measured63763300599

Writers and second labellers are agents of one model family, so their agreement shows that rows read unambiguously to that family; it does not show the labels are right. An auditor from another model labelled a seeded sample of 60 rows from each holdout without seeing any label and differed from the bank on 3 of 660 rows: at most 1.2% at 95% confidence.

Qualification

Measured for model 47ed8d6e03938bfc (float32-cpu), the dtype the hosted service runs, at an error budget of 3% with 95% confidence.

QuestionHoldout rowsCoverage
asn_unit600not qualified
carrier_outcome60021.7%
cost_usable59918.2%
destination_outcome600not qualified
discovered_stock60033.5%
handover_by_cutoff600not qualified
movement_during_count60030.2%
order_change_kind60060.2%
same_order60027.2%
scan_is_serial600not qualified
weight_measured59930.9%

Certified options

An option decides only when its probability clears the threshold. Decided is the number of holdout rows in that region; the bound is the 95% upper limit on its error rate.

QuestionOptionThresholdDecidedErrorsBound
carrier_outcomenot_performed0.93872613002.3%
cost_usableno0.99268310902.7%
discovered_stockyes0.99317820112.3%
movement_during_countyes0.98884718112.6%
order_change_kindadditional_units0.99522712702.3%
order_change_kindcancellation0.99181610102.9%
order_change_kindrevised_quantity0.90927413302.2%
same_orderno0.99918716312.9%
weight_measuredyes0.99532418512.5%

Verdicts on held-out rows

Each holdout row becomes its decision's request. Settled counts verdicts that need no person or lookup and match the verdict the agreed label gives, first with every question abstaining and then with the model. Unsafe counts the model's allow or replay verdicts where the label gives something else.

QuestionDecisionRowsSettled by rulesSettled with modelUnsafe
asn_unitreceive600000
carrier_outcomecarrier60001300
cost_usablevariance59901090
destination_outcomereplay600000
discovered_stockvariance600000
handover_by_cutoffhandover600000
movement_during_countcount600000
order_change_kindreserve60003610
same_orderwave60001621
scan_is_serialserial6001151150
weight_measuredpack59901850

The method: five thresholds per option are proposed on calibration rows and tested on holdout rows from the narrowest to the widest, each with a one-sided 95% Clopper-Pearson bound on the error rate inside the region it decides. Testing stops at the first threshold whose bound misses the error budget, so the claim holds at 95% for the whole sequence, and the widest threshold that passed is used. An option with no passing threshold never decides.