Severity ordering and confidence-segment review bands
A real-data evaluation of the word model with severity as the primary queue ordering, development-only diagnostics, and a stricter rule for assigning the priority label. Every moderation action still requires human confirmation.
What changed
The primary ordering is severity: for each comment, take the maximum over labels of
calibrated probability × that label's severity weight, then review the highest-scoring comment first.
The priority_review label no longer guarantees earlier service.
Band-first priority, probability and FIFO remain comparisons on the same admitted comments and arrival streams.
For each label, the priority threshold is the lowest declared score edge for which every segment above it meets the 95.00% empirical agreement target with at least 30 rows. These are score ranges, not statistical confidence intervals. Human-review thresholds and identity-term subgroup thresholds retain cumulative precision selection.
The rule is a project extension inspired by score-based triage in Thomas et al., Supporting Human Raters. The paper does not prescribe this all-segments rule or these bin edges. The edges were designed after earlier test diagnostics were inspected.
| Label | Cumulative priority | Segment priority | Human review |
|---|---|---|---|
| toxic | 0.8474 | 0.9800 | 0.6575 |
| severe_toxic | disabled | disabled | disabled |
| obscene | 0.9258 | 0.9950 | 0.6575 |
| threat | disabled | disabled | disabled |
| insult | 0.9904 | 0.9950 | 0.9179 |
| identity_hate | disabled | disabled | disabled |
Disabled means that no label-specific threshold qualified; other labels or policy rules can still admit a comment. The table shows population thresholds. Subgroup thresholds are stored in the source record.
Band precision and coverage
| Rows | Priority rule | Priority count | Coverage of all rows | Priority precision |
|---|---|---|---|---|
| Selection, in sample | Cumulative | 1,628 | 5.10% | 97.24% |
| Selection, in sample | Segment | 1,060 | 3.32% | 99.53% |
| Selection, cross-fitted | Segment | 1,032 | 3.23% | 99.52% |
| Scored test, reused | Cumulative | 4,624 | 7.23% | 76.30% |
| Scored test, reused | Segment | 2,823 | 4.41% | 87.42% |
Precision here means that at least one original Jigsaw label is positive, after rules and subgroup thresholds. It is not an independent measure of whether a comment deserves review. Cumulative and segment test arms use the same saved model scores and selection data; neither fits thresholds on test labels.
The 5-fold development diagnostic refits thresholds on the other folds before routing each held-out fold. It estimates sensitivity to threshold selection within the same distribution; it is not a formal bound on optimism or evidence of generalization to a new distribution.
Where high-risk comments go
| Test arm | Band | Comments | High-risk comments | High-risk share of band |
|---|---|---|---|---|
| Cumulative | priority_review | 4,624 | 837 | 18.10% |
| Cumulative | human_review | 1,688 | 79 | 4.68% |
| Cumulative | allow | 57,666 | 164 | 0.28% |
| Segment | priority_review | 2,823 | 714 | 25.29% |
| Segment | human_review | 3,489 | 202 | 5.79% |
| Segment | allow | 57,666 | 164 | 0.28% |
High risk uses the saved policy's true-label harm proxy and the simulation threshold
harm proxy >= 5.0. A larger high-risk share in a smaller band can coexist
with fewer high-risk comments in that band. 1,801 test comments change bands
between these two rules. The two rules admit exactly the same test comments. Since severity ignores the review band, changing only these bands cannot change its queue results.
Queue results
These rates are admitted review jobs per hour, not all incoming platform comments. Capacity is 120 jobs/hour with 4 reviewers, 2.0 minutes per job and an 8-hour shift. Cells show mean ± standard deviation across 5 paired simulation seeds.
| Review arrivals/h | Ordering | High-risk completed | Harm proxy/reviewer-hour | High-risk wait p90, min | High-risk unfinished |
|---|---|---|---|---|---|
| 108 | severity | 122.8 ± 13.1 | 50.65 ± 3.07 | 1.00 ± 0.10 | 0.2 ± 0.4 |
| 108 | priority | 122.8 ± 13.1 | 50.66 ± 3.09 | 1.06 ± 0.09 | 0.2 ± 0.4 |
| 108 | prob | 122.8 ± 13.1 | 50.66 ± 3.09 | 1.19 ± 0.11 | 0.2 ± 0.4 |
| 108 | fifo | 122.6 ± 13.2 | 50.62 ± 3.09 | 5.95 ± 2.73 | 0.4 ± 0.5 |
| 180 | severity | 184.6 ± 10.0 | 71.04 ± 2.09 | 1.64 ± 0.28 | 16.2 ± 4.5 |
| 180 | priority | 184.2 ± 10.1 | 70.90 ± 2.14 | 2.01 ± 0.29 | 16.6 ± 4.4 |
| 180 | prob | 176.2 ± 10.9 | 68.49 ± 2.56 | 2.29 ± 0.66 | 24.6 ± 4.1 |
| 180 | fifo | 134.0 ± 9.1 | 55.76 ± 2.14 | 140.04 ± 8.54 | 66.8 ± 9.6 |
At 180 review arrivals/hour, the paired mean difference, severity minus band-first priority, is 0.40 high-risk completions and 0.144 harm proxy per reviewer-hour. These are the current bands; the earlier v2 comparison is a different experiment.
Wait p90 includes jobs that started, including reviews still in progress. Never-started jobs are excluded from waiting-time quantiles and remain in the unfinished count. These seeds characterize simulated arrivals on this fixed corpus; they do not measure uncertainty over new corpora. Harm is a label-weighted proxy, not measured real-world harm.
Evidence and limits
Recorded GitHub real-evaluation run: success. The record includes the workflow jobs and artifact hashes. A successful run does not mean every diagnostic met a precision target or that fairness was established. Per-band test precision gates remain unset.
For priority-review false discoveries, the identity-term subgroup ratio has 95% interval [0.720, 1.294], against a limit of 1.250. Its recorded-metric status is inconclusive. FDR concerns false discoveries among flagged comments; it does not replace the separately recorded FPR on clean comments.
The public test set has been inspected repeatedly. This is a benchmark rerun and a descriptive comparison, not a fresh holdout. The segment rule concentrates the priority band; it does not establish an improvement in model recall or actual human review value.
Two clean retrains with the same data, split, configuration and recorded dependency versions differed in two admitted comment IDs. Those membership changes also change which comments the queue's random indices select. Compare orderings within this saved run; differences between retrains cannot be attributed to the ordering alone. The retraining check found different vocabulary selections at the feature cap; the exact platform cause remains unresolved.
Keep the word model and human confirmation. Jev evaluation remains unattempted following the reported registration limit. An independent human evaluation of review worthiness remains deferred and not started because of annotation workload. Original Jigsaw labels are retained.
Reproduce this step
# Download and collect the recorded CI run
gh run download 35857528118 --name evaluation-210d9857220b23d175d4aff5a54ab221ad86280b --dir reports/ci-35857528118
gh run view 35857528118 --json headSha,status,conclusion,url,jobs >reports/ci-35857528118-ci.json
PYTHONPATH=. python record/collect_priority.py reports/ci-35857528118 --ci-metadata reports/ci-35857528118-ci.json
python record/render.py
# Development diagnostics without scoring test rows
python scripts/run_pipeline.py --config configs/dev.yaml
# Rerun the benchmark protocol from the recorded code revision
python scripts/run_pipeline.py --config configs/baseline.yaml
REVIEW_ROUTER_EVAL_REPORT=reports/<run>/report.json pytest -q tests/test_gate.py
python scripts/render_results.py reports/<run>
Run 20260923T115815829404Z-human-review-baseline, commit 210d9857220b23d175d4aff5a54ab221ad86280b,
clean working tree.
The Step 6 source record contains policy, thresholds, development diagnostics,
the matched-score test comparison and CI provenance. Steps 1 to 5 retain their saved results.
Collected artifact SHA-256 hashes
manifest.json:d468260855d625fbda8048bffaa59c892394e2fe38b96b4d1009c7e6eb8dfc63report.json:82d529e1f20efa2dee8426175e13191f6d7b2f2d80a97e31d821d935be190393policy.yaml:3df259614b78297f94c1465c767c8386ef9ca13ea43242e2f72c7c47ded27b2econfig.yaml:801efb58457b838a7316359cc89b91670d762379269720fe6d548611f137162dpredictions.csv:438029e7d44f3b333240b3c536e97d5ecdb1349009851abf85e7c18a3e111f09simulation_jobs.csv:4516e2fab3f04ec0366dc496771246ed0163532fcca5560989c08094b1ba6349