Step 6Policy v3

Severity ordering and confidence-segment review bands

A real-data evaluation of the word model with severity as the primary queue ordering, development-only diagnostics, and a stricter rule for assigning the priority label. Every moderation action still requires human confirmation.

Recorded conclusion. On the same scored test rows, segment selection changes priority precision from 76.30% to 87.42%, while the priority band shrinks from 4,624 to 2,823 comments. The current review pool contains 6,312 comments and captures 916 of 1,080 high-risk comments. The two rules admit exactly the same test comments. Since severity ignores the review band, changing only these bands cannot change its queue results.

What changed

The primary ordering is severity: for each comment, take the maximum over labels of calibrated probability × that label's severity weight, then review the highest-scoring comment first. The priority_review label no longer guarantees earlier service. Band-first priority, probability and FIFO remain comparisons on the same admitted comments and arrival streams.

For each label, the priority threshold is the lowest declared score edge for which every segment above it meets the 95.00% empirical agreement target with at least 30 rows. These are score ranges, not statistical confidence intervals. Human-review thresholds and identity-term subgroup thresholds retain cumulative precision selection.

The rule is a project extension inspired by score-based triage in Thomas et al., Supporting Human Raters. The paper does not prescribe this all-segments rule or these bin edges. The edges were designed after earlier test diagnostics were inspected.

LabelCumulative prioritySegment priorityHuman review
toxic0.84740.98000.6575
severe_toxicdisableddisableddisabled
obscene0.92580.99500.6575
threatdisableddisableddisabled
insult0.99040.99500.9179
identity_hatedisableddisableddisabled

Disabled means that no label-specific threshold qualified; other labels or policy rules can still admit a comment. The table shows population thresholds. Subgroup thresholds are stored in the source record.

Band precision and coverage

RowsPriority rulePriority countCoverage of all rowsPriority precision
Selection, in sampleCumulative1,6285.10%97.24%
Selection, in sampleSegment1,0603.32%99.53%
Selection, cross-fittedSegment1,0323.23%99.52%
Scored test, reusedCumulative4,6247.23%76.30%
Scored test, reusedSegment2,8234.41%87.42%

Precision here means that at least one original Jigsaw label is positive, after rules and subgroup thresholds. It is not an independent measure of whether a comment deserves review. Cumulative and segment test arms use the same saved model scores and selection data; neither fits thresholds on test labels.

The 5-fold development diagnostic refits thresholds on the other folds before routing each held-out fold. It estimates sensitivity to threshold selection within the same distribution; it is not a formal bound on optimism or evidence of generalization to a new distribution.

Where high-risk comments go

Test armBandCommentsHigh-risk commentsHigh-risk share of band
Cumulativepriority_review4,62483718.10%
Cumulativehuman_review1,688794.68%
Cumulativeallow57,6661640.28%
Segmentpriority_review2,82371425.29%
Segmenthuman_review3,4892025.79%
Segmentallow57,6661640.28%

High risk uses the saved policy's true-label harm proxy and the simulation threshold harm proxy >= 5.0. A larger high-risk share in a smaller band can coexist with fewer high-risk comments in that band. 1,801 test comments change bands between these two rules. The two rules admit exactly the same test comments. Since severity ignores the review band, changing only these bands cannot change its queue results.

Queue results

These rates are admitted review jobs per hour, not all incoming platform comments. Capacity is 120 jobs/hour with 4 reviewers, 2.0 minutes per job and an 8-hour shift. Cells show mean ± standard deviation across 5 paired simulation seeds.

Review arrivals/hOrderingHigh-risk completedHarm proxy/reviewer-hourHigh-risk wait p90, minHigh-risk unfinished
108severity122.8 ± 13.150.65 ± 3.071.00 ± 0.100.2 ± 0.4
108priority122.8 ± 13.150.66 ± 3.091.06 ± 0.090.2 ± 0.4
108prob122.8 ± 13.150.66 ± 3.091.19 ± 0.110.2 ± 0.4
108fifo122.6 ± 13.250.62 ± 3.095.95 ± 2.730.4 ± 0.5
180severity184.6 ± 10.071.04 ± 2.091.64 ± 0.2816.2 ± 4.5
180priority184.2 ± 10.170.90 ± 2.142.01 ± 0.2916.6 ± 4.4
180prob176.2 ± 10.968.49 ± 2.562.29 ± 0.6624.6 ± 4.1
180fifo134.0 ± 9.155.76 ± 2.14140.04 ± 8.5466.8 ± 9.6

At 180 review arrivals/hour, the paired mean difference, severity minus band-first priority, is 0.40 high-risk completions and 0.144 harm proxy per reviewer-hour. These are the current bands; the earlier v2 comparison is a different experiment.

Wait p90 includes jobs that started, including reviews still in progress. Never-started jobs are excluded from waiting-time quantiles and remain in the unfinished count. These seeds characterize simulated arrivals on this fixed corpus; they do not measure uncertainty over new corpora. Harm is a label-weighted proxy, not measured real-world harm.

Evidence and limits

Recorded GitHub real-evaluation run: success. The record includes the workflow jobs and artifact hashes. A successful run does not mean every diagnostic met a precision target or that fairness was established. Per-band test precision gates remain unset.

For priority-review false discoveries, the identity-term subgroup ratio has 95% interval [0.720, 1.294], against a limit of 1.250. Its recorded-metric status is inconclusive. FDR concerns false discoveries among flagged comments; it does not replace the separately recorded FPR on clean comments.

The public test set has been inspected repeatedly. This is a benchmark rerun and a descriptive comparison, not a fresh holdout. The segment rule concentrates the priority band; it does not establish an improvement in model recall or actual human review value.

Two clean retrains with the same data, split, configuration and recorded dependency versions differed in two admitted comment IDs. Those membership changes also change which comments the queue's random indices select. Compare orderings within this saved run; differences between retrains cannot be attributed to the ordering alone. The retraining check found different vocabulary selections at the feature cap; the exact platform cause remains unresolved.

Keep the word model and human confirmation. Jev evaluation remains unattempted following the reported registration limit. An independent human evaluation of review worthiness remains deferred and not started because of annotation workload. Original Jigsaw labels are retained.

Reproduce this step

# Download and collect the recorded CI run
gh run download 35857528118 --name evaluation-210d9857220b23d175d4aff5a54ab221ad86280b --dir reports/ci-35857528118
gh run view 35857528118 --json headSha,status,conclusion,url,jobs >reports/ci-35857528118-ci.json
PYTHONPATH=. python record/collect_priority.py reports/ci-35857528118 --ci-metadata reports/ci-35857528118-ci.json
python record/render.py
# Development diagnostics without scoring test rows
python scripts/run_pipeline.py --config configs/dev.yaml
# Rerun the benchmark protocol from the recorded code revision
python scripts/run_pipeline.py --config configs/baseline.yaml
REVIEW_ROUTER_EVAL_REPORT=reports/<run>/report.json pytest -q tests/test_gate.py
python scripts/render_results.py reports/<run>

Run 20260923T115815829404Z-human-review-baseline, commit 210d9857220b23d175d4aff5a54ab221ad86280b, clean working tree. The Step 6 source record contains policy, thresholds, development diagnostics, the matched-score test comparison and CI provenance. Steps 1 to 5 retain their saved results.

Collected artifact SHA-256 hashes
  • manifest.json: d468260855d625fbda8048bffaa59c892394e2fe38b96b4d1009c7e6eb8dfc63
  • report.json: 82d529e1f20efa2dee8426175e13191f6d7b2f2d80a97e31d821d935be190393
  • policy.yaml: 3df259614b78297f94c1465c767c8386ef9ca13ea43242e2f72c7c47ded27b2e
  • config.yaml: 801efb58457b838a7316359cc89b91670d762379269720fe6d548611f137162d
  • predictions.csv: 438029e7d44f3b333240b3c536e97d5ecdb1349009851abf85e7c18a3e111f09
  • simulation_jobs.csv: 4516e2fab3f04ec0366dc496771246ed0163532fcca5560989c08094b1ba6349