Experiment recordv0.1.0

Same model and reviewers. Different review priorities.

review-router is a reproducible benchmark for human-review routing. It compares admission policies and queue ordering under fixed reviewer capacity.

v0.1.0 release and downloads · GitHub repository

Read the article: Same model, different review queue: which risks get reviewed first?

A walkthrough of the four queue orderings, the gains and displaced work under fixed capacity, and why waiting-time statistics need to be read alongside counts of jobs that never started.

What the queue changes, and what it costs

At 180 admitted review jobs per hour, the default ordering completes 53.75 more high-risk reviews per shift than FIFO and 53.75 fewer other reviews. These are means across 20 paired simulation seeds, with four reviewers, two minutes per job and an eight-hour shift.

Completed reviews per shiftFIFODefault severity ordering
High-risk reviews135.60189.35
Other reviews819.65765.90
Total reviews955.25955.25

Ordering reallocates fixed capacity. High risk means a positive severe toxicity, threat or identity hate label; other reviews include all remaining queued jobs. The bounded weight sensitivity experiment uses 20 seeds and retains the declared default weights. Step 6 below preserves the earlier five-seed baseline. Neither experiment is an independent evaluation on new comments or a measurement of production review value.

The released method

The scorer uses frozen word TF-IDF features, six logistic classifiers and Platt calibration. The router assigns comments to allow, ordinary review or priority review. Both review bands enter one queue, ordered by the largest calibrated label probability times its severity weight. The priority label alone does not determine queue position. Every moderation action requires human confirmation.

The saved benchmark routes 6,312 of 63,978 test comments to review, or 9.87%. Jigsaw labels remain proxies for review value. The method and its limits are fixed in the release notes; the scorer freeze record explains what stays fixed.

How the method developed

Steps 1 and 2 document the retired automatic-action experiment and its 99% precision target. Later steps require human confirmation. Historical pages keep their original results; Steps 1 to 5 are in Chinese, and Step 6 is in English.

  1. Step 1: baseline Chinese archive
    Build a reproducible pipeline for scoring, thresholds, rules and queue simulation.
  2. Step 2: first results Chinese archive
    Inspect the first results, retire R103 and diagnose the unmet automatic-action target.
  3. Step 3: human confirmation Chinese archive
    Require human confirmation and count both review bands against reviewer capacity.
  4. Step 4: feature comparison Chinese archive
    Audit false positives and compare word features with word plus character features.
  5. Step 5: Jev pilot Chinese archive
    Prepare a paired pilot. Registration was unavailable; no live Jev evaluation was attempted.
  6. Step 6: severity and confidence segments English
    Compare severity ordering and confidence segments in the archived five-seed baseline.
  7. Freeze the scorer
    Hold the vocabulary, classifier and calibration parameters fixed for router comparisons.
  8. Check weight sensitivity and release v0.1.0
    Test five declared weight vectors, report the capacity tradeoff and retain the default.

Jev remains untested because the project owner reported that registration was at capacity. The prepared pilot does not establish its accuracy, latency or benefit.

Independent human evaluation remains deferred

A comment may deserve review because it is ambiguous or needs context, even if its original label is negative. An independent assessment could record review worthiness, urgency, missing context, final decisions and handling time. This work has not started because of the annotation workload. Original Jigsaw labels remain unchanged.

The label-quality research note, in Chinese, explains why labels deserve scrutiny and why that alone does not prove a particular label is wrong.

Reproduce the release

Start with the versioned model and evaluation archives and the setup and restoration protocol. The archives preserve model bytes, predictions, thresholds and configuration without relying on expiring CI artifacts.

Regenerate the saved record pages
python -m pip install -e ".[ml,analysis]"
python record/diagnose.py --render-only
python record/render.py

These commands render saved records and figures. They do not train the model or rerun the benchmark.

Archived Step 6 provenance before the scorer freeze: run 20260923T115815829404Z-human-review-baseline, commit 210d9857220b23d175d4aff5a54ab221ad86280b. Its policy and results are in priority_run.json. Earlier runs are in runs.json; the release sensitivity protocol and results are in severity-sensitivity.json.

Project code is licensed under MIT. Third-party data and dependencies retain their own terms.