6 Matching Annotations
  1. Sep 2026
    1. quick votes map from Strong No through Strong Yes to 10, 30, 50, 70, and 90

      My sense is that it might be clearer just to show the numbers for quick votes (rather than the signs that may be read differently by people who don't read instructions carefully) (?)

    2. Calibration and stability answer different questions. Calibration asks whether AI ratings line up with human judgments. Stability asks whether the same papers keep similar ratings when the model is run again or a substantively equivalent instruction is reworded. A system can be stable but wrong, so stability is a reproducibility diagnostic rather than an accuracy claim. Our pilot covers 24 selected papers with 192 repeated ratings and remains exploratory; it is not a full-dashboard reliability estimate.

      I might misunderstand this, but should there be an assessment of prompt and run stability (rather than just an explanation of what it means)?