I wonder if we should normalize this in a few ways, at least as an alternative measure.
I suspect the AI's distribution of ratings may have different than the human distribution of ratings overall and, the "bias" may also differ by category.
Actually, that might be something to do first -- compare the distributions of (middle -- later more sophisticated) ratings for humans and for LLMs in an overall sense.
One possible normalization would be to state these as percentiles relative to the other stated percentiles within that group (humans, LLMs), or even within categories of paper/field/cause area (I suspect there's some major difference between the more applied and niche-EA work and the standard academic work (the latter is also probably concentrated in GH&D and environmental econ). On the other hand, the systematic differences between LLM and human ratings on average might also tell us something interesting. So I wouldn't want to only use normalized measures.
I think a more sophisticated version of this normalization just becomes a statistical (random effects?) model where you allow components of variation along several margins.
It's true the ranks thing gets at this issue to some extent, as I guess Spearman also does? But I don't think it fully captures it.