structured measurement schemas (Asirvatham, Mokski, and Shleifer 2026), iterative quality-checking workflows (Zhang and Abernethy 2025), or the kind of prompt-robustness engineering motivated by specification-search concerns (Asher et al. 2026)—should improve further.
Of course we want to look at these carefully before we praise them. I'm not super familiar with what each of these things are. And I'm not sure that I would state it so strongly that it will necessarily improve on this. There may be countervailing constraints and limitations. ... Taking a look at the AMS abstract, I don't quite see that this is the same sort of thing we're trying to do.