I found your presentation below on this rather convincing. This also comports with what I've heard from other EAs (although perhaps the same circle of conversation). We need better evidence on which of these AW improving technologies actually reduce animal suffering. I'm in some discussions about possibly building and funding an evaluation service for specific tools and approaches (maybe something between The Unjournal and a fast-review journal, also inspired by Rapid Reviews Infectious Diseases).
Very small sample sizes do not always mean lack of inference. For instance, in very predictable contexts without a lot of noise, like, let's say, Newtonian physics, even a few data points could help us narrow our beliefs substantially. I like Richard McElreath's example about how you can substantially update on the share of a planet that is water by simply, even in the first few random samples, choosing a single point on a planet's sphere.
More intuitive -- if I ask 4 people to taste a drink and they all wince deeply in pain and disgust, I'm going to be highly confident it tastes bad. If all 4 smile and praise it, I'll be fairly confident that it's at least tolerable.
But I don't know that that is the case here. There might in fact be a lot of uncertainty and heterogeneity. What I wonder is whether the sample sizes observing the behavior and bioindicators of these fish are very expensive, or whether it could easily be scaled up with just a small amount of money. as a non-biologist, it seems intuitive that it should be cheap, but I might be missing something