Reviewer #3 (Public Review):
The authors describe a method for fitting a simple, separable function of contrast and cone excitation to a set of fMRI data generated from large, unstructured chromatic flicker stimuli that drive the L- and M- cone photoreceptors across a range of amplitudes and ratios. The function is of the form of a scaled ellipse – hereafter referred to as a 'Quadratic Color Model' (QCM). The QCM fits 6 parameters (ellipse orientation, ellipse elongation, and 4 parameters from a non-linear, saturating (Naka-Rushton) contrast response curve. The QCM fits the dataset well and the authors compare it (favorably) to a 40-parameter GLM that fits each separate combination of chromatic direction and contrast separately.
The authors note three things that 'did not have to be true' (and which are therefore interesting):
1) The data are well-fit by a separable ellipse+contrast transducer - consistent with the idea that the underlying neuronal computations that process these stimuli combine relatively independent L-M and L+M contrast.
2) The short axis of the QCM tends to align with the L-M cone contrast directing (indicating that this direction is one of maximum sensitivity and the L+M direction (long axis) is least sensitive. This finding is qualitatively consistent with psychophysical measurements of chromatic sensitivity.
3) Fit parameters do not change much across the cortical surface – and in particular they are relatively constant with respect to eccentricity.
This is a technically solid paper – the data processing pipeline is meticulous, stimuli are tightly-calibrated (the ability to apply cone-isolating stimuli to fovea and periphery simultaneously is an impressive application of the 56-primary stimulus generator) and the authors have been careful to measure their stimuli before and after each experimental session. I have a few technical questions but I am completely satisfied that the authors are measuring what they think they are measuring.
The analysis, similarly, is exemplary in many ways. Robust fitting procedures are used and model performance and generalizablility are evaluated with a leave-run-out and leave-session-out cross validation procedures. Bootstrapped confidence intervals are generated for all fits and analysis code is available online.
The paper is also useful: it summarises a lot of (similar) previous findings in the fMRI color literature going back to the late 90s and points out that they can, in general, be represented with far fewer parameters than conditions. My main concerns are:
1) Underlying mechanisms: The QCM is a convenient parameterization of low spatial-frequency, high temporal-frequency L-M responses. It will be a useful tool for future color vision researchers but I do not feel that I am learning very much that is new about human color vision. The choice to fit an ellipse to these data must have been motivated at least in part by inspection. It works in this case (possibly because of the particular combination of spatial and temporal frequencies that are probed) but it is not clear that this is a generic parametric model of human color responses in V1. Even very early fMRI data from stimuli with non-zero spatial frequency (for example, Engel, Zhang and Wandell '97) show response envelopes that are ellipse-like but which might well also have additional 'orthogonal' lobes or other oddities at some temporal frequencies.
2) Model comparison: The 40-parameter GLM model provides a 'best possible' linear fit and gives a sense of the noisiness of the data but it feels a little like a strawman. It is possible to reduce the dimensionality of the fit significantly with the QCM but was it ever really plausible that the visual system would generate separate, independent responses for each combination of color direction and contrast? I suspect that given the fact that the response data are not saturating, it would be possible to replace the Naka-Rushton part of the model with a simple power function, reducing the parameter space even further. It would be more interesting to use the data to compare actual models of color processing in retina/V1 and, potentially, beyond V1.
3) Link to perception. As the authors note, there is a rich history of psychophysics in this domain. The stimuli they choose are also, I think, well suited to modelling in the sense that they are likely to drive a very limited class of chromatic cells in V1 (those with almost no spatial frequency tuning). It is a shame therefore that no corresponding psychophysical data are presented to link physiology to perception. The issue is particularly acute because the stimulus differs from those typically used in more recent psychophysical experiments: it flickers relatively quickly and it has no spatial structure. It may, however, be more similar to the types of stimuli used prior to the advent of color CRTs : Maxwellian view systems that presented a single spot of light.