Reassessing Choice Probability: What 59 Macaque Studies Tell Us About Decision-Related Activity in Visual Cortex

  1. Department of Brain and Cognitive Sciences, University of Rochester, Rochester, United States
  2. Center for Visual Science, University of Rochester, Rochester, United States
  3. Program in Neuroscience, Harvard Medical School, Boston, United States
  4. Department of Neurobiology, Harvard Medical School, Boston, United States
  5. Department of Neuroscience, School of Medicine and Public Health, University of Wisconsin–Madison, Madison, United States
  6. Graduate School of Frontier Biosciences, The University of Osaka, Osaka, Japan
  7. Center for Information and Neural Networks (CiNet), National Institute of Information and Communications Technology, Osaka, Japan
  8. Department of Neuroscience, University of Pennsylvania, Philadelphia, United States
  9. Center for Perceptual Systems, The University of Texas at Austin, Austin, United States
  10. Fuster Laboratory for Cognitive Neuroscience, Departments of Psychiatry and Biobehavioral Sciences & Ophthalmology, UCLA, Los Angeles, United States
  11. Institute of Biology, Otto-von-Guericke-University Magdeburg, Magdeburg, Germany
  12. Leibniz-Institute for Neurobiology (LIN), Magdeburg, Germany
  13. Department of Physiology, Anatomy and Genetics, University of Oxford, Oxford, United Kingdom
  14. Department of Marketing, Wharton Neuroscience Initiative, Wharton School, University of Pennsylvania, Philadelphia, United States
  15. Department of Computer Science, Rochester Institute of Technology, Rochester, United States
  16. Center for Perceptual Systems, Departments of Neuroscience & Psychology, The University of Texas at Austin, Austin, United States
  17. Laboratory of Sensorimotor Research, National Eye Institute, National Institutes of Health, Bethesda, United States
  18. Department of Neurology and Neurosurgery, Montreal Neurological Institute, McGill University, Montreal, Canada
  19. School of Cognitive Sciences, Institute for Research in Fundamental Sciences (IPM), Tehran, Iran
  20. Department of Integrative Physiology, Graduate School of Medicine, University of Yamanashi, Yamanashi, Japan
  21. Center for Neuroscience, New York University, New York, United States
  22. Gonda Multidisciplinary Brain Research Center, Bar-Ilan University, Ramat Gan, Israel
  23. Department of Computer Science, University of Rochester, Rochester, United States
  24. Department of Physics and Astronomy, University of Rochester, Rochester, United States

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Timothy Hanks
    University of California, Davis, Davis, United States of America
  • Senior Editor
    Michael Frank
    Brown University, Providence, United States of America

Reviewer #1 (Public review):

This meta-analysis addresses long-standing questions about the reliability and interpretation of choice probability in macaque visual areas, and provides some important findings (e.g., the cross-study consistency of the CP-sensitivity relationship, V1 distinctiveness). However, the evidence for several claims is incomplete: the analysis does not consider the statistical dependence of data points from the same studies and monkeys, and both the bistable-stimulus effect and the stimulus-duration effect rely on interpretive assumptions.

Strengths:

The paper's transparency about its own limitations is a genuine strength. Several sections of the paper and the supplement report null results (task exposure, lapse rate, eccentricity) rather than omitting them. This kind of self-scrutiny is uncommon in meta-analyses and substantially increases confidence in the parts of the analysis that do hold up.

Weaknesses:

(1) No mixed/hierarchical statistical models for nested data. The paper considers 150 data points from 59 studies and treats them as independent samples, though many share monkeys and brain areas. This reduces the confidence in the reported p-values. A standard way of dealing with this would be to use mixed-effect models with random intercepts rather than OLS.

(2) Evidence for one of the main findings in the abstract ("First, CPs were higher in tasks involving bistable percepts, reinforcing the link between CP magnitude and subjective perception.") is weak. This effect relies entirely on studies using bistable rotating cylinder stimuli performed in a single lab (lines 702 - 711). I would suggest making this more explicit in the abstract / discussion and in Figure 8b,c.

(3) The interpretation of main drivers of CP is unclear. The discussion summarizes the 4 main drivers of CP as "four systematic drivers of this variability: neuronal sensi758tivity, brain area, stimulus duration, and task type." The independent contribution of task type is however, questionable. In line 623, it is stated that the difference between coarse and fine discrimination can be entirely explained by the difference in sensitivity (explained possibly by differences in optimizing the stimuli). Again, detection tasks (line 658) show a trend for higher CP because most studies used tailored stimuli from single-recording experiments. Bistable task: see point 2. Thus, all "task effects" can be attributed to confounds, and the claim of the "four drivers of CP" should be revised.

(4) The paper could be improved by a Discussion that synthesizes the results in a concise manner. Now it seems more like a re-iteration of the results. Overall, I appreciate that the paper is thorough and discusses many of the caveats. However, those are somewhat buried in the long subsections, and I fear that the quick reader may walk away with a stronger impression of "four robust independent drivers" than the text, read carefully, actually supports.

(5) Datapoints are not weighted according to their standard error (common practice in meta-analysis is inverse-variance weighting). The concern is that underpowered studies with high variance (e.g., due to a low number of recorded neurons) have the same impact as well-powered studies, and this may change some of the estimates. For example, Supplementary Figure 4 shows that mean CP values decrease with statistical power of the study, consistent with the concern. If SEMs are not available, could the authors show the robustness of the results by weighting by sample size as a partial check?

(6) Interpretation of feed-forward vs. feedback origin of CP [Disclosure: I am an author of Wimmer et al. 2015.]. This paper presents a mechanistic network model of area MT and a decision area that decomposes CP into two components with distinct time courses, and, directly relevant to Section 2.6, shows how a combination of an early feedforward and a late feedback component can produce a roughly time-invariant (flat) CP. This is a specific, quantitative instance of the "sustained plateau" pattern the authors themselves note is inconsistent across studies (lines 480-486) but don't develop further. Engaging with this model in Section 2.1/2.6 would let the authors contrast their duration-effect interpretation against an explicit dynamical model rather than the generic feedforward/feedback dichotomy in Figure 7a.

(7) Reaction-time experiments. I am worried that differences in CP in RT vs. fixed duration tasks (Supplementary Figure 12) could have an influence on the main regression analysis (because RT experiments are mostly from detection tasks, and because RT experiments presumably include less of a post-decision period). Could this factor be included in the main analysis?

Reviewer #2 (Public review):

Summary:

This manuscript by Pletenev et al. provides a meta-analysis of 59 published studies on decision-related activity in the macaque visual cortex. The work does not contain original research material, but by conducting extensive analyses and compilations of previously published data, it provides several new insights not already conveyed in recent reviews on this topic.

Strengths:

The work is scholarly and helps organize and integrate a broad set of findings. A difficulty in such an undertaking is making sure that the original studies are accurately characterized and tabulated. The lead authors should be commended for the rigorous approach they've taken; the inclusion of many of the authors of the original studies as co-authors provides additional reassurance that trends visible across studies reflect an accurate quantification of what each study has shown.

Weaknesses:

My sole scientific concern is on the 'duration' section (lines 487 to 536). Specifically, the authors relate the dependence of choice probability (CP) on time to two competing (but not mutually exclusive) views of how CPs arise-the 'feedforward' vs 'feedback' views. The basis for the predictions in this section was not clear to me. In particular, it was not obvious that the predictions fully considered all the relevant factors. For instance, did the feedforward predictions consider how the response covariance depends on duration and how this would affect CP values (Equation 1)? (See for example Figure 4 of Kang and Maunsell, 2012, JNP.)

I would suggest either explaining the basis of the predictions much more carefully or moving the predictions to a supplementary section where they can be presented in more detail (i.e., more convincingly). Alternatively, since the section is inconclusive in the end (there is not strong evidence in favor of FF or FB), the authors could mention the theoretical predictions in passing only, i.e., much more briefly, just to make the reader aware that the different theories can provide predictions about the duration dependence.

Reviewer #3 (Public review):

Summary:

This study presents a comprehensive meta-analysis of choice probability (CP), a classic metric of the relationship between single-neuron responses and an animal's perceptual judgment. The authors compiled data from 59 macaque electrophysiology studies and identified several factors that consistently influence CP magnitude, including neuronal sensitivity, brain area, stimulus duration, and task type. These results provide evidence that helps settle several long-standing debates about how CP should be interpreted.

Strengths:

This work is a rare example of meta-analysis in macaque electrophysiology, focusing on an important and long-debated metric, choice probability (CP), in the study of sensory and decision-making mechanisms. CP has been measured across many studies, but its interpretation remains contentious because it depends on numerous task and recording factors in addition to sensory and decision-making mechanisms. Individual macaque studies also typically include few subjects, and CP effects are generally small, which prevents any single study from drawing strong conclusions about general patterns. The authors identified an ideal use case for meta-analysis and combined fragmented findings from individual primate studies into a coherent picture of which factors matter most for CP and which do not. This work can also serve as a guide for future meta-analyses of macaque electrophysiology data.

Weaknesses:

While the survey and discussion of CP's interpretation are comprehensive, the paper would benefit from a clearer conclusion on why measuring CP remains important and what future directions could make CP more useful for revealing sensory and decision-making mechanisms.

Author response:

We thank the Editors and Reviewers for their encouraging evaluation and constructive feedback. We are glad that they recognized the value of this systematic synthesis in reconciling disparate findings across many studies on an important question, and in providing new insights that help address long-standing debates around the interpretation of CP values. They also acknowledged our rigorous data curation involving many original study authors, and our transparency in reporting limitations and null results.

To address their constructive recommendations, we will provide additional hierarchical regression analyses where possible, state some limitations more explicitly, and condense the Discussion section to improve focus and readability.

Regarding the hierarchical nature of the data, we distinguish three potential hierarchical levels: studies, monkeys, and neuronal samples. Because individual monkeys contribute roughly one observation per study and cannot be tracked across publications, animal-level random effects are statistically unidentifiable. We will address the rare cases where identical neuronal pools were evaluated across multiple task conditions by providing a sensitivity analysis restricted to one data point per unique neuronal sample. At the study level (median 2, range 1–7 observations per study), we will present linear mixed-effects models with random study intercepts.

On the bistability findings, we agree with Reviewer 1's concern about the limited number of studies. We noted in the Results and Discussion that this effect currently relies on rotating-cylinder paradigms and emphasized the need for CP to be quantified with other forms of bistable stimuli. We will make this limitation explicit in the Abstract and Figure 8 caption. We will revise the Discussion to emphasize the three primary drivers (neuronal sensitivity, brain area, and stimulus duration) and treat the bistable stimulus effect separately as a distinct finding.

Regarding Reviewer 1's concern about reaction-time experiments: because reaction-time (RT) paradigms are heavily confounded with detection tasks in the existing literature (89% of detection tasks are RT tasks, and 67% of RT tasks are detection ones), including RT as a separate factor introduces near-complete collinearity. We will make this constraint and the inability to statistically disentangle them explicit in the main text.

Addressing Reviewer 2's concern regarding the CP–duration predictions, we agree that they depend on specific modeling assumptions. While our predictions—for the feedforward framework in particular—reflect standard models from the literature, we will explicitly acknowledge that alternative feedforward assumptions—such as duration-dependent response covariance—could alter the expected relationship.

In response to Reviewer 3’s comments on the rationale and future utility of measuring CP, we will revise the Discussion to emphasize that while the interpretation of CP has evolved from feedforward readout to include feedback mechanisms, our results confirm that CP remains a robust neural correlate of subjective perception. Although the exact mechanisms linking CP to perception remain unresolved, this ambiguity does not justify abandoning the metric; rather, CP remains an indispensable tool, provided it is supplemented with additional analyses as outlined in our recommendations.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation