Figures and data

Conceptual framework of choice-related metrics and the rationale for a meta-analytic approach.
(a) In perceptual decision tasks, sensory neuron responses reflect both the external stimulus and the animal’s choice. Sensitivity and preference characterize the neuron-stimulus relationship: sensitivity quantifies the magnitude of response modulation, while preference indicates the stimulus category eliciting higher firing rates (ideally both assessed when choice is held constant). Choice Correlations and Choice Probability (CP) quantify the relationship between neuronal activity and the subject’s choice for a fixed stimulus. Choice correlations measure the correlation between trial-by-trial neuronal firing rates and the subject’s binary choice (coded as +1 or -1), using a fixed sign convention for choice across the entire population. CP differs in two key respects: (1) it is calculated as an ROC area, and (2) the “positive” choice class is defined by the neuron’s own stimulus preference (its influence on CP is indicated by the dashed line). CP is shaped by covariability —the trial-by-trial correlation between neuronal responses to a fixed stimulus (Equation (1)). Mean CP —the core metric of our meta-analysis— is a population-level measure indicating how strongly, on average, neurons increase their firing when the subject selects the neuron’s preferred category. (b) Conceptual illustration of the complementary strengths of this meta-analysis versus individual experimental studies. The y-axis (Amount/resolution of data) reflects both the number of data points and the ability to capture simultaneous population dynamics. The x-axis represents experimental diversity, specifically the number of different behavioral tasks and individual subjects tested. Studies based on modern recording techniques offer superior data resolution but typically rely on few tasks and few animals (high y-axis, low x-axis). Conversely, our meta-analysis aggregates mean CP across 59 studies to span diverse tasks and subjects (high x-axis, low y-axis). Thus, while limited in the amount of data, it maximizes generalizability by overcoming subject-specific variance and enabling the isolation of experimental factors that impact the choice signal in sensory neurons. Future meta-studies may overcome these data limitations by pooling raw datasets rather than relying on summary statistics (see Section 3.7).

Schematic representation of factors hypothesized to influence Choice Probability (CP).
Arrows denote putative causal relationships between experimental design choices, intermediate variables, and core factors influencing CP (see Tables 2 and 3 for details). Interaction effects between variables are not shown. Two central variables –decoder type/readout w and covariability C (enclosed by the dashed outline), linking CP to perception through Equation (1) – mediate the effect of all other factors on choice correlations. Gray boxes indicate variables that can be experimentally observed. Thick blue borders highlight variables that showed a statistically significant relationship with mean CP in our meta-analysis. For a more complete schematic incorporating additional factors (see Fig. S1).

Mean CP values decline over time, primarily due to reduced neuronal sensitivity in later studies.
(a) Relationship between mean CP and year of publication. Each point represents a study-level mean (typically per monkey, brain area, and task). Points are slightly jittered horizontally for visibility. The thick black line shows the overall regression (β = 0.0019, p < 0.001, N = 150); colored lines show within-area regressions. (b) Same as (a), but showing the partial regression after controlling for neuronal sensitivity, brain area, stimulus duration, and task type. No significant relationship remains, regardless of how we control for sensitivity (either via stimulus tailoring, black line, p = 0.68, or the inverse N/P ratio, not shown, p = 0.55, also see Section 2.4). (c) Relationship between normalized neuronal sensitivity and year of publication. Sensitivity is approximated by the inverse of the mean neurometric-to-psychometric threshold ratio (inverse N/P ratio; see Section 2.4.1). A strong decrease in sensitivity is observed, likely reflecting methodological shifts toward less stimulus tailoring to individual neurons.

Choice of sensitivity metric and sampling strategy shapes the CP–sensitivity relationship.
All panels present simulated data. (a) Toy simulation showing how neuronal informativeness metrics yield distinct relationships with CP. Left: CP vs. neurometric threshold. Right: CP vs. sensitivity (inverse threshold). Open circles are simulated neurons; large filled circles are study averages. Solid lines denote ground truth relationship. Theoretical models predict a linear CP–sensitivity relationship, but a hyper-bolic relationship with threshold. (b) Toy simulation illustrating how stimulus tailoring impacts sampled sensitivity and CP. Grey points represent the full hypothetical neuronal population. Red points denote a simulated study tailoring stimuli to individual neurons (sampling the high-sensitivity tail); blue points denote a study without tailoring (random sample with a minimum response threshold). Despite sampling differences, both studies yield similar CP–sensitivity slopes (solid black lines) that closely match the line connecting their averages (large colored circles). Insets: marginal distributions. (c) Task-aligned feedback inflates sensitivity estimates. This effect is illustrated using simulations from the hierarchical inference model of Haefner et al. (2016). Normalized sensitivity is estimated either without feedback (resembling passive viewing data, open circles), or with feedback (resembling active task data, filled circles). CP values remain identical across conditions. However, feedback elevates estimated sensitivity—especially for already-sensitive neurons (rightward shift)—causing the CP–sensitivity slope to become shallower (solid vs. dashed line).

Positive relationship between CP and sensitivity both within and across studies, with a positive offset across studies.
(a) Conceptual illustration of two edge cases. Left: the slope of the across-study mean CP–mean sensitivity relationship closely matches the average within-study slopes, suggesting a shared mechanism. Right: no relationship across studies, implying different sources of variability. Circles represent study-level means; shaded gray region – schematic 95% confidence interval of the across-study relationship; short thin lines – within-study slopes. (b) Empirical mean CP–mean sensitivity relationship across studies. Each point is a study-level mean. The x-axis shows normalized sensitivity approximated by the inverse of the mean neurometric-to-psychometric threshold ratio (inverse N/P ratio). Short colored segments passing through the points denote within-study slopes for the subset of studies where data were available (N = 42). The text insert reports the median and interquartile range (25%–75%) of these within-study slopes: 0.06 (0.01-0.12). (c) Same data as in (b), but overlaid with across-study relationships. The thick black line represents the across-study relationship with sensitivity as the sole predictor with slope of 0.05 (p = 0.003, N = 88). A multivariate model incorporating brain area, stimulus duration, and task type yielded a similar sensitivity slope of 0.04 (not shown; p = 0.02, N = 88). Thin colored lines show regressions within brain areas with at least 6 data points. Notably, the across-study slope estimate falls within the distribution of the within-study slopes.

Lowest mean CP in V1 but no consistent increase with hierarchical level across brain areas.
(a) Schematic illustration of two competing hypotheses about how CP varies across visual areas. Top: the decoder or feedback can target any sensory area, resulting in comparable CP values across areas if tasks are well matched to the recorded neurons. Bottom: the decoder or feedback accesses only the most downstream area, leading to increasing CP with hierarchical level. (b) Mean CP values across brain areas organized by hierarchical position across the ventral and dorsal streams. We assigned area V3/V3A to the dorsal stream, although the classification of its constituent parts remains a subject of ongoing debate (Felleman and Van Essen, 1991; Lyon and Kaas, 2002; Kravitz et al., 2011; Rosenberg et al., 2023). Diamonds and error bars represent medians and interquartile ranges for areas with more than five observations. Asterisks indicate significant differences between two brain areas based on Mann–Whitney tests (** p < 0.01, *** p < 0.001); p < 0.05 results are omitted as they would not survive correction for multiple comparisons. (c) Regression coefficients capturing the contribution of each brain area to mean CP, controlling for sensitivity (either inverse mean N/P ratio or tailoring technique), stimulus duration, and task type. V1 served as the baseline category and is not shown. Filled and open circles show regression coefficients using, respectively, the inverse N/P ratio and tailoring technique to control for sensitivity. Error bars indicate standard errors. Coefficients are only shown for brain areas with at least four data points. Asterisks denote significance relative to V1 (*V3/V3A (N = 8), MT (N = 74), MST/MSTd (N = 19), and CIP (N = 4). In our dataset, some areas—particularly IT, MT, and CIP, exhibited higher mean CPs than V1; however, beyond that, we found no evidence that later stages of the visual hierarchy are associated with higher CPs (Fig. 6b). Although a subset analysis of single-electrode recordings showed a similar overall pattern, it unexpectedly revealed lower mean CP values for MST compared to MT, despite MST occupying a higher level in the dorsal stream (Fig. S7). This surprising result can be explained by the lower task sensitivity in MST compared to MT (Fig. S9). Notably, V1 and MT had similar sensitivity, so the CP difference between them cannot be attributed to this factor.

Mean CP increases with stimulus duration, aligning with feedback-based model predictions.
(a) Schematic illustrating model predictions for the relationship between mean CP and stimulus duration. Feedforward models predict either no relationship (in classic integration-to-bound models) or a negative relationship (due to memory and attention constraints). Conversely, models incorporating feedback predict a positive relationship: hierarchical inference models predict this due to a larger number of samples acquired during inference, whereas models with post-decision feedback predict this due to increased post-decision time. (b) The empirical relationship between mean CP and stimulus duration across studies exhibits a positive slope in a simple linear regression with duration as the sole predictor (black line; slope = 0.019, p = 0.001). This positive trend remains robust in regressions controlling for sensitivity, brain area, and task type (not shown; inverse N/P ratio as sensitivity control: slope = 0.022, p = 0.001; tailoring as sensitivity control: slope = 0.019, p = 0.002). Data points are jittered horizontally for visibility. Thin colored lines show pairwise regressions within each brain area. Regression lines only shown for brain areas with more than 6 points.

While mean CP varies systematically with task type, only bistable stimuli showed an independent effect.
(a) Representative stimuli for coarse-discrimination, fine-discrimination, detection, and bistable stimulus tasks; in the latter, a single image evokes two distinct perceptions. (b) Mean CP values across four task categories. Diamonds and error bars represent medians and interquartile ranges. Asterisks indicate significant differences between two task types based on Mann–Whitney tests (** p < 0.01, *** p < 0.001); p < 0.05 results are omitted as they would not survive correction for multiple comparisons. While the difference in mean CP between coarse and fine-discrimination did not reach significance in the full dataset, it is significant within the subset of single-electrode studies (Fig. S13). (c) Regression coefficients estimating the effect of each task type on mean CP, controlling for sensitivity, brain area, and stimulus duration. Coarse-discrimination tasks serve as the baseline category and is not shown. Filled circles indicate coefficients from models using the inverse mean N/P ratio; open circles from models using tailoring technique. Error bars indicate standard errors. Asterisks denote statistical significance relative to coarse-discrimination tasks (* p < 0.05, ** p < 0.01). Only bistable tasks showed a consistently significant positive effect across both models. Conversely, the lower mean CP observed for fine discrimination under the tailoring control model is likely an artifact, as fine-discrimination tasks inherently exhibit smaller inverse N/P ratios than coarse-discrimination tasks even when stimulus tailoring is matched (see text).



Summary of papers used in the meta-analysis (see also full dataset).
Pts: number of data points; Pts comb.: number of data points for which the mean CP was estimated from pooled data across monkeys; Monk.: number of monkeys; Cond.: number of task conditions; N/P: whether N/P ratio was reported (if mixed, shown as “+; -”); Stim. Dur. (s): stimulus duration; Task: task type.


Explanation of terms used in Fig. 2 and Fig. S1.
3Grouped under the “Learning and Engagement” category in Fig. 2.



Annotation of the arrows in Fig. 2 and Fig. S1.
The “Sign” column indicates the expected direction of influence. A missing sign means that the direction cannot be specified, either because the relationship is non-monotonic or because one of the variables lacks a hierarchical ordering.

Linear regression results predicting mean CP using inverse mean N/P ratio, brain area, stimulus duration, and task type.

Linear regression results predicting mean CP using tailoring-based sensitivity proxy, brain area, stimulus duration, and task type.

Expanded schematic of hypothesized factors influencing Choice Probability (CP).
This supplementary diagram extends the framework presented in Fig. 2 by incorporating additional elements: (1) potential confounds and pre-motor signals, and (2) experimental factors that could not be evaluated in this meta-analysis due to data limitations (Psych. Kernel, Reward Scheme). It was proposed that at least part of the observed choice correlations may be explained by pre-motor signals during saccade planning (Laamerad et al., 2024; Zhang and Gu, 2026)(lower right part of the diagram). Methodological choices that date back to Britten et al. (1996) such as CP estimation procedures (“Grand CP”) and stimulus control techniques (“frozen noise”, eye movement control) may influence CP estimates (Nienborg and Cumming, 2009; Kang and Maunsell, 2012; Herrington et al., 2009) (lower left of the diagram). Additionally, compared to Fig. 2 the “Learning and Engagement” node has been subdivided into “Task learning,” “Perceptual learning,” and “Engagement.” Arrows denote putative causal relationships between experimental design choices, intermediate variables, and core factors driving CP (see Tables 2 and 3 for details). Gray boxes denote experimentally observable variables. Thick blue borders highlight variables that exhibited a statistically significant relationship with mean CP in our analysis.

Differences in mean CP between monkeys within the same study.
Each point shows the difference between two monkeys, computed by subtracting the lower CP value from the higher. Vertical jitter was added for visualization. The black curve is a density function

Mean CP is lower in multi-electrode studies and declines over time in single-electrode studies.
(a) Mean CP by recording technique. Multi-electrode studies report significantly lower CP values (*** p < 0.001). (b) Relationship between mean CP and year of publication, shown separately for single-electrode (filled symbols, solid regression line) and multi-electrode studies (open symbols, dashed line). A significant decline is observed even when restricting the analysis to single-electrode studies (β = −0.0015, p = 0.02).

Mean CP values decrease with statistical power of the study.
(a) Relationship between mean CP and the number of neurons contributing to the estimate, restricted to single-electrode studies to reduce confounding. A negative trend (slope = 0.0003, p = 0.02, N = 109) suggests that low-powered studies (fewer neurons) tend to report inflated CP values. (b) same as (a) but restricted to multi-electrode studies. No trend is observed (p = 0.96, N = 37) (c) Funnel plot showing standard error (y-axis) versus mean CP (x-axis). Restricted to single-electrode studies to reduce confounding. The solid vertical line marks the meta-analytic mean CP (0.52), computed as an inverse-variance-weighted average. The shaded triangle represents the 95% expected region assuming this true mean. An asymmetry is evident, with higher-error studies skewed toward elevated CP values, consistent with publication bias (Egger’s test: p = 0.01, N = 51). (d) same as (c) but restricted to multi-electrode studies. Similar asymmetry as for single-electrode studies is observed (Egger’s test: p < 0.001, N = 24).

Effect of tailoring techniques on neuronal sensitivity (a, b) and mean CP (c, d).
Left column (a, c) shows the effect of tailoring two types of stimulus parameters: task-relevant (“task”) and task-irrelevant but response-modulating (“non-task”). X-axis labels represent combinations of tailoring choices. For example, the label ‘not fit/single neuron’ indicates that the task parameter was not tailored, whereas the non-task parameter was optimized for each individual neuron. Note that in panel a, nearly all available sensitivity data come from single-electrode studies, so tailoring to population-level properties is mostly absent. Right column (b, d) shows the effect of tailoring stimulus size relative to the receptive field. Asterisks indicate significant differences between two categories based on Mann–Whitney tests (* p < 0.05, ** p < 0.01, *** p < 0.001). For a, c p < 0.05 results are omitted as they would not survive correction for multiple comparisons.

Relationship between mean CP and the within study CP–sensitivity correlation coefficient (Pearson or Spearman).
For comparability, coefficients based on neuronal thresholds were sign-inverted. The black line shows the overall regression (slope = 0.10, p = 0.001, N = 41); the thin golden line shows the regression within MT (data from other areas were too sparse for separate analysis).

Mean CP values across brain areas for single-electrode recordings
Diamonds and error bars represent medians and interquartile ranges for areas with more than five observations. Asterisks indicate significant differences between two brain areas based on Mann–Whitney tests (** p < 0.01, *** p < 0.001); p < 0.05 results are omitted as they would not survive correction for multiple comparisons.

Mean CP values across brain areas within the same study and monkey.
Lines connect data points from the same monkey within each study.

Comparison of task sensitivities between brain areas
Asterisks indicate significant differences between two brain areas based on Mann–Whitney tests (** p < 0.01, *** p < 0.001); p < 0.05 results are omitted as they would not survive correction for multiple comparisons.

Relationship between mean CP and stimulus duration for reaction time tasks

Relationship between mean CP and normalized sensitivity for fine-vs. coarse-discrimination tasks.
Linear regression analyses were performed separately for the two task types: solid line – coarse- and dashed line – fine-discrimination task). A separate model with interaction terms revealed no significant difference in slope or intercept: Δβ = βfine − βcoarse = −0.03, p = 0.3; Δβ0 = −0.007, p = 0.7.

Relationship between Mean CP values with whether experiment have reaction-time design.
The empty circles represent detection tasks; solid circles – all the other task types. In RT tasks (N = 24), 16 data points are from the detection tasks, 8 –from coarse-discrimination ones.

Mean CP values across task types for single-electrode recordings
Diamonds and error bars represent medians and interquartile ranges. Asterisks indicate significant differences between two task types based on Mann–Whitney tests (** p < 0.01, *** p < 0.001).

Relationship between stimulus duration and task type.
Asterisks indicate significant differences between two task types based on Mann–Whitney tests (** p < 0.01, *** p < 0.001).

Fine-discrimination tasks show substantially lower sensitivity, which may account for their reduced mean CP.
Normalized neuronal sensitivity (inverse mean N/P ratio) across task types. Diamonds and error bars represent medians and interquartile ranges with more than five observations. Asterisks indicate significant differences between two task types based on Mann–Whitney tests (** p < 0.01, *** p < 0.001).

Relationship between mean CP values and task exposure.

Relationship between Mean CP values and monkey’s lapse rate.

Relationship between Mean CP values and predictability of targets.
No pairwise comparisons were statistically significant (Mann-Whitney test, corrected for multiple comparisons).

Relationship between mean CP values and stimulus noise structure.
Mean CP values are shown for conditions using frozen vs. random stimulus seeds. Lines connect data points from the same study. Wilcoxon signed-rank tests have not revealed statistical difference (p = 0.3).

Relationship between mean CP values and CP estimation method.
Lines connect data points from the same study. Wilcoxon signed-rank test has not revealed statistical difference (p = 0.4)

Relationship between Mean CP values and method of neuron preference estimation.
Mann–Whitney test revealed that studies with preferences estimated from the task trials exhibited lower mean CP compared to those that used preferences estimated from passive viewing sessions (* asterisk in the figure, β = 0.018, p = 0.03), but this effect appears to be driven by other factors (see section 5.3).

Relationship between Mean CP values and task variable.
Asterisks indicate significant differences between two task variables based on Mann–Whitney tests (** p < 0.01, *** p < 0.001); p < 0.05 results are omitted as they would not survive correction for multiple comparisons. Note that the observed difference is largely driven by the effect of brain area, as task characteristics and brain area are strongly coupled.

Relationship between mean CP values and task variable within area MT.
Asterisks indicate significant differences between task conditions based on Mann–Whitney tests (*** p < 0.001); results with p < 0.05 are omitted as they would not survive correction for multiple comparisons. Unlike the across-area analysis, the observed difference between motion and depth discrimination tasks in MT cannot be attributed to brain area and remains significant after controlling for sensitivity, stimulus duration, and task type (see section 5.3).

Hierarchical inference model simulations with task-aligned feedback and incomplete learning predict within- and across-study CP–normalized sensitivity slopes of approximately 0.15.
Normalized sensitivity (x-axis) is defined as the d′neuron/ d′behavior, where d′neuron is estimated during task performance and reflect both feedforward and feedback influences (see Section 2.4.1). Simulations utilized the hierarchical inference model of Haefner et al. (2016) with 256 sensory neurons randomly divided into 8 “studies” of 32 neurons each (indicated by color). Feedback strength was controlled by the parameter δ and set to 0.08—the maximum value used in Haefner et al. (2016)—which corresponds to proficient but incomplete task learning. Small open circles represent individual neurons, and large filled circles denote study means. Thin lines indicate within-study regressions; the thick black line represents the across-study regression.

Relationship between mean CP values and stimulus eccentricity.

Relationship between stimulus duration and year of publication.

Relationship between neural sensitivity and number of neurons recorded in the study (only single electrode recordings).

Changes in the distribution of brain areas studied over time.
Proportion of data points from each brain area is shown across three publication periods, illustrating shifts in focus within the CP literature.