Reviewer #2 (Public review):
Summary:
Renard, Foustoukos and colleagues present a study of rapid sensorimotor learning in the mouse barrel cortex. Head-fixed water-restricted mice already trained on an auditory detection task are introduced to a novel C2 whisker stimulus, and the authors show that reward-paired mice acquire the whisker-lick association within a single behavioral session, with the two groups (rewarded vs non-rewarded) diverging behaviorally within ~22 whisker trials and ~14 minutes. Both pharmacological inactivation of wS1 across Days 0/+1/+2 and optogenetic inactivation on Day 0 impair whisker-guided performance, while fpS1 manipulations do not, establishing that wS1 activity is required for whisker-guided behavior during the initial learning period. Longitudinal two-photon imaging of GCaMP6f-expressing L2/3 neurons across five days (-2 to +2 relative to whisker introduction) reveals a bidirectional, reward-dependent reorganization of population responses to passive whisker stimuli: rewarded mice show enhancement, non-rewarded mice show suppression. The authors use a logistic-regression decoder trained to discriminate pre- vs post-learning passive trials and then project Day 0 active whisker trials onto this learning axis; the projection rises monotonically across Day 0 in R+ mice and is significantly correlated with behavioral performance, with no such trajectory in R- mice. Finally, the authors detect reactivation events during catch trials by template-matching to the average passive whisker response, and show that on Day 0, the neurons most positively modulated by learning (LMI-positive) participate in these reactivations more than LMI-negative neurons in R+ but not R- mice. The authors interpret this as evidence that online, reward-gated reactivations may act as an upstream selection mechanism for which neurons undergo learning-related plasticity, operating on the minutes-timescale of within-session learning. There is much to like in this paper, with some moderate-to-major concerns that could largely be addressed with re-analysis or re-framing.
Strengths:
The single-session learning paradigm is a key aspect of this paper, given the rapid learning observed. Coupled with the R+ and R- design, there's a lot to like with the behavioral approach. The bidirectional response change across these R+ and R- groups (enhancement vs suppression) is also a nice finding.
The causal manipulations demonstrate that the imaged region is used during the task. By doing both pharmacological and optogenetic inactivation, each with a control in the spatially adjacent region (fpS1), the authors make a strong case that wS1 activity is necessary for whisker-guided behavior during the initial learning period (though see below about the limitations of the current approach).
The longitudinal two-photon imaging of the same L2/3 neurons across five days underlies essentially every neural analysis in the paper and enables the single-cell LMI and population-trajectory analyses.
The pathway-specific analysis in Figure 3 - figure supplement 2 is very interesting, but not much time is spent on it (lines 151-155). The dissociation between wS2-projecting neurons (which show learning-related enhancement in R+ and suppression in R-) and wM1-projecting neurons (which do not) is (in my opinion) a nice instance of projection specificity - it also aligns with the known routing of task-relevant whisker information through the wS1→wS2 pathway. I would encourage the authors to motivate this experiment in the main text rather than leaving it all to the discussion (lines 256-262).
The methods are generally well documented and easy to follow.
Weaknesses:
(1) Conflation of de novo association learning with generalization from auditory pre-training.
All mice have already learned a task structure with the auditory task - "detect the salient sensory cue → lick → reward". Under these conditions, the rapid emergence of licking to the whisker stimulus could reflect either de novo formation of a whisker-specific association or generalization of an instrumental policy to a novel salient cue. The manuscript frames the result as the former ("acquisition of a novel sensorimotor association"), but the experiment cannot distinguish between the two alternatives. This distinction between de novo learning and generalization may have a meaningful impact on the interpretation, though it doesn't impact the specific results. It would be helpful for the authors to discuss the two possibilities and generally consider the contribution of generalization from auditory pre-training to Day 0 performance.
Relatedly, the R- group is introduced (lines 69-74) and later used (lines 244-247) as a passive-exposure control that rules out representational drift. While R- group is an important control for repeated whisker stimulation and task context, it does not appear to be a pure passive-exposure control: Figure 1B shows that on Day 0 the mice lick more to the R- stimulus than with no stimulus and then extinguish that licking by Day 1. Thus, one possibility is that R- mice actively learn to suppress licking to an unrewarded stimulus (whisker) in a context where other stimuli (auditory) remain rewarded. This would be a different cognitive operation (response suppression) from a purely passive exposure condition. The manuscript therefore lacks a true passive-exposure baseline, and several claims that rely on R- as such a baseline (including that bidirectional changes are reward-driven rather than reflecting passive drift, lines 244-247) need to be reframed.
(2) The inactivation experiments establish that wS1 is necessary on Day 0, but they cannot separate detection, acquisition, and expression.
Both the muscimol manipulation (whole session, Days 0/+1/+2) and the optogenetic manipulation (0.1 s before stimulus onset through the 1 s reporting window) silence wS1 during the moments when the whisker stimulus must be detected for a successful trial. Under these conditions, impaired performance could reflect that the animal cannot detect the stimulus, cannot express the learned response on that trial, or cannot acquire the association. These are causally distinct processes, and the manuscript currently treats them as equivalent.
Specifically, on Day +1 of the opto experiment (light off), do mice learn at the same rate as a naive Day 0 cohort (e.g., the R+ imaging mice on Day 0), or is performance already higher than the naive group? If higher than the naïve group, this would suggest that there is learning occurring and would suggest that something that may have been acquired during Day 0 inactivation, even if it could not be expressed.
(3) The interpretation of the LMI-participation correlation is complicated by the peaked LMI distribution and neuron-level pooling.
Two related issues arise from the results shown in Figure 4I. First, the LMI distribution in Figure 3F (and visible in 4I) is sharply peaked near zero. The reported r = 0.24 in R+ mice is therefore difficult to interpret biologically because the distribution is dominated by near-zero-LMI neurons and the slope may be disproportionately influenced by neurons in the tails. The key claim is better tested by comparing significantly LMI-positive, LMI-negative, and non-modulated neurons. The authors do address this in Figure 4J - showing that participation rate rises across days for significantly LMI-positive R+ neurons (p = 5×10⁻⁴) but not for LMI-negative neurons (p = 0.05) - but this analysis is not the lead result. To my understanding, Figure 4J is more interpretable and should be the key piece of data supporting their claim.
Second, the p-value of p = 1×10^-41 in Figure 4I comes from treating thousands of neurons pooled across 19 mice as independent observations. Neurons within an animal are correlated through shared behavioral state, shared imaging session, and circuit-level interactions, so it would be helpful to consider a different statistical unit of comparison (FOV, animal, etc). For example, a linear mixed-effects model with mouse as a random effect could work.
(4) The reactivation-LMI relationship is partially circular, and the framing in the abstract could be more constrained.
The "reactivation template" is the trial-averaged passive whisker-evoked population vector from each session, and reactivations are detected as moments in catch-trial activity that correlate with this template above a shuffled threshold. This approach is reasonable, but it means that the reactivation-LMI relationship is not fully independent of template construction, and the framing in the abstract blurs that line. LMI-positive neurons are defined as neurons whose passive whisker-evoked responses increase from pre- to post-learning. Therefore, neurons with strong whisker responses, or neurons that become stronger components of the whisker-evoked template across learning, may be more likely to contribute to template-matching events by construction. Thus, the LMI-participation relationship could partly reflect template weighting or sensory-response amplitude, rather than showing that reactivation events selectively recruit neurons for future learning-related plasticity. It would be helpful and more reassuring if the authors could control for each neuron's whisker-template weight, baseline whisker responsiveness, and overall calcium event rate when relating LMI to reactivation participation.
A complementary unsupervised approach could also help: rather than starting from the whisker template, one can derive co-activity assemblies directly from spontaneous activity (e.g., via PCA or ICA on the catch-trial population activity), and then ask, separately, whether any of these assemblies overlap with the whisker ensemble. The interesting test is then whether whisker-like assemblies become more frequently expressed across Day 0 in R+ but not R- mice, and whether LMI-positive neurons are preferentially loaded onto these whisker-like assemblies. This logic inverts the current pipeline and can be complementary to the current analysis. By identifying structure in nominally spontaneous activity first and then comparing to the whisker response, this could help avoid the circularity in which the template both defines the events and contains the cells being tested. The Figure 4 - figure supplement 1B partial-correlation analysis is a step in this direction but addresses only spontaneous firing rate, not template coupling. Without such a complementary approach, the authors may want to clarify that the reactivation detection is anchored to a template defined in part by the same cells whose participation is being tested.
(5) The reactivation-as-selection-mechanism interpretation is not supported by the current data.
The Discussion (lines 278-281) acknowledges that the authors have not shown necessity, but the end of the intro and part of the discussion (Lines 275-277) frame reactivations as a "reward-gated selection mechanism" for plasticity. An equally plausible alternative is that neurons whose synaptic inputs or intrinsic excitability have been potentiated by reward-driven learning will simply co-fire more often during quiet periods - meaning reactivations would be a consequence of plasticity that has already occurred rather than a mechanism that selects which neurons to potentiate. The current data cannot distinguish these.
A separate concern is the use of the term "spontaneous." The authors' usage is defensible in one sense - catch trials are stimulus-free, so the activity is not externally driven. However, "spontaneous" in the reactivation literature typically connotes offline, internally generated activity during quiet wakefulness or sleep, which carries different implications for plasticity than activity during active task engagement. Catch trials in this paradigm occur within the behavioral session, with the animal still engaged in the task, potentially anticipating reward or licking. The authors should either acknowledge this distinction in the text or qualify the term - "within-session" or "inter-trial" reactivations would be more accurate and would avoid borrowing the conceptual weight of the offline-replay literature.
The authors should also clarify whether catch-trial activity around licks (false alarms, anticipatory licks) is excluded from the reactivation analysis, and whether reactivation rates depend on recent reward, recent whisker trial outcome, or behavioral state. Specificity controls - template-matching with shuffled templates and with auditory templates - would help establish that detected events reflect whisker-specific patterns rather than generic high-coactivity moments.
(6) Motor, lick, and behavioral-state confounds in the neural analyses are not fully addressed.
I have two specific concerns. First, for the Day 0 active-trial projection, mean whisker reaction times in Figure 1 - figure supplement 1G are around 350-500 ms, but the distributions extend into the 0-300 ms analysis window. The correlation between the projection trajectory and the behavioral learning curve (Figure 4E, lines 196-198) is the key piece of evidence that the neural shift tracks learning. However, on hit trials the lick may fall within or close to the analysis window, so a motor confound could in principle contribute to the rising projection. The authors could repeat the projection using an earlier/shorter window, exclude trials with early licks, or regress out lick timing. It would be helpful to better understand whether this effect is, in part, driven by licking activity.
Second, the central evidence for representational reorganization (Figure 3) rests on a post-session passive epoch in which 50 whisker stimulations are delivered after "task disengagement" (lines 131, 387-389). The concern is that the brain state during this epoch is unlikely to be matched across groups or across days. R+ mice receive additional water rewards on whisker trials, whereas R- mice receive rewards only on auditory trials. This could lead to systematic differences in satiety, arousal, and disengagement state during the passive block. Because cortical sensory responses are strongly modulated by arousal, some of the apparent learning-related enhancement (R+) or suppression (R-) of passive whisker responses across days could reflect systematic state differences during the passive epoch rather than plasticity. The disengagement criterion ("stopped licking in all trial types") is also qualitative - no consecutive-miss or time-window threshold is specified - so the epoch may begin at slightly different behavioral states across mice. To resolve this, the authors could (i) specify the disengagement criterion quantitatively and (ii) compare pupil diameter and whisker self-motion (if available) across R+ vs R- and across days during the passive epoch.