1. Sep 2026
    1. Reviewer #2 (Public review):

      Summary:

      Grichine et al. investigate platelet-mediated fibrin compaction using human donor platelets and propose a novel mechanistic model in which platelets generate contractile forces and wind fibrin fibres into compact, coiled structures. Using a combination of 2D spreading assays, 3D clot imaging via expansion microscopy, live-cell imaging, and computational modelling, the authors present evidence of cage-like fibrin architectures, coiled fibre morphologies, and platelet-centred "rosette" structures that are present during fibre compaction. They suggest the involvement of actomyosin in fibre compaction and, overall, the study addresses an important and longstanding question in thrombosis and haemostasis while offering a conceptually novel perspective on clot compaction.

      Strengths:

      The integration of multiple imaging modalities is a notable strength. In particular, the 2D fibre-retraction assay provides a useful model for understanding the spatiotemporal dynamics of platelet-mediated fibrin compaction, which could be applied to other systems and may yield detailed mechanistic insights into biological processes. The live-imaging approaches are particularly well executed and provide valuable dynamic insights into fibre accumulation and compaction.

      Weaknesses:

      The primary weakness of the paper is the absence of direct evidence demonstrating the mechanism of fibre compaction via cytoskeletal swirling. Consequently, the relationship between platelet dynamics and fibrin organisation, including coordinated measurements of platelet motion and fibre rearrangement, is not directly assessed (perhaps due to technical barriers). However, the paper does provide solid evidence through myosin inhibition and computational modelling, demonstrating how platelets might mediate fibre compaction.

      Comments on revised version.

      Overall, the study addresses an important question in thrombosis and haemostasis and introduces a potentially impactful conceptual framework for understanding clot compaction. The imaging approaches and datasets presented will be valuable to the community, particularly to researchers interested in platelet mechanics and fibrin organisation. The possibility that fibres can be compacted extracellularly through cytoskeletal swirling represents a compelling and relatively unexplored mechanism. Therefore, this paper does a good job of establishing a thought-provoking mechanism with solid supporting evidence, although a direct demonstration of the underlying molecular mechanism requires further investigation.

    2. Reviewer #3 (Public review):

      Summary:

      This work aims to understand the mechanisms that platelets use to interact with and compact fibrin fibers during clot formation. This is an important process during wound healing and recent work has demonstrated that platelets play a critical role in generating the force required to drive accumulation of fibrin. The authors argue that current models are insufficient to account for the observed reduction in clot volume and propose that platelets actively 'wind-up' these fibers by undergoing myosin-dependent rotation. While interesting, the experiments performed by the authors do not directly test this mechanism and further evidence is required to support their claims.

      Weaknesses:

      (1) The motivation to switch from the system used in Figure 1 and 2 to the '2D fiber-retraction assay' is not clear. While the authors state that this system has 'reduced complexity' the differences between these assays appears to disrupt the 'cage-like' organization of fibrin around platelets shown in Figure 1 and 2 (compare images in Figure 2 with those in Figure 4). An in-depth comparison of two methods is needed to support the conclusions from the 2D system. Furthermore, the change in plasma volume (Figure 2 vs Figure 7) should also be tested - the authors state that this increases fibrin fiber formation, but this is not quantified or demonstrated in the figures. Notably, this appears to change the morphology of the fibrin fibers shown (comparing Figure 2 and Figure 7).

      (2) It is unclear how the classification of platelets as 'fiber-winding' versus 'fiber compaction' differs in Figure 2. The criteria used for these classifications should be stated. Further, it seems premature to characterize fibers as wound without having established this earlier in the manuscript.

      (3) Is the 'gearwheel' different from the 'cage' of fibrin fibers? They appear similar, but it is difficult to distinguish between these with only qualitative descriptions of these phenotypes.

      (4) The quantification of platelet extensions in Figure 9 is confusing. While the those in 9A are clear, those in 9B are not. For instance, what is the difference between #7 and #8 in the middle panel of 9B? It does not seem like #8 is labeling an extension.

      (5) It is unclear what the modeling accomplishes as there is no comparison between the results of these simulations and their experiments.

      (6) The data presented in Figure 12 provides the most direct support for their mechanism, but falls short of directly testing their claims. These experiments should be repeated to include blebbistatin to test the contribution of myosin and include quantitative rather than qualitative comparisons of these experiments.

      Comments on revised version:

      The manuscript is substantially improved in clarity and organization. The authors have adequately addressed most of my concerns regarding presentation and interpretation through revisions to the text and figures. However, my primary mechanistic concerns remain unresolved. Although the proposed model is now presented more cautiously, the revised manuscript still does not directly test the central mechanism, and the conclusions therefore remain insufficiently supported.

    3. Author response:

      The following is the authors’ response to the original reviews.

      We thank the editor and the reviewers for their time and efforts to evaluate our manuscript. We have taken into account all the comments and revised the manuscript accordingly which has considerably strengthened the message.

      Several parts of the manuscript, have been extensively rewritten to add explanations and clarify our hypothesis and claims, this also led us to add four new references.

      In addition, we have added the following new figures:

      - New part of figure 2. Figure 2E shows a platelet in a constrained clot after 4h of retraction with the fibrin cage around the platelet center still present. The actin staining of the platelet shows that radial actin fibers are present in each bulb extending to the platelet center. This observation supports our hypothesis that in each bulb an individual cytoskeletal swirling could take place resulting in the accumulation of fibrin fibers at the base of each bulb.

      - Modification of figure 3, to include the criteria used to define four categories of platelets and associated fibers in the 2D fiber retraction assay (new Fig. 3C).

      - New figure 13, illustrating the quantification of fibrin fibre compaction mediated by platelets in the 2D fiber retraction assay and the rotational movement of a fiber mass (video 9).

      - New supplementary figure 1, showing the result of a new model simulation in the absence of cytoskeleton swirling. Under this condition the fibrin fiber does not loop around the platelet bulb.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper reports a previously unrecognized mechanism by which platelets compact fibrin fibers during clot retraction. Rather than simply pulling on fibers, the authors propose that platelets generate swirling motions that wind and loop fibrin into dense structures.

      While the results are intriguing, the underlying physical mechanism remains unexplained. In particular, it is unclear how platelets generate swirling motion capable of inducing fibrin coiling, especially when suspended in 3d fibrin mesh. This raises concerns about the conclusions.

      The reviewer is right, it is difficult to imagine how platelets in a 3D fibrin mesh can accumulate fibers at the base of their extensions to form a cage-like fiber organisation around the center of the platelets. We therefore developed the 2D fibre-retraction assay, which we believe provides important insight for the coiled fiber accumulations above spread platelets in the 2D situation but also provides a framework for interpreting similar processes that may occur within a 3D clot. In response, we have placed greater emphasis on clarifying and strengthening the comparison between the potential mechanistic aspects in the 2D and 3D assays, in order to better support our proposed model (see Results, section: "Platelets, spread on a 2D surface, organize fibers above them", last paragraph). In addition, the Ideas and Speculations section of the discussion has been extensively rewritten to provide more detailed explanations about the potential mechanism leading to fibrin fiber accumulations around platelet bulbs in a 3D fibrin mesh.

      Also, does fibrin have inherent chirality or structural asymmetry that could promote coiling independently of platelet activity?

      Yes, double-stranded fibrin protofibrils have a helical twist [1]. Furthermore, a clot formed in the absence of platelets and other cellular components shows intrinsic tensile forces [2]. However, we show that inhibition of actomyosin actions prevents fibrin fiber accumulation in the 2D fibre-retraction assay providing evidence that platelet actions are necessary to observe the coiled fibers above spread platelets. This has been accentuated in the revised version and three references have been added.

      Furthermore, platelet retraction typically involves platelet aggregation rather than isolated cells, and it is unclear how fibrin coiling would proceed in clustered platelets.

      Under the in vitro fiber retraction conditions used in our study (constrained or unconstrained clots or even in the 2D assay) individual platelets are homogenously distributed within the forming clot or on the coverslip. Therefore, there are no big platelet aggregates or clusters of platelets under our experimental conditions and the results can only demonstrate how individual platelets act on fibrin fibers. This point has been emphasized in the revised version (Discussion, third paragraph).

      Reviewer #2 (Public review):

      Summary:

      Grichine et al. investigate platelet-mediated fibrin compaction using human donor platelets and propose a novel mechanistic model in which platelets generate contractile forces and wind fibrin fibers into compact coiled structures. Using a combination of 2D spread assays, 3D clot imaging via expansion microscopy, live-cell imaging, and computational modelling, the authors present evidence of cage-like fibrin architectures, coiled-fibre morphologies, and platelet centred "rosette" structures present during fibre compaction. They further suggest that actomyosin-driven cytoskeletal dynamics, potentially involving rotational or swirling motion, underlie this proposed winding mechanism, analogous to DNA looping and compaction. The study addresses an important and longstanding question in thrombosis and hemostasis and offers a conceptually novel perspective on clot compaction.

      Strengths:

      The integration of multiple imaging modalities is a notable strength of this paper. In particular, the 2D fiber-retraction assay provides a useful model for understanding the spatio-temporal dynamics of platelet-mediated fibrin compaction, which can be applied to other systems and may yield detailed mechanistic insights into biological processes. The live-imaging approaches are particularly well executed and offer valuable dynamic insight.

      Weaknesses:

      The primary weakness of this paper lies in its descriptive nature and its reliance on correlative rather than causal evidence. Several interpretations are not uniquely supported by the data presented. For example, the categorisation of fibrin accumulation in 2D assays as "fiber winding" and "fibre compaction" remains descriptive without establishing winding as a mechanism.

      When introducing the 2D fiber-retraction assay (figure 3) in the revised version, we now only mention the terms fiber accumulation and compaction to better align with the level of evidence, since wound-up fibers cannot be distinguished in this figure. The criteria to establish the four categories of platelets and associated fibers in the 2D fiber retraction assay have now been included in figure 3C.

      Nevertheless, coiled fibers above spread platelets are clearly visible in figure 4 and 8 and dynamic fiber rotations or winding-up are observed in figure 12 and video 9. These observations have been presented more cautiously, as indicative rather than definitive evidence of a winding mechanism.

      Alternative mechanisms, such as circular bundling, stacked fibers under tension, or fibrin crosslinking-induced aggregation, are neither excluded nor investigated.

      For fibrin fiber bundling, staggered or crosslinked protofilaments no platelet actions are necessary as described previously [2,3]. Since we observed a clear difference between +/- blebbistatin conditions in the 2D fiber-retraction assay, the fiber compaction we observe depends on platelet actions. Consequently, we consider these alternative mechanisms unlikely based on our data. This has been stated explicitly in the results section and discussion and three references have been added.

      Although the authors present compelling live imaging, establishing winding as a dynamic phenotype would require quantitative analyses, such as measuring angular velocities and coiling rates.

      We have incorporated quantitative measurements (new figure 13) about platelet mediated fibrin fiber compaction and angular rotation velocities to complement the observations obtained from live imaging. It is important to note, however, that angular velocities and coiling rates are likely influenced by the number of fiber–fiber contacts present at the time coiling occurs. Specifically, an increased number of contacts is expected to elevate tension within the network, thereby modulating the forces generated by platelets and, consequently, affecting both velocity and coiling dynamics.

      The use of a second fluorophore-labelled fibrin population could further strengthen evidence for rotational dynamics.

      These live videos are quite difficult to acquire because of the following reasons:

      - Small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time needed to adjust parameters for image acquisition, necessitates an arbitrary choice of the acquisition window and only one acquisition (90 min) per sample preparation is possible.

      - Furthermore, the laser-induced illumination can perturb the observed processes. We therefore use high-spatial-resolution 3D confocal time-lapse imaging, performed in photon-counting mode with very low laser excitation.

      For these reasons, the use of additional markers would be technically challenging and could perturb the delicate equilibrium and dynamics of the process under investigation.

      Similarly, the inference of rotational contractility or actomyosin "swirling", based on chiral actin organisation and blebbistatin treatment, is not sufficiently supported to conclude that platelets actively wind or loop fibrin fibers.

      Importantly, in the 2D fiber-retraction assay, we do not propose that the rotational actomyosin activity leads to a contractility of the platelets which would allow fiber retraction. Rather, we suggest that cytoskeletal actomyosin swirling (as demonstrated for nucleated cells by Bershadsky's team) can induce rotational dragging of extracellular bound fibrin fibers around the pseudonucleus of spread platelets thereby promoting accumulation of fibrin fibers (shown in figure 12C, video 9, third panel). Consistent with this interpretation, inhibition of myosin by blebbistatin prevents the accumulation of fibrin fibers above spread platelets in the 2D fibre retraction assay (Fig. 3).

      The mathematical model, while complementary and well-constructed, relies on multiple assumptions and lacks predictive validation.

      We thank the reviewer for this insightful comment and acknowledge that the proposed model relies on several important assumptions. In our view, the most significant assumption is that integrin molecules undergo rotational downstream motion as a consequence of their coupling to the swirling cytoskeleton. To assess the necessity and impact of this assumption, we provide an additional simulation performed in absence of the cytoskeletal swirling. Under this condition the fibrin fibers are not looped around the platelet bulb (this result has been added as supplementary figure 1). This analyses also provides further validation of the proposed model and underlying mechanism. At the same time, it is important to emphasize that the primary purpose of the model was to examine whether the hypothetical swirling dynamics of the cytoskeleton, together with the associated receptors, could in principle reproduce the experimentally observed fibrin organization.

      Appraisal:

      While the authors successfully document intriguing fibrin architectures and provide a compelling descriptive framework, they do not fully demonstrate a mechanistic model of active fibrin winding by platelets. The conclusions regarding platelet-driven winding and rotational dynamics are not sufficiently supported by direct or quantitative evidence. To substantiate these claims, the study would benefit from experiments that directly link platelet dynamics to fibrin organisation, including coordinated measurements of platelet motion and fibre rearrangement. As it stands, the results are suggestive but do not definitively support the proposed mechanism.

      Discussion and Impact:

      Despite these limitations, the study addresses an important question in thrombosis and hemostasis and introduces a potentially impactful conceptual framework for understanding clot compaction. The imaging approaches and datasets presented will be valuable to the community, particularly for researchers interested in platelet mechanics and fibrin organisation. However, the overall impact will depend on whether the proposed mechanism can be more rigorously validated. In its current form, the study presents an interesting and thought-provoking model, but would benefit from either stronger experimental support for the proposed mechanisms or a more cautious interpretation of the findings.

      We agree that the proposed mechanism requires further validation. In the revised version we have added a new result (figure 2E) showing that radial actin filaments are present in each bulb of a platelet in a constrained clot, supporting the possibility that rotational cytoskeletal movements could take place in individual bulbs. In a new figure 13, we have also quantified fiber compaction and the angular velocity of a rotating fibrin mass observed in video 9. Furthermore, in the revised manuscript, we present a more cautious and explicitly hypothesis-driven interpretation of the mechanism. We hope that the publication of our observations will be of interest to researchers in the field of thrombosis and clot mechanics who possess the specialized tools and expertise necessary to rigorously evaluate and either substantiate or refute the proposed mechanistic model.

      Reviewer #3 (Public review):

      Summary:

      This work aims to understand the mechanisms that platelets use to interact with and compact fibrin fibers during clot formation. This is an important process during wound healing, and recent work has demonstrated that platelets play a critical role in generating the force required to drive the accumulation of fibrin. The authors argue that current models are insufficient to account for the observed reduction in clot volume and propose that platelets actively 'wind up' these fibers by undergoing myosin-dependent rotation. While interesting, the experiments performed by the authors do not directly test this mechanism, and further evidence is required to support their claims.

      We do not "propose that platelets actively 'wind up' these fibers by undergoing myosin-independent rotation" of the whole platelet, but rather of the cytoskeleton winding-up extracellular fibrin fibers attached to integrin receptors.

      Weaknesses:

      (1) The motivation to switch from the system used in Figures 1 and 2 to the '2D fiber-retraction assay' is not clear. While the authors state that this system has 'reduced complexity', the differences between these assays appear to disrupt the 'cage-like' organization of fibrin around platelets shown in Figures 1 and 2 (compare images in Figure 2 with those in Figure 4). An indepth comparison of two methods is needed to support the conclusions from the 2D system.

      We agree that the cage-like fibrin organization around platelets is disrupted in the 2D fibre-retraction assay when platelets are completely spread on the coverslip before they have encountered fibrin fibers (Fig. 4). This has been explicitly stated in the revised version. However, some platelets in the 2D fiber-retraction assay form the same number of extensions as platelets in a 3D clot (Fig. 9 A, B) and are not completely spread on the glass surface. For these platelets a cage-like fibrin organisation is retained under the 2D conditions (Fig. 5 and 6). Nevertheless, the fiber density at the base of the bulbs is higher in the 2D assay than under the constrained 3D clot retraction conditions (Fig. 1C and Fig. 2), probably because in the 2D condition the fibers are less constrained and readily available for compaction.

      Furthermore, the change in plasma volume (Figure 2 vs Figure 7) should also be tested - the authors state that this increases fibrin fiber formation, but this is not quantified or demonstrated in the figures. Notably, this appears to change the morphology of the fibrin fibers shown (comparing Figure 2 and Figure 7).

      We thank the reviewer for raising this point. We would like to clarify that Figure 2 and Figure 7 correspond to two distinct experimental setups: the constrained clot retraction assay (Figure 2) and the 2D fiber-retraction assay (Figure 7). As such, they are not directly comparable. We understand, however, that the reviewer is likely referring to the apparent differences between Figures 3–6 (lower plasma volume, higher fiber density) and Figures 7–8 (higher plasma volume, lower apparent fiber density).

      The reduced number of visible fibers in the latter condition is not solely a consequence of plasma volume per se, but rather results from the formation of a labile fibrin gel at higher plasma concentrations, which is lost during the fixation and aspiration steps. This effect was initially observed across samples from two donors with differing plasma fibrinogen levels. In one case, an unusually low fibrinogen concentration allowed the addition of higher plasma volumes without inducing gel formation. In contrast, in the other sample, a more typical fibrinogen level resulted in gel formation under the same conditions.

      Importantly, we performed all experiments using matched donor plasma and platelets. As a result, the precise fibrinogen concentration could not be determined prior to experimentation. Nonetheless, post hoc measurements confirmed that fibrinogen levels in most donor samples fell within the normal physiological range, which allowed us to always use the same plasma volumes for low and high plasma concentrations (4ul/ml PBS and 7 ul/ml PBS, respectively) except for one donor as mentioned above.

      (2) It is unclear how the classification of platelets as 'fiber-winding' versus 'fiber compaction' differs in Figure 2. The criteria used for these classifications should be stated. Further, it seems premature to characterize fibers as wound without having established this earlier in the manuscript.

      The reviewer probably refers to figure 3 and he is right; it is premature to mention fiber winding at this stage of the results section (see our response to reviewer #2). In the revised version, we have modified figure 3 to include the criteria used to classify the platelets into four different categories (Fig. 3C).

      (3) Is the 'gearwheel' different from the 'cage' of fibrin fibers? They appear similar, but it is difficult to distinguish between them with only qualitative descriptions of these phenotypes.

      The "gearwheel" is observed for completely spread platelets in the 2D fiber-retraction assay and a figure illustrating our hypothetical speculations to compare the 2D gearwheel with the 3D clot situation is presented in the discussion under the "Ideas and Speculations" paragraph (now Fig. 14). We have given a more comprehensive explanation of the proposed mechanism in the revised version.

      (4) The quantification of platelet extensions in Figure 9 is confusing. While those in 9A are clear, those in 9B are not. For instance, what is the difference between #7 and #8 in the middle panel of 9B? It does not seem like #8 is labeling an extension.

      For the platelet shown in the middle panel of Figure 9B, the extensions cannot be clearly distinguished in the MIP (Maximum Intensity Projection) image because extension #8 is positioned above extension #7 and is therefore superimposed in the projection. However, the two extensions can be differentiated when examining the 3D image stack (Video 4, upper panel). As indicated in the figure legend, the number of extensions was determined manually by scrolling through the z-stack image sequence. In the revised version, we will also define the abbreviation “MIP” as Maximum Intensity Projection.

      (5) It is unclear what the modeling accomplishes, as there is no comparison between the results of these simulations and their experiments.

      We thank the reviewer for this valuable concern. We chose not to combine the experimental fibrin organization and the modeling results within the same figure panel, as the resulting image would be too complex and difficult to interpret. We have, however, added a supplementary figure 3 showing the results of a new simulation in the absence of cytoskeletal swirling. Under these conditions no winding of the fibrin fiber around the platelet bulb can be observed. It is also important to emphasize that the comparison between the model and the experimental data was intended to be primarily qualitative rather than quantitative.

      (6) The data presented in Figure 12 provides the most direct support for their mechanism, but falls short of directly testing their claims. These experiments should be repeated to include blebbistatin to test the contribution of myosin and include quantitative rather than qualitative comparisons of these experiments.

      As mentioned already above, these live videos are quite tricky to acquire because of the following reasons: - small platelet size

      - Heterogeneity of platelets within the population (10 d half-life, old platelets may not be able to compact fibers efficiently).

      - The speed of the process and the time required to optimize imaging parameters, necessitate the selection of an arbitrary acquisition window. Consequently, only a single acquisition of approximately 90 min can be performed per sample preparation, with no guarantee that relevant platelet-fibrin interactions can be acquired in the acquisition window.

      - Furthermore, after blood donation, the first sample is usually ready to be acquired around 3 pm, acquisition time 90 min. At least 10 successful acquisitions per condition would be required to ensure statistical robustness, but maximal 4 can be acquired per donor, because platelet samples start to deteriorate within twelve hours after blood donation.

      Taken together, the intrinsic heterogeneity of the platelet population, the low likelihood of capturing informative events, and the limited availability of suitable imaging resources at our institute render a robust and quantitative comparison between conditions with and without blebbistatin extremely challenging, if not impractical, within a reasonable timeframe.

      In accordance with the reviewer's request, we have added a new figure 13 to the revised version, presenting quantitative data on the platelet-mediated fibre compactions and the speed of angular fibrin rotations observed in video 9.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Throughout the manuscript, it is difficult to map the data presented in the figures to the text in the results section. Often, many subpanels are referred to collectively (for example, 'Fig 4 AE and animation, Video 3' on line 150), and the reader is left to piece together how this data fits into the statements in the results section. More guidance from the authors would help to understand the connection between these data and their conclusions.

      In the revised version, we have provided clearer explanations to make it easier to understand the conclusions drawn from the data. Concerning the indication "Fig 4 A-E and animation, Video 3" just means that platelets shown in panels A-E of figure 4 can also be visualized in the animation video 3. We have also put an effort to clearly indicate which figure part is presented in the associated video.

      There are also many figures that contain redundant information. The authors should consider revising these figures and including some of these repeated images as supplemental figures.

      As noted by the reviewers, our study provides predominantly qualitative observations essentially because it is not obvious to choose parameters which would be pertinent and could be quantified accurately using expansion microscopy. A quantitative analysis would allow to show the quantification and a representative image to describe the phenotypes of platelet-mediated fibre organisations. Without a quantitative analysis, we consider it more appropriate to provide multiple examples, enabling the reader to assess the consistency as well as the variability across repeated observations.

      Additional References

      (1) Jansen KA, Zhmurov A, Vos BE, et al. Molecular packing structure of fibrin fibers resolved by X-ray scattering and molecular modeling. Soft Matter. 2020;16(35):8272-8283.

      (2) Spiewak R, Gosselin A, Merinov D, et al. Biomechanical origins of inherent tension in fibrin networks. J Mech Behav Biomed Mater. 2022;133:105328.

      (3) Ramanujam RK, Lavi Y, Poole LG, Bassani JL, Tutwiler V. Understanding blood clot mechanical stability: the role of factor XIIIa-mediated fibrin crosslinking in rupture resistance. Res Pract Thromb Haemost. 2025;9(4):102871.

      (4) Gaertner F, Ahmad Z, Rosenberger G, et al. Migrating Platelets Are Mechano-scavengers that Collect and Bundle Bacteria. Cell. 2017;171(6):1368-1382 e1323.

    1. Within clinical practice, these cited value tables are conventionally treated as a good surrogate for the true conversion value5. However, these values have been co-opted to produce metrics used in research despite a lack of clear methodological quality assurance.

      I think you are implying the clinical use case came before the research one - is that right and it's not the other way around?

      would maybe change "lack of clear methodological assurance" to "without additional methodological quality assurance"

    2. depending on the reference.

      Does this need another sentence setting the context for why these are used?

      ..to produce equi-analgesic doses measured in OME or MME depending on the reference. This is commonly conducted to enable effective comparisons within research studies (as an outcome measure), or clinically, to guide opioid switching in acute or chronic pain settings.

    1. eLife Assessment

      This valuable study examines how the rodent prelimbic cortex represents learned and generalized threat over time and identifies distinct stable and dynamic neuronal populations that contribute to these representations. The evidence is convincing, supported by longitudinal calcium imaging, appropriate control groups, and sophisticated analyses showing that the relevant neural signals cannot be explained simply by freezing behavior. The work provides a conceptual framework for understanding how stable threat-related representations supporting memory generalization and discrimination can be maintained despite ongoing changes in neuronal ensemble composition.

    2. Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure.

      To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Comments on revised version.

      The authors have addressed my previous concerns well, and the revised manuscript is substantially improved. In particular, the additional analyses strengthen the conclusion that prelimbic cortical activity reflects learned threat value rather than simply freezing behavior, while the revised framing and additional controls clarify the interpretation of the longitudinal neural dynamics. This paper represents an important contribution to our understanding of the neural mechanisms supporting aversive learning, memory, and generalization.

    3. Reviewer #2 (Public review):

      The authors have substantially revised the paper in response to the original review, which is greatly appreciated. It is clear that it will eventually make a nice contribution to the literature. This being said, the following points are somewhere between major and minor in term of their implications for interpretation of the study results. If they were to be addressed, the paper would be again improved.

      There are a few remnants of the past language that are not helpful re interpretation of the study results: 1) "Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subpopulations that integrate information across tones to encode their learned threat-value."; and 2) "Together, these findings suggest that the PL integrates sensory similarity with learned threat value to generate stable representations that support adaptive generalization and discrimination." Neither of these statement follows what has been shown in the study, even with inclusion of the results from the GLM analysis (see point 4 below).

      (1) This paragraph in the Discussion is difficult to follow: "Generalization has traditionally been explained by perceptual similarity (Shepard, 1987), whereby stimuli resembling a conditioned cue recruit overlapping sensory representations and evoke similar behavioral responses (Corches et al., 2019; Grosso et al., 2018). Although perceptual similarity clearly influences the extent of generalization, accumulating evidence indicates that it cannot fully account for generalized responding (Verra et al., 2026). More recent frameworks propose that associative learning assigns learned value to novel stimuli by integrating their sensory similarity with previous experience, allowing behavior to scale according to predicted biological significance (Verra et al., 2026; Zaman et al., 2023). Our findings provide a neural framework consistent with these ideas. Sensory similarity promoted consistent neuronal population responses across tones, whereas associative learning organized these responses into graded representations that tracked learned threat value across the stimulus continuum. Thus, sensory similarity appears to define the neuronal substrate upon which associative learning constructs value-based representations that support graded behavioral generalization."

      While the revisions have removed the many unnecessary references to inference and integration, this paragraph seems like it is adhering to the original idea of how the authors wished to present their work. If the authors wished to talk about something more than perceptual similarity in the context of generalization, they should have used a task that lends itself to a more-than-perceptual-similarity explanation. Again, the inclusion of the GLM analysis is suggestive for some of what the authors wish to say, but doesn't justify the statements that: "Sensory similarity promoted consistent neuronal population responses across tones, whereas associative learning organized these responses into graded representations that tracked learned threat value across the stimulus continuum." In short, the analysis does not substitute for the design that could have and should have been used to assess learned threat value independently of sensory similarity.

      (2) The next paragraph in the Discussion is also confusing. "Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance (Mau et al., 2020; Zaki & Cai, 2024). Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure (Lacagnina et al., 2019; Mau et al., 2020; Sangha, 2015; Zaki & Cai, 2024). Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations (Bouton et al., 2021), whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients (Dunsmoor & LaBar, 2013; Herzog et al., 2021; Jenkins & Harrison, 1960; Lommen et al., 2017). Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles."

      The issue with repeated testing is *not* caused by extinction per se. The issue is that non-reinforcement across the repeated testing should differentially affect the CS+ and CS-. Specifically, it should extinguish responding to the CS- stimulus at a rate that matches its distance from the CS+, thereby sharpening the CS+ versus CS- discrimination in precisely the ways that have been observed. Ergo, the repeated testing *is* a problem for inferences that might be drawn about the way that generalization gradients change with time; and *is* a problem for statements regarding "dynamic reorganization of cortical activity patterns over time." There is nothing in the study that allows one to comment on the reorganization of cortical activity patterns over time. The reorganization can and should be attributed to the repeated testing, which is confounded with time. Nonetheless, the reorganization must be due to the repeated testing and NOT time as the present findings are inconsistent with the well-documented broadening of generalization gradients with time.

      (3) In the next paragraph, the authors state: "At the same time, narrower generalization gradients and improved discrimination across retrieval sessions suggests ongoing memory updating. These observations are consistent with contemporary theories proposing that systems consolidation and retrieval-dependent updating are complementary processes through which memories continue to evolve after learning (Mau et al., 2020; Tome et al., 2024; Zaki & Cai, 2024)."

      In general, I'm not sure why one would invoke systems consolidation or retrieval-induced reconsolidation as an explanation for any of the present findings: they are not explanations of much at all. In this specific text, the authors seem to be implying an updating process that occurs independently of what is learned across the repeated sessions of testing. Why? The changes that occur in the behaviour and neuronal representations are perfectly explicable in terms of additional learning that occurs - of the sort that I hope to have made clear in my previous comment. Why invoke more than what is needed to explain the observed pattern of results?

      (4) Re the GLM analysis - The authors write that: "the fact that the GLM analysis indicates that these neurons reflect learned threat value more than freezing behavior, suggests that they encode an abstract property of the learned stimulus rather than simply mirroring behavioral output."

      This is fine if freezing fully indexes the state of conditioned fear and there are no other behaviours in which animals express their fear. If, however, fear is expressed in a range of other behaviours that are likely coordinated by the PL (e.g., startle, vigilance, scanning, orienting to source of danger), this interpretation of the GLM analysis is unwarranted. This is an important point and would be worth noting somewhere in the paragraph where the statement appears.

    4. Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under examined in the field.

      Comments on revised version.

      The authors have convincingly and thoroughly addressed my concerns. I have no further issues regarding this study.

    5. Author response:

      The following is the authors’ response to the original reviews.

      Public review:

      Reviewer #1 (Public review):

      Summary:

      The authors combine discriminative auditory fear conditioning with longitudinal in vivo calcium imaging to ask how prelimbic (PL) representations of learned and generalized threat evolve across recent and remote memory time points. Using two different CS+ frequencies and a no-shock control group, they report that PL population activity tracks graded behavioral generalization, that population similarity is highest for tones eliciting strong threat responding, and that distinct subnetworks can be identified that appear to encode tone-specific sensory features versus learned threat-related response structure. To my knowledge, this may be the first study to comprehensively examine neural encoding of fear generalization in prelimbic cortex (PL). The manuscript is ambitious and technically interesting, and several aspects are potentially important. In particular, the suggestion that neurons showing graded, learning-related response patterns become selectively stabilized over time is intriguing. The inclusion of two CS+ training conditions and a no-shock control also strengthens the case that at least some of the reported effects are related to associative learning rather than simple sensory differences. However, in its current form, the manuscript does not yet fully support the strength of the conceptual claims. Several issues limit confidence in the interpretation, including the possibility that repeated testing itself contributes to changes across days, uncertainty about the relationship between neural activity and freezing behavior, limited quantitative documentation of longitudinal cell registration, and a number of problems in figure clarity and statistical framing. Overall, the study contains promising observations, but the claims should be narrowed, and several analyses or controls would be needed to fully support the proposed framework.

      Detailed Comments

      (1) A general concern is that the repeated test procedure itself may contribute to extinction. Because the animals are exposed to multiple CS frequencies across multiple test days, and each tone is presented three times per session, some of the reported changes in behavior and neural activity across days could reflect extinction or repeated nonreinforced retrieval rather than the passage of time per se. This is especially relevant given that the manuscript makes claims about recent versus remote representations and representational drift over 30 days. At a minimum, the authors should discuss this limitation explicitly and temper claims about time-dependent changes. Ideally, they would include a control group in which animals are tested only once or twice (e.g., at an early and later time point with fewer CS frequencies), or a reduced-frequency testing design that minimizes extinction while still allowing evaluation of recent versus remote memory.

      We agree with the reviewer that repeated testing is an inherent limitation of longitudinal memory studies and may itself contribute to neural changes across sessions. Repeated retrieval can induce memory updating (reconsolidation) or extinction, the latter involving the formation of a new association between the CS+ and safety. Although memory updating may have contributed to the ensemble reorganization observed here, several aspects of our findings argue against extinction as the primary explanation for the observed neural changes.

      First, we observed substantial neuronal ensemble turnover beginning with the first retrieval session. This early turnover is consistent with previous observations in the prefrontal cortex [1, 2] and with growing evidence that cortical memory representations remain dynamic throughout systems consolidation [3, 4]. Longitudinal studies have shown that neurons are continuously recruited into and removed from cortical memory ensembles while memory expression remains stable [1-4].

      Second, we calculated discrimination ratios to quantify discrimination of each tone relative to the CS+ across retrieval sessions (Figure S1). These analyses showed that discrimination increased, rather than decreased, over successive retrieval sessions, a pattern inconsistent with the behavioral profile expected if repeated testing had induced extinction.

      Finally, one of the most novel findings of our study is that ensemble turnover does not affect all neuronal populations equally. The graded neurons identified by our clustering analysis maintained their identity and functional organization across retrieval sessions, and their activity was better explained by tone threat value than by freezing behavior (Figure 8). This selective stability indicates that ensemble reorganization is not a uniform process but instead preferentially affects specific neuronal subpopulations while preserving a stable threat-value generalization gradient. Thus, although repeated retrieval may contribute to ongoing ensemble reorganization, our results demonstrate that this process is selective and largely spares the neuronal subpopulations that encode graded threat-value representations.

      Accordingly, we have revised the Discussion to explicitly acknowledge these points as follows:

      “The ensemble turnover observed here is consistent with previous studies demonstrating dynamic reorganization of cortical activity patterns over time [1-3, 5]. Such reorganization has been proposed to provide flexibility by allowing new information to be incorporated into existing cortical representations while preserving stable behavioral performance [4, 6]. Several mechanisms could contribute to this turnover, including systems consolidation, retrieval-induced reconsolidation or memory updating, and repeated nonreinforced stimulus exposure [4, 6-8]. Although our experiments cannot distinguish between the first two possibilities, the behavioral data argue against extinction as the primary explanation. Extinction is generally associated with the formation of new CS+-safety associations [9], whereas discrimination ratios increased across retrieval sessions, indicating that animals progressively improved their discrimination between threat-associated and safe stimuli rather than acquiring generalized safety responses. This pattern is consistent with previous work showing that discrimination learning sharpens stimulus representations and narrows behavioral generalization gradients [10-13]. Importantly, turnover was not uniform across the population. Graded neurons retained remarkably consistent response profiles across retrieval sessions, and their activity remained more strongly associated with learned threat value than with freezing behavior. These observations indicate that stable components of the population code can coexist with extensive reorganization of surrounding neuronal ensembles.” Pg. 19

      (2) More generally, some of the reported learning-related neural differences may be driven by behavioral differences, particularly freezing, rather than by learning or generalization per se. For example, animals that freeze more to certain frequencies may show corresponding neural response differences simply because freezing alters PL activity. The authors should examine this possibility more directly. Analyses testing whether recorded cells encode freezing behavior, or whether tone frequency-related neural differences remain robust when comparing high- and low-freezing epochs, would help determine whether the reported effects reflect learned stimulus value rather than behavioral state differences.

      This is an important point, which was also highlighted by the other reviewers. To directly address this concern, we implemented the generalized linear model (GLM) analysis suggested by Reviewer 3. We modeled the neuronal activity time series using both tone identity and freezing behavior as simultaneous predictors. Because tone identity was fixed across trials whereas freezing varied from trial to trial, the GLM allowed us to dissociate their independent contributions to neuronal activity.

      As described in the original submission, freezing was estimated from the miniscope's onboard inertial measurement unit (IMU), which measures body acceleration along three axes. Rather than classifying freezing using a fixed threshold, we estimated the continuous probability of freezing from the accelerometer signal using a Gaussian mixture model. This probabilistic estimate was incorporated directly into the GLM together with tone identity, providing a conservative test of whether neuronal activity was better explained by freezing behavior or by the auditory stimulus.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves (Figure 4) and to the graded and frequency-selective neuronal subpopulations identified by our clustering analysis (Figure 8). Across both experimental groups and all analyses, the median regression coefficients (β) associated with tone identity were consistently larger than those associated with freezing, indicating that tone identity contributed more strongly to neuronal activity. Moreover, tone coefficients exhibited graded monotonic profiles that closely tracked the learned threat value of each tone, with graded neurons showing the strongest gradients (Figures 4a, 8a, and 8e). Consistent with previous reports [14, 15] freezing accounted for a modest but significant component of PL activity. However, only 6–8% of graded neurons were classified as freezing-dominant, indicating that for the vast majority of these neurons, tone identity was the stronger predictor. Together, these findings demonstrate that the graded representation of learned threat value persists after accounting for freezing behavior, supporting our conclusion that PL activity reflects learned threat value rather than merely the behavioral expression of fear.

      (3) A central feature of the manuscript is the analysis of neural response properties over an extended period of time, up to 30 days after learning. However, aside from a brief mention in the Methods that spatial registration was used, the manuscript provides very little quantitative information about this critical aspect of the study. The paper would be strengthened by including explicit metrics describing longitudinal cell tracking, such as the number and proportion of ROIs retained across all sessions, distributions of spatial-footprint correlations or centroid distances across days, and representative examples of matched imaging fields over time. Without this information, it is difficult to assess how strongly the longitudinal claims are supported.

      We thank the reviewer for this suggestion. We now include measures of registration quality in the resubmission. Specifically, we calculated shifts in centroid distances, proportion of ROIs retained across all sessions, and representative examples of matched imaging fields over time (Fig, S3).

      (4) The text states that "Figs. 1c and 1d show GCaMP6f expression in PL, representative calcium footprints, and activity traces". However, the figure as presented does not clearly show all of these elements, at least not in a way that matches the description in the Results. The correspondence between text and figure should be corrected.

      We corrected correspondence between text and Figure.

      (5) The labeling of Figure 2a is insufficient for interpretation. The legend states that the panel shows raster plots of sound responsiveness, but the axes and scaling are not clearly defined. It is not clear from the figure what the x-axis represents, whether the y-axis corresponds to individual neurons, where the CS period occurs, or what the activity scale at the right denotes. Also, the term 'rasters' implies that spikes were analyzed. It seems that the spike inference approach (CASCADE) was only used for later analyses. Perhaps 'heat-plot' would be more accurate here? Generally, this figure should be annotated more clearly so that the reader can understand it without referring back to the Methods.

      We clarified the labelling of the Figure 2a and call the graphs “activity-plots”.

      (6) In relation to Figure 3, the analysis of population-averaged responses across tone frequencies is useful, but the manuscript would be stronger with additional statistical analyses across time and across groups. For example, if the authors want to argue that learning induces graded changes in neural responses and that these evolve across time, they should directly compare within-group responses across days and also compare matched frequencies between the conditioned groups and the no-shock controls. These analyses would help establish whether the observed differences are genuinely learning dependent and whether they change significantly over time.

      For Figure 3, we maintained the previous one-way ANOVAs assessing changes in AUC per day to be able to note significance on the Figure panels. However, we added a three-way mixed-effects analysis, using group (CS15, CS3, no shocks), frequency (3, 7, 11, 15), and day of testing (2, 15, 30) as variables, with frequency and day of testing as repeated measures. The results were described as follows (statistical details Table S1):

      “To determine how AUC varied across groups over time, we performed a three-way mixed-effects ANOVA with group (CS+15, CS+3, and no shock), frequency (3, 7, 11, and 15 kHz), and time (test days 1, 15, and 30) as factors, with repeated measures on frequency and time. For positive responder neurons, the analysis revealed significant main effects of group (p < 0.001) and time (p < 0.05), as well as a significant group × frequency interaction (p < 0.001), whereas the time × frequency and group × time × frequency interactions were not significant (p > 0.05; Table S2a). Tukey-corrected post hoc comparisons showed that, in the CS+15 group, AUC differed between all frequency pairs except 11 and 15 kHz (p < 0.05). In the CS+3 group, the AUC at 3 kHz differed from those at 7, 11, and 15 kHz (p < 0.05), whereas no significant frequency differences were observed in the no-shock controls (p > 0.05). For negative responder neurons, the only significant effect was a time × frequency interaction (p < 0.01). Tukey-corrected simple-effects analyses revealed that, on day 30, the AUC at 15 kHz differed from those at 3, 7, and 11 kHz (p < 0.05; Table S2b). Because this pattern was observed across all experimental groups, including the no-shock controls, it is unlikely to reflect associative learning. These results indicate that although the AUC exhibited modest changes over time, these changes were not group-specific and therefore do not support learning-dependent alterations in neuronal responses. Together, these results show that despite substantial neuronal turnover, PL population responses encode generalization gradients, closely matching behavioral expression.” Pg. 9

      (7) The inclusion of two different CS+ frequencies and a no-shock control is a strength of the study and substantially improves the interpretation that graded neural responses are related to learning and generalization rather than to simple sensory processing or passage of time. That said, I am not entirely comfortable with the use of the term "inference" throughout the manuscript. What is being measured here appears closer to sensory generalization than inference in a stronger cognitive sense. The current task does not clearly require that animals infer hidden structure or stimulus value through abstract reasoning; rather, the generalized stimulus may simply be treated as similar to the conditioned cue. The terminology should therefore be reconsidered or softened.

      We thank the reviewer for appreciating the strengths of the experimental design and for this thoughtful suggestion regarding terminology. We agree that the term inference may overstate the cognitive processes engaged by the current task. Accordingly, we revised the terminology throughout the manuscript to describe these effects as graded generalization of threat value across stimuli. The new GLM analyses further support this interpretation by demonstrating that, in the conditioned groups, neuronal activity at both the population and single-neuron levels is explained substantially better by tone identity than by freezing behavior (Figures 4 and 8). We therefore retained the term threat value, as our results indicate that PL activity primarily reflects learned threat value rather than simply the expression of freezing behavior, but removed inference.

      (8) I also found the use of the term "valence" somewhat problematic. The manuscript appears to use valence to refer to graded responding across tones with different aversive significance, but valence typically refers more broadly to distinctions between appetitive and aversive value. Here, terms such as "threat value," "aversive value," may be more precise. The authors should consider revising this language throughout.

      We corrected the language and replaced valence for “threat value”

      Reviewer #2 (Public review):

      Summary:

      The following points are those that occurred to me across readings of the paper. They are listed in what I take to be the order of their significance. Many of the points relate to the loose use of language and invocation of concepts that are not warranted, given the study design and results obtained.

      Major Comments:

      (1) The concept of ensemble turnover is interesting - the way it is introduced and discussed implies some type of spontaneous change in the neural underpinnings of fear discrimination and generalization in the PL. But, of course, every trial involves an opportunity to learn about the threat CS or the generalization test stimuli, and I am troubled by the thought that stability in the neural underpinnings of fear discrimination and generalization will actually reflect the level of defensive behaviours evoked on different trial types and/or the discrepancy between those behaviours and the outcome of a given trial in the generalization test. That is, stability in the neural underpinnings may be related to an animal's certainty or uncertainty in the contingency between a stimulus and danger; or, put another way, an animal's confidence that danger will or won't occur given the presence of some stimulus. This is not uninteresting. It is, however, not considered anywhere in the paper, which is overloaded with references to inferred threat values and integration of information across different types of stimuli. The protocol is not one that requires inference about anything or integration across anything.

      We thank the reviewer for this thoughtful comment. We agree that our original wording may have implied that turnover was a spontaneous process. Repeated retrieval provides opportunities for updating the learned contingencies associated with both the conditioned and generalization stimuli, and therefore changes in ensemble composition across sessions need not arise independently of experience. We also agree that the stability of graded neuronal representations may be related to the animal's certainty about the learned contingencies. However, in our data the graded neuronal population remained remarkably stable across retrieval sessions, whereas changes occurred primarily within the dynamic, frequency-selective neuronal populations. This suggests that stable ensembles preserve representations of learned threat value while updating is concentrated in a distinct neuronal subpopulation. We have now incorporated these ideas into the Discussion.

      (2) I appreciate the link to Gu and Johansen in paragraph 3 of the Introduction, but the type of generalization under investigation here is not the same as the type of 'generalization' studied by Gu and Johansen [who used a sensory preconditioning protocol]. Nonetheless, the authors have forced the language used by Gu and Johansen into their paper, and this has created tension [at least for this reader] as the concepts introduced by Gu and Johansen [inference, integration] are simply not relevant given the generalization protocol used here. Here are a few examples of points where the tension might interfere with a reader's understanding:

      We thank the reviewer for these specific criticisms. We revised the manuscript throughout to remove or redefine terms like "inferred valence" and "integration," replacing them with clearer, more accurate descriptions of gradient generalization of threat value. Below we address each point raised by the reviewer regarding terminology clarifications.

      (a) 'We hypothesized that generalization to novel stimuli depends on stable subnetwork organization that enables comparisons between learned and inferred valence, as well as population-level features that reduce variability across related representations.'

      I understand the words in the hypothesis, but can't form a representation of what is being said because of the reference to terms that stand in need of clarification [inferred valence, variability across related representations], but, ultimately, won't be clarified. This needs to be re-expressed so that the reader can appreciate what is being said.

      (a) We hypothesized that the PL generates representations of learned threat value that support threat generalization and discrimination, and that these representations emerge from the coordinated activity of stable and dynamic neuronal subnetworks, preserving consistent relationships among stimuli despite ongoing cellular turnover.

      (b) 'Our results show that stable cortical subnetworks integrate the emotional "gist" of memory and inferred valence for novel cues over time, despite ongoing ensemble reorganization, and that population-level firing rate similarity across stimulus presentations determines threat generalization.'

      Again, what does this mean? How is the gist of a memory integrated with inferred valence for novel cues over time? The statement simply doesn't make sense. This needs to be rewritten for clarity.

      (b) The summary statement was rewritten: " Together, these findings provide a neural framework for understanding how the PL supports adaptive threat generalization and discrimination.” pg. 4

      (c) 'In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting the contingency learned valence as well as the inferred valence of novel tones across testing days...'.

      Can this be rewritten as 'In CS<sup>+</sup>15 mice, positively modulated sound-responsive neurons exhibited graded activity to the tone CS and its variants that were used to assess generalization.'? The overloading of the text with references to 'contingency learned valence' and 'inferred valence' is unnecessary and makes it much harder to understand what has been shown in the results.

      We adopted the reviewer's suggested rewording: " In CS<sup>+</sup> 15 mice, positively modulated sound-responsive neurons exhibited graded tone activity reflecting learned contingency value across testing days" pg. 9

      We will systematically review the entire manuscript to ensure consistency with this revised framing.

      (3) Re the same passage of text as in 2c:

      Is it the case that these neurons are simply tracking the expression of freezing to the various tones? The same question applies to the results obtained for the CS+3 mice. If this is the case, then why should the results be taken to support the banner statement that 'Sound-modulated PL population responses encode learned and inferred valence' - these analyses do not support that statement. And, as indicated, I don't believe that the language of learned and inferred valence is appropriate to such statements, given the nature of the protocol used and results obtained. It is a study looking at how populations of neurons in the PL respond during presentations of auditory stimuli that were subject to discriminative conditioning, and during tests of generalized freezing to other [intermediate] auditory stimuli.

      The reviewer is correct that the graded population responses observed in PL could reflect freezing behavior across tone frequencies rather than encoding an abstract threat-value representation. This important concern was also raised by other reviewers. To address it directly, we followed Reviewer 3’s suggestion and implement a Generalized Linear Model (GLM) using the time series activity derived from the Ca2+ signals, with both tone identity and freezing behavior included as predictors. This analysis allowed us to dissociate the respective contributions of tone frequency and freezing to the graded neural responses. Based on the outcome of this analysis, we concluded that tone identity was a stronger predictor of neuronal activity than freezing. These results are summarized in Figures 4 for all cells contributing to population responses and Figure 8 for the main neuron types identified in the clustering analysis (frequency-selective and graded neurons). All details of this extensive new analysis are shown in red in the revised resubmission.

      In addition, we revised the text to remove the terminology of “learned and inferred valence” throughout the manuscript.

      (4) It is stated that:

      'In no-shock controls, although both positive and negative responses were present, population activity was not modulated by tone frequency or valence'.

      What does this mean? I can understand that population activity was not modulated by tone frequency. But what does it mean to say that it was not modulated by valence? Why should it have been when none of the tones were conditioned in this group and, hence, mice were responding to all the tones equally? And given that this is true, I don't understand the use of 'valence' here, or the subsequent statements in this paragraph that 'graded responses require associative learning' and that 'PL population responses encode graded sound-valence associations that reflect both learning and inference, closely matching behavioral generalization.' The latter statement is particularly unwarranted and, again, highlights a major issue with the paper. It could and should be rewritten as 'PL population responses reflect behavioral generalization.' There is nothing in the additional language that adds to the reader's understanding of what has been shown. The reference to 'graded sound-valence associations that reflect both learning and inference' is completely unwarranted, given the nature of this study. It is anathema to the vast literature on stimulus generalization. If the authors wished to make statements of this sort, they should have taken a different approach, perhaps using protocols like those featured in Gu and Johansen.

      We thank the reviewer for this helpful comment. We agree that our use of the term valence in describing the no-shock controls was imprecise. Because none of the tones was associated with reinforcement in this group, there was no learned valence that could modulate neuronal activity. Our intention was simply to convey that, although both positive and negative sound-responsive neurons were present, the population responses did not vary systematically across tone frequencies. We have revised this section accordingly.

      We also agree that our original wording overstated the interpretation of the graded population responses. Our data do not demonstrate that associative learning is required for sound responsiveness itself; rather, they show that associative learning is required for the emergence of graded population responses that distinguish tones according to their learned threat value. We have revised the text to make this distinction explicit.

      Finally, we agree that our previous references to "learning and inference" were not justified by the behavioral paradigm. We have removed this language throughout the manuscript and now describe the findings more directly as graded representations of learned threat value that closely parallel the observed behavioral generalization gradients.

      (5) The section titled, 'Consistently active neurons preserve valence representations as newly recruited neurons sharpen remote memory traces' ends with the following summary:

      'Together, these results indicate that consistently active neurons maintain stable representations of learned and inferred sound associations across time, whereas neurons recruited after conditioning progressively acquire graded tuning at later retrieval stages. This dynamic refinement suggests that cortical memory representations become increasingly selective during systems consolidation, while a stable neuronal subpopulation preserves the core emotional content of the memory.'

      Once again, the summary is not in keeping with the results obtained. The 'dynamic refinement' of representations is far more likely to reflect the repeated testing across days 1, 15, and 30 rather than anything to do with systems consolidation - at the very least, it is the simplest interpretation of the results. The impact of repeated testing is evident in the sharpening of generalization gradients over time, which is contrary to what is otherwise observed in the literature - the incredibly well -documented broadening of generalization gradients with time. Given this impact of repeated testing, surely the changes in the neuronal population that underlie performance are more likely to reflect the learning that occurs on days 1, 15, and 30, which is reflected in reduced freezing to the non-conditioned tones. If this is a reasonable take on the results, then I don't see the basis for invoking systems consolidation at all, and I don't see the basis for inferring a stable neuronal subpopulation that preserves the emotional content of the memory. Rather, non-reinforced presentations of 'never-reinforced' tones result in recruitment of additional neurons that result in suppression of freezing responses to those stimuli.

      We thank the reviewer for this thoughtful comment. We agree that repeated retrieval is an inherent limitation of longitudinal memory studies and that repeated non-reinforced presentations of the tones provide opportunities for memory updating. Accordingly, we have revised the Discussion to explicitly acknowledge that repeated retrieval may contribute to the ensemble reorganization observed across sessions through memory updating or reconsolidation processes (Discussion, pg. 19).

      We also agree that the progressive sharpening of the behavioral generalization gradients across retrieval sessions is consistent with memory updating. Both the behavioral data (increased discrimination ratios) and the neuronal data (progressively sharper population generalization gradients among neurons active after conditioning) indicate that the memory representation became more precise over time. We now discuss this possibility explicitly in the revised Discussion. We also agree that fear generalization often broadens with time; however, this is not universal. Under discriminative conditioning paradigms, repeated retrieval can instead produce progressively narrower generalization gradients [11]. We have revised the Discussion to clarify this distinction and added the appropriate references (pg. 19).

      While the reviewer's interpretation is therefore plausible, we do not believe it fully accounts for our observations. If repeated non-reinforced presentations were the sole driver of the observed neuronal changes, one might expect a more uniform reorganization across the neuronal populations engaged by the task. Instead, the reorganization was highly selective. Neurons encoding graded threat value remained remarkably stable across retrieval sessions, whereas neuronal turnover occurred primarily within the frequency-selective subpopulations. Thus, although repeated retrieval may update the memory representation, the neuronal substrate supporting graded threat-value coding is largely preserved while refinement occurs within a distinct neuronal subpopulation.

      Moreover, we observed substantial neuronal turnover beginning with the first retrieval session, consistent with previous longitudinal studies showing that cortical memory ensembles remain dynamic despite stable memory [1-4]. This early emergence of turnover suggests that repeated testing alone is unlikely to account for the continuous population dynamics observed throughout the experiment.

      Rather than viewing these findings as evidence exclusively for either memory updating or systems consolidation, we believe they are more consistent with current models proposing that these processes occur in parallel. Several influential frameworks argue that memories are continuously modified through retrieval while simultaneously undergoing systems-level reorganization [4, 6, 16, 17].We have therefore revised the Discussion to interpret the longitudinal changes more conservatively as reflecting the combined influence of retrieval-dependent memory updating and systems-level reorganization.

      In summary, we have revised the manuscript to better acknowledge the contribution of repeated retrieval while emphasizing what we believe is the principal finding of our study: despite substantial turnover within the overall ensemble, the neuronal population encoding graded threat value remained remarkably stable, whereas refinement occurred primarily within dynamic frequency-selective neuronal populations.

      (6) In the section titled, 'Population vector similarity at stimulus onset determines degree of generalization', it is stated that:

      'Because population similarity peaked shortly after stimulus onset, we quantified similarity during the first 5 s after tone onset relative to the CS<sup>+</sup>. In CS<sup>+</sup>15 mice, population similarity was highest for 15/15 and 15/11 tone pairs with no differences between them.'

      Isn't this consistent with the view that the population response in the PL simply reflects the level of freezing? Freezing to the 15-15 and 15-11 tones is most likely to be similar on their first presentation prior to the effects of extinction on the 11 Hz tone; hence the results obtained. That is, these results appear to clearly indicate that neuronal responses in the PL reflect the degree of stimulus generalization, as evidenced in freezing behavior. Given all that we know about the involvement of the PL in expressing fear responses, it is not appropriate to claim that 'population vector similarity at stimulus onset *determines* the degree of generalization. The PL responses simply reflect the varying levels of performance displayed to the different types of tones. What have I missed that could be taken to support additional statements?

      We agree that, because population similarity is highest for the 3/3, 15/15, and 15/11 tone pairs and freezing is also greatest for these same stimuli, the neural data could, in principle, reflect a correlate of behavioral expression rather than an independent representation of learned threat value.

      To directly address this possibility, we implemented a generalized linear model (GLM) to dissociate the contributions of tone identity and freezing behavior to neuronal activity. Across all analyses, tone identity consistently explained substantially more variance in neuronal activity than freezing behavior. Importantly, this finding held not only for the full population of sound-responsive neurons used to generate the population similarity analyses (Figure 4), but also for both the stable graded neurons and the dynamic tone-selective neuronal populations identified by our clustering analysis (Figure 8). Thus, although freezing behavior contributes modestly to PL activity, it cannot account for the enhanced similarity of population vectors across stimulus presentations or the graded population responses that form the basis of our conclusions.

      In addition, the temporal dynamics of the population vector similarity analysis are not entirely consistent with the interpretation that PL activity simply reflects the expression of freezing behavior. Population vector similarity peaked during the first 5 seconds following tone onset, whereas freezing occurred intermittently throughout the tone presentations. Although this temporal relationship does not establish causality, it is consistent with the interpretation that PL activity reflects the learned threat value associated with each tone rather than merely tracking the magnitude of freezing.

      Finally, we have revised the manuscript to more clearly acknowledge the correlational nature of these analyses. Specifically, we now state that population vector similarity is associated with, rather than determines, the degree of threat generalization.

      Later in the same section, it is stated that 'population-level similarity at stimulus onset scales with behavioral threat generalization and is maximal for tones associated with robust threat responses.' For simplicity and, therefore, clarity, this should be rewritten as 'population-level similarity at stimulus onset reflects behavioral threat generalization.'

      We made this correction. (“These findings indicate that population-level similarity at stimulus onset scales with behavioral threat generalization”. pg. 13)

      (7) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'Our previous analyses show that learned and inferred associations are represented at the population level. However, these results do not resolve whether graded responses arise from pooled activity of frequency-selective neurons or from subnetworks encoding integrated learned valence across tones.'

      What does it mean to say 'integrated learned valence across tones'? As it presently stands, the meaning of the phrase is unclear. It only makes sense if one supposes that generalized freezing responses to the 11 and 7 kHZ tones reflect separate associations between those tones and the aversive foot shock US. This supposition is inconsistent with the rich literature on generalization of Pavlovian conditioned fear responses. Specifically, it is inconsistent with the many theories of fear generalization, which attribute the reduction in fear as one moves away from the specific conditioned stimulus to a decrement in the ability of the test stimulus to activate the trained CS-US association. My strong impression is that the authors would do well to ground their findings in theories of stimulus/fear generalization, of which there are many. This would better serve the results obtained [and the reader's appreciation of them] - at present, the unnecessary invocation of concepts does very little to enhance the reader's appreciation or understanding of what has been found in the study.

      We agree that the phrase "integrated learned valence" is unnecessarily opaque and we replaced it with more precise language “Our previous analyses demonstrated that threat-value generalization gradients are represented at the population level. However, these findings do not reveal how these representations arise. Specifically, the observed population gradients could emerge either from the pooled activity of frequency-selective neurons that respond to individual tones or from neuronal subnetworks that integrate information across tones to encode their learned threat-value.” (Pg. 13)

      (8) Another example of what has been a common theme in this review:

      '...we hypothesized that the PL active ensemble segregates into functionally distinct subnetworks: one encoding tone-specific sensory features with dynamic characteristics, and another responding to all frequencies encoding stable core memory content and inferred emotional valence.'

      What does it mean to say 'all frequencies encoding stable core memory content and inferred emotional valence'? Do the authors mean to say '...and another that tracks freezing/defensive responses regardless of whether they were elicited by the trained CS or one of the generalization test stimuli'?

      We thank the reviewer for pointing out that this section was unclear. We agree that our original wording was imprecise and could be interpreted as implying cognitive processes that were not directly tested in the present study. Accordingly, we have revised the terminology throughout the manuscript. We no longer refer to "inferred emotional valence" or "core memory content" and instead describe these neurons more specifically as exhibiting graded representations of learned threat value.

      This is not the interpretation we intended. To determine whether these neurons primarily reflected defensive behavior rather than learned stimulus value, we implemented a Generalized Linear Model (GLM) that dissociates the contributions of tone identity and freezing behavior to neuronal activity. Across the entire neuronal population, as well as within the stable graded and dynamic tone-selective neuronal subpopulations, tone identity consistently explained substantially more variance than freezing behavior (Figures 4 and 8). Furthermore, after accounting for freezing, the regression coefficients of the graded neurons continued to follow the learned threat value of the tones, exhibiting opposite monotonic gradients in the CS+15 and CS+3 groups. If these neurons simply tracked defensive behavior irrespective of the stimulus presented, this relationship would not be expected to persist after accounting for freezing. We therefore conclude that the activity of this stable neuronal subpopulation is better explained by graded representations of learned threat value than by defensive behavior alone, and we have revised the manuscript accordingly.

      (9) It is stated that - 'Graded clusters encode emotional valence but constitute only a fraction of the active population; yet valence coding at the population level remains accurate and precise. This indicates that neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.'

      What does this mean? Are the authors trying to say that - 'Some clusters of PL neurons track freezing responses. In spite of the fact that these are only a fraction of the total active neuronal population, the population-level response of PL neurons also tracks the levels of fear to the trained tone and its variants used in the test for generalization.' If this is what one wants to say, then the final statement in the reproduced section does not follow. That is, there is no indication that 'neurons newly recruited into the population-likely frequency-selective and organized within learning-independent clusters-can be shaped by associative processes through modulation of firing activity.' As noted, the characteristics of other ensembles that become active across the repeated tests on days 1, 15, and 30 are more likely to reflect learning from non-reinforcement that occurs within and across those sessions. Perhaps this is what is meant by the phrase, 'shaped by associative processes'? If so, it should be stated explicitly instead of left to the reader to work out.

      We thank the reviewer for highlighting that this section was unclear. We agree that the original phrasing was insufficiently precise. Our intention was to convey that only a subset of PL neurons displays graded tuning that tracks behavioral generalization across tones. Nevertheless, despite constituting only a fraction of the total active population, this graded coding is also reflected at the population level. This observation led us to hypothesize that neurons recruited into the active population after conditioning— likely dynamic, frequency-selective neurons—also contribute to these graded population responses through modulation of their firing rates.

      The reviewer correctly notes that the phrase "shaped by associative processes" was too vague. By this we meant that the firing properties of these neurons are modified by the animal's associative history, including both the original conditioning experience and any retrieval-dependent updating that may occur during subsequent test sessions. We have revised the manuscript to make this interpretation explicit rather than leaving it to the reader to infer.

      To test this hypothesis, the GLM analysis we implemented dissociated the contributions of tone identity and freezing behavior to neuronal activity. After accounting for freezing, tone identity (i.e., learned threat value) remained a significant predictor of neuronal responses. Importantly, this was also true for the dynamic, frequency-selective neurons (Fig. 8e–f), indicating that these neurons contribute to population-level representations of learned threat value through firing-rate modulation rather than simply reflecting defensive behavior.

      To clarify our interpretation, we have rewritten the relevant section as follows:

      "Graded clusters encode generalization gradients but constitute only a subset of the active neuronal population. Nevertheless, population-level representations, which incorporate all active neurons, remain robust and accurately preserve these gradients. This observation led us to hypothesize that neurons recruited over time (e.g., dynamic, frequency-selective cells) also contribute to threat-value representations. Consistent with findings in the hippocampus showing that neurons can encode task contingencies through firing-rate modulation despite responding selectively to a single location (Gagliardi et al., 2024; Huxter et al., 2003; Sanders et al., 2019), we tested whether dynamic, frequency-selective clusters exhibited firing-rate differences proportional to learned threat value." (page 15)

      Regarding the reviewer's suggestion that the characteristics of the newly recruited neurons may reflect learning during repeated non-reinforced test sessions, we agree that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population, and we now explicitly acknowledge this possibility in the Discussion. However, we do not believe that our findings are fully explained by repeated non-reinforced retrieval alone. First, no-shock control animals underwent the same repeated testing but failed to develop graded neuronal representations, indicating that repeated exposure in the absence of associative learning is insufficient to account for the observed changes. Second, both behavioral discrimination and the corresponding population-level neural gradients became progressively sharper over time, consistent with refinement of learned threat representations rather than an effect of repeated testing alone, which must lead to extinction.

      In summary, we thank the reviewer for highlighting both the ambiguity of our original wording and an important alternative interpretation. In response, we have clarified the text to explicitly define what we mean by associative processes, added a GLM analysis demonstrating that the newly recruited neurons encode learned threat value beyond freezing behavior, and revised the Discussion to acknowledge that retrieval-dependent memory updating likely contributes to the reorganization of the dynamic neuronal population.

      (10) The following points all relate to the Discussion and reiterate many of the points above.

      (a) 'A subset of neurons remains consistently active across sessions, preserving core components of the memory trace and supporting inference of emotional valence for novel sounds, while neurons recruited after conditioning progressively acquire valence selectivity at remote time points.'

      'Inference of emotional valence' is unclear and unwarranted for all of the reasons provided above regarding the use of language.

      We modified the language as stated in the prior points.

      (b) '...Our data reconcile these views by demonstrating that cortical representations of emotional valence emerge rapidly after learning and persist within stable subnetworks, even as the broader population undergoes substantial turnover. This architecture preserves core mnemonic content while allowing flexibility in the surrounding ensemble.'

      These statements assume that the PL neuronal responses reflect something more than the levels of freezing behavior to the different stimuli; what are the grounds for this assumption?

      We incorporated the new GLM analysis to address this point and conclusions.

      (c) 'Importantly, these subnetworks encode both learned contingencies and the inferred valence of novel stimuli along a graded representational axis, suggesting that strong recurrent connectivity provides a stable scaffold for emotional memory representations.'

      What is a graded representational axis, and what part of the first statement suggests that 'strong recurrent connectivity provides a stable scaffold for emotional memory representations'? If the authors' goal was to make statements about emotional memory representations vis-à-vis emotional memory content, they should have used protocols that allowed them to probe such content. The auditory fear conditioning protocol used here [followed by tests for generalization to other auditory stimuli that differ in frequency from the conditioned tone] is not one that lends itself to analysis of emotional memory representations or content.

      We agree that the term "graded representational axis" was insufficiently defined and could be interpreted in multiple ways. Because this terminology was not essential to our conclusions, we have removed it and instead describe the observed phenomenon as a graded population representation of learned threat value across tone frequencies. We also removed the statement suggesting that recurrent connectivity provides a stable scaffold for these representations, as this mechanistic interpretation is not directly supported by our data.

      We also agree that some sections of the manuscript overstated the scope of our conclusions and have revised the wording accordingly. Our study uses neuronal activity recorded during memory retrieval after learning, an approach widely used in studies of systems consolidation to infer how learned information is represented within neural populations. Accordingly, we have revised the manuscript to explicitly state that our findings pertain to neural representations of learned threat value during memory retrieval rather than the broader content of emotional memories.

      Finally, we agree that our data are correlational and do not establish the causal role of the neuronal representations we identify. Throughout the manuscript, we now refer more precisely to population- and single-neuron correlates of learned threat value during memory retrieval following auditory fear conditioning.

      (d) 'Dynamic tone-selective responsive neurons emerge independently of learning, as they are present in both control and experimental mice, reflecting pre-existing PL sensory-driven properties (Hockley & Malmierca, 2024; Zikopoulos & Barbas, 2006).'

      Maybe. They are also likely to have developed as a consequence of the repeated testing on days 1, 15, and 30, which involved intermixed exposures to the tones of different frequencies. That is, rather than 'pre-existing PL sensory-driven properties', the responses of these neurons might reflect the emergence of discrimination between the various tones across testing, and greater suppression of freezing to the non-trained tones compared to the trained tone across the various test intervals.

      We thank the reviewer for this thoughtful comment. Our interpretation that these neurons reflect preexisting sensory-driven properties of PL cortex is based on two observations. First, tone-selective neuronal clusters were present in both conditioned and no-shock control animals, consistent with previous reports of sensory responsiveness in PL cortex [18, 19]. Second, these responses were already present during the first retrieval session, when the intermediate frequencies were presented for the first time. Thus, they cannot be explained by repeated exposure to those tones across subsequent test sessions.

      We therefore interpret the frequency-selective response properties as pre-existing features of PL circuitry that are present independently of conditioning. In contrast, associative learning modifies the firing activity of these neurons, allowing them to contribute to graded representations of learned threat value. This interpretation is supported by our GLM analysis, which showed that, after accounting for freezing, tone identity significantly predicted the activity of frequency-selective neurons in conditioned animals but not in no-shock controls. Thus, while the frequency-selective response properties are present independently of learning, associative learning modifies how these neurons encode learned threat value. We have revised the manuscript to clarify this distinction.

      Reviewer #3 (Public review):

      Summary:

      Normandin et al. explore the coding of stimuli predicting an aversive event in the prelimbic cortex. Stimuli could either be explicitly paired, explicitly unpaired, or novel but with an inferred association with the aversive event (generalization). Long-term tracking of GCaMP-positive neurons allowed them to examine how coding evolves out to a month following training. In general, they found two types of ensemble codes. One was ensembles coding for each stimulus independently, but with enhanced responding to the one eliciting a freezing response. The other was ensembles that responded to all stimuli in proportion to their similarity to the stimulus paired with the aversive event, either increasing or decreasing their activation with the degree of freezing elicited by a stimulus. Importantly, this second set of ensembles was more stable across days, potentially providing a memory trace.

      Strengths:

      (1) The authors track ensembles in prelimbic cortex over long time scales, providing valuable information on the consolidation of neural codes.

      (2) Neural coding of generalization is examined, which is under-examined in the field.

      We thank the reviewer for appreciating our design to track ensembles over time and the relevance of studying the neural substrates of generalization.

      Weaknesses:

      (1) Difficult to determine if responses treated as encoding stimulus valence are driven instead by the behavior that the stimulus elicits, freezing.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an alternative interpretation is that the graded neuronal responses may partially reflect freezing-related activity rather than representations of learned threat value. In the revised manuscript, we acknowledge that previous studies have identified PL neurons whose activity tracks freezing independently of stimulus identity or associative content. To directly address this possibility, we implemented the reviewer's suggestion by fitting a generalized linear model (GLM) to the neuronal activity time series derived from the Ca<sup>2+</sup> signals, using tone identity and freezing behavior as predictors. Because tone identity is fixed across trials, whereas freezing varies both during tone presentation and across trials (see below our answer to the Recommendations to Authors), this approach allowed us to dissociate their respective contributions to neuronal activity. We are grateful for this excellent suggestion, which has substantially strengthened both the manuscript and the conclusions that can be drawn from our data. The new analyses are summarized in Figures 4 and 8.

      In the points below we summarize the new findings.

      (2) The study implies that the identified ensembles are causally related to valence memory, but no experimental interventions are performed to justify this.

      We appreciate the reviewer's point. We agree that our data are correlational in nature and that establishing a causal relationship between identified ensembles and valence memory would require experimental interventions such as combinations of optogenetic and two-photon manipulations, which are beyond the scope of the present study but represent an important direction for future work.

      We examined inter-individual variability in freezing relative to the proportion of graded cells but the number of mice used in this study (CS+3= 5 and CS+15=7) did not give us enough power to reach significance.

      Therefore, we modified the manuscript terminology accordingly, replacing causal language with phrasing that accurately reflects the correlational nature of our conclusions.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Many sections of the paper should be rewritten along the lines that I have suggested in my public review and below.

      Minor Comments:

      (1) INTRO - 'This broad accessibility reduces spatial specificity and increases learning variability...'.

      Broad accessibility of what, exactly? And how does the 'broad accessibility' reduce spatial specificity and increase learning variability? That is, I do not understand what the terms 'reduced spatial specificity' and 'increased learning variability' refer to at this point in the first paragraph...

      We rewrote the introduction and discussion to address the points raised by the reviewer.

      (1) INTRO - 'The prelimbic cortex (PL) contributes to the expression (Burgos-Robles et al., 2009; SierraMercado et al., 2011; Sotres-Bayon & Quirk, 2010) and the proper discrimination and generalization of threat memories (Rosas-Vidal et al., 2025; Stujenske et al., 2022).'

      What is achieved by calling it 'proper' discrimination and generalization? Can't one simply say that the PL contributes to the expression, discrimination, and generalization of threat memories?

      This was corrected.

      (3) INTRO - '... and that population-level firing rate similarity across stimulus presentations determines threat generalization'.

      Or, alternatively, that generalization of conditioned freezing responses from the tone CS to variants along the dimension of Hz values is reflected in systematic changes in the firing rate of PL neuronal ensembles; when the test stimulus is similar to the conditioned stimulus, the two elicit similar behavioural responses and evoke a similar population-level firing rate in the respective PL neuronal ensembles.

      We have revised the Introduction as stated above. However, as discussed in our detailed responses, freezing behavior cannot fully account for the observed patterns of PL activity.

      (4) METHODS - 'Memory retrieval was tested on days 1, 15, and 30 after conditioning to probe early, long-term, and remote memory (Bontempi et al., 1996). During retrieval, mice were tested in a novel context with the CS<sup>+</sup>, CS1<sup>-</sup>, and two intermediate frequencies (7 and 11 kHz), presented in semirandom order, with each tone repeated three times (Fig. 1a).'

      Why was testing conducted in a different context than that of conditioning? This is likely to result in an underestimation of generalization to the different tones...

      In tone fear conditioning, it is always customary to test in a different context to dissociate conditioning to the context vs conditioning to the tones, which usually take place simultaneously in the same context [20]. Therefore, testing generalization in a novel context gives the correct estimate of generalization to the tones in the absence of contextual conditioning confounds. Please note that while overall freezing levels may be lower in a novel context due to the absence of contextual conditioning, the relative generalization gradient across tones — which is what your study measures — is unlikely to be systematically distorted by context change.

      (5) RESULTS - 'No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that freezing reflected associative learning.'

      The inference doesn't follow from the result described. Was there more freezing among animals in the shocked groups compared to those in the no-shock group? I presume so - my point is that this comparison is the one that most directly speaks to the presence or absence of associative learning.

      Experimental animals exhibited not only higher overall freezing but also graded freezing responses across tone frequencies. It is important to note that no-shock controls did not display this pattern, ruling out the possibility that the different frequencies themselves elicited graded behavioral responses. To clarify this point, we revised the sentence as follows: "No-shock control mice showed no significant differences in freezing across frequencies on any testing day (p > 0.05; Fig. 1b, right), confirming that the graded freezing patterns resulted from associative learning rather than the acoustic properties of the tones." (Pg. 6)

      (6) 'Across animals and sessions, we identified distinct neuronal populations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound (Fig. 2b)...'

      To be clear, do you mean to say that there were distinct neuronal populations that consistently [i.e., across all three sessions] increased their responses to the tones [positive modulation], decreased their responses to the tones [negative modulation], showed variable responses to the tones [mixed responses], and did not respond to tones [not modulated]?

      The sentence refers to neuronal populations identified within each recording session based on their responses to the tones, not to neurons that maintained the same response profile across all three sessions. We have revised the text to make this distinction explicit. The only stable patterns across sessions were observed in graded neurons that were stable across retrieval.

      “Across animals, we identified distinct neuronal subpopulations showing positive modulation, negative modulation, mixed responses, or no consistent response to sound in each session (Fig. 2b)” Pg. 7

      (7) What does 'active' mean in relation to Figure 2? Does this refer to neurons that displayed either positive responses, negative responses, and/or mixed responses? In the text, it is stated that 'Sound responder neurons were classified using a test that detected modulation based on magnitude relative to baseline variability, allowing reliable identification of both transient and sustained responses while remaining robust to noise...'

      I can't work out if this is the same classification criteria used for the determination of positive modulation, negative modulation, and mixed responding.

      We thank the reviewer for pointing out this ambiguity. In Figure 2, the term "active" referred to neurons that exhibited significant sound-evoked modulation and were subsequently classified as showing positive, negative, or mixed responses. Thus, active and sound-responsive refer to the same population of neurons. We removed the word active to avoid confusion.

      The reference to transient and sustained responses describes the temporal profile of the calcium signals rather than separate response categories. Some neurons exhibited brief calcium transients that rose and decayed rapidly, whereas others displayed sustained activity throughout the tone presentation. The sound-response detection algorithm was designed to reliably identify both temporal response profiles. We have revised the manuscript to make these definitions explicit. We modified the sentence as follows: “Sound-responsive neurons were identified using a statistical test that detected activity modulation relative to baseline variability, allowing reliable identification of responses while remaining robust to noise. This approach was effective for neurons exhibiting either brief calcium transients that rose and decayed rapidly or sustained activity throughout the tone presentation.” Pg. 7-8

      (8) 'A moderate proportion of neurons was present across all retrieval sessions, with no differences between groups (p > 0.05).'

      Do you mean to say that 'A moderate proportion of neurons was ACTIVE across all retrieval sessions, with no differences between groups (p > 0.05)'?

      We replaced the word present and replaced it with “active”. Pg. 8

      (9) In the section titled, 'Different subnetworks encode acoustic versus learned properties of sound association', it is stated that:

      'If neurons encoding graded responses carry core mnemonic information, they should exhibit enhanced stability over time. To test this hypothesis, we quantified the proportion of registered neurons that retained their cluster identity across at least two retrieval sessions and compared these values to a shuffled null distribution (10,000 iterations), with multiple comparisons controlled using the BenjaminiHochberg procedure.'

      What does the comparison to the shuffled null distribution tell us exactly? I accept that some neurons were stable positive responders across at least two sessions. The comparison to the shuffled null distribution creates a false impression about the robustness of this stability or the 'enhanced stability over time'.

      Our intention in comparing the observed stability to a shuffled null distribution was to evaluate whether the proportion of neurons retaining cluster identity exceeded chance levels expected from random assignment. The shuffled distribution therefore provides a statistical baseline against which the observed degree of stability can be evaluated. We agree, however, that the wording “enhanced stability over time” may be confusing regarding this finding. We rephrased this paragraph to clarify that a subset of neurons retained cluster identity across all retrieval sessions at levels greater than expected by chance as follows:

      “These data demonstrate that graded clusters remain consistently active at levels exceeding chance, preserving their cellular identity and providing a stable representation of learned contingencies and generalization gradients.” Pg. 15.

      (10) ABSTRACT. The abstract states that, 'Stimulus-evoked population similarity scaled precisely with behavioral generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing, indicating that population-level similarity predicts inferred threat.'

      I believe that the sentence could be rewritten as, 'Stimulus-evoked population similarity reflected the degree of generalization, and consistent population states emerged only for tones associated with shock or those eliciting strong generalized freezing.'

      We revised the text according the reviewer’s suggestion; however, we had to shorten the sentence due to word limits. “Population similarity tracked behavioral generalization, whereas consistent population states emerged only for shock-associated or highly generalized tones.” Pg. 2

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The ensembles with graded activation in proportion to stimulus valence are described at various points in the manuscript as "maintaining the emotional 'gist'", "preserving core components of the memory trace", and "preserving core components of the memory trace". This conclusion is premature because there is an alternative interpretation. The graded response ensembles would also be consistent with coding for the freezing behavior itself, irrespective of the specific memory or stimulus association that drives it. An ensemble that encodes a behavior in this way would not be considered mnemonic, just as motor neurons in the spinal cord are not, even if they may fire during a conditioned response. Indeed, previous work has identified neurons in the prelimbic cortex that encode freezing independently from the stimuli that signal an aversive outcome (e.g., Kyriazi, Headley, and Pare 2020; Casanova, Pouget, ..., Vetere 2024).

      There are two ways the authors can address this point.

      (a) Fit a generalized linear model to the time series of inferred spiking activity from the Ca2+ signal and include stimuli and freezing as predictors. Since freezing behavior is inconsistent across trials, while stimulus presence is fixed, they can be disassociated. If, after accounting for freezing, responsiveness neurons still show a graded coding of stimuli that agrees with inferred aversiveness, this would strengthen their claim that they have identified an ensemble that corresponds with mnemonic or salience aspects of the stimuli.

      (b) Conduct no further analysis but cover the issue in the discussion as a limitation to their study and to dampen some of the language throughout the manuscript that implies that a memory trace has been identified.

      We thank the reviewer for this thoughtful and constructive comment. We agree that an important alternative interpretation is that graded-response ensembles could reflect freezing-related activity rather than representations of learned threat value. To directly address this possibility, we implemented a Generalized Linear Model (GLM) analysis, as suggested by the reviewer. The GLM was fitted to the activity of every sound-responsive neuron included in the population analyses and simultaneously incorporated tone identity and continuous freezing probability (derived probabilistically from miniscope acceleration) as predictors, allowing us to quantify their independent contributions to neuronal activity.

      We want to note that freezing was quantified from the miniscope's inertial measurement unit (IMU) using a two-component Gaussian mixture model applied to the log-transformed body-acceleration signal. Rather than classifying freezing with a binary threshold, we used the posterior probability of the low-movement state as a continuous freezing regressor. This approach captures graded variations in immobility and provides a more conservative test of tone encoding, because it accounts for more behaviour-related variance than a binary classifier, making it more difficult to detect an independent contribution of tone identity.

      We applied the GLM both to all sound-responsive neurons contributing to the population response curves and separately to the identified frequency-selective and graded neuronal subpopulations. Across all analyses, tone identity consistently explained neuronal activity better than freezing. Furthermore, the freezing-corrected tone β coefficients scaled with learned threat value, with graded neurons exhibiting the strongest monotonic gradients, indicating that they provide the most robust representation of learned threat value. These findings demonstrate that the graded coding of learned threat value persists after accounting for freezing behavior and therefore cannot be explained simply by the behavioral expression of fear. The new analyses are presented in Figures 4 and 8. Notably, although freezing-dominant neurons were present in both the tone-selective and graded populations, they represented only a small fraction of each group and were least prevalent among graded neurons (6–8%), further supporting the conclusion that graded neurons primarily encode learned threat value.

      In addition, we revised the manuscript to avoid language implying that these neuronal populations constitute a mnemonic trace. Instead, we consistently describe them as encoding learned threat value, a more accurate interpretation that is directly supported by the new GLM analyses.

      (2) The title makes a seemingly causal claim by using the term 'arise', "Learned and inferred valence arise from interactions between stable and dynamic subnetworks". While it is true that the authors show that both stable and dynamic ensembles encode valence, they do not demonstrate that the behavioral expression of valence depends on these codes, nor their interaction. Experimentally testing this is beyond the scope of this study (holographic two-photon stimulation of transient and stable ensembles?), but they may be able to get closer to it by examining inter-individual variability. The authors could measure the proportion of neurons in each subject that participate in the stable (graded responding) and dynamic (stimulus-specific) ensembles, and see if they predict individual differences in the expression of freezing behavior or its generalization. Indeed, this correlation may change across testing days.

      We agree that the term “arise” in the title may imply a stronger causal relationship than is directly supported by the present data. We modified the title in the resubmission as follows: “Complementary stable and dynamic prelimbic ensembles encode learned threat value underlying generalization and discrimination”

      The new LGM analysis confirms that a large proportion of neural activity can be predicted by tone threat value; therefore, we think this title fully captures our findings.

      We also appreciate the reviewer's suggestion to examine inter-individual variability. In the revised manuscript, we tested whether the proportion of graded neurons correlated with freezing behavior. However, the limited number of experimental animals in each experimental group provided insufficient statistical power to reliably assess this relationship. Accordingly, we revised the manuscript to clarify that our conclusions are based on correlational observations rather than causal inferences.

      In summary, we revised the title and related language throughout the manuscript to avoid implying causal mechanisms beyond the scope of the current experiments.

      Minor points:

      (1) I was surprised by the absence of an ensemble in the No-shock group that responded uniformly to all stimuli. Can the authors confirm this?

      Yes, we confirm this finding. It was unexpected to us as well. We would like to clarify, however, that some control neurons may have responded to more than one frequency, but these responses were too infrequent or too weak to be classified as a distinct graded neuronal population by our clustering algorithm. Thus, while broadly responsive neurons may have been present in the control group, they did not form a robust, identifiable ensemble comparable to that observed after fear conditioning.

      (2) Several different approaches were used to analyze the same Ca2+ responses to stimuli across testing days. These were the "Sound responder classification", "Average stimulus-aligned trace procedure", "Population similarity over time across tone pairs", and the construction of "Stimulus response vectors". These feature differing alignment/binning/interpolation, normalization, and response quantification procedures, and it is unclear why they cannot all be in agreement, at least when it comes to alignment and normalization.

      We thank the reviewer for this careful reading of our Methods. All analyses were performed on the same underlying calcium imaging dataset, but they were designed to address different aspects of the data and therefore required different preprocessing steps. The analyses share a common initial pipeline leading to the calcium traces (all z-scored across the session). Differences in subsequent processing (e.g., use of ΔF/F versus CASCADE-deconvolved activity, normalization, baseline correction, temporal binning, and interpolation) were introduced only when required by the specific analysis method.

      To make this clearer, we have substantially revised the Methods. We added a new overview of preprocessing section that summarizes the common preprocessing pipeline and explicitly distinguishes the shared steps from those that are analysis-specific. We also included a summary table describing the input signal (ΔF/F or CASCADE-deconvolved activity), normalization procedure, and temporal processing used for each analysis. Finally, the individual Methods sections were revised to eliminate redundancies and more clearly describe the steps to avoid confusion. We hope these revisions make the rationale for the different preprocessing procedures and the overall analytical workflow more transparent (Pg. 22-23)

      (3) In the methods section "Window-wise response quantification" the Ca2+ signal was baselinesubtracted and divided by the standard deviation in the baseline across trials, but that data was already presumably z-normalized to the baseline of each trial ("Data alignment and normalization"). This second step of normalization seems excessive. Why is it not sufficient to just take the average peri-stimulus response across the z-normalized trials from the "Data alignment and normalization" section? This is simpler and would capture the effect size of the response relative to baseline.

      We thank the reviewer for this careful observation.

      The two operations are also not the same normalization applied twice; they standardize different sources of variability. During the alignment step, each trial is z-scored relative to its own baseline by dividing by the standard deviation of that trial's baseline across time. This places all trials on a common within-trial scale before averaging. In the window-wise step, the trial-averaged, baseline-subtracted response is expressed relative to the standard deviation of the per-trial baseline levels across trials—a distinct quantity that reflects trial-to-trial baseline stability rather than within-trial fluctuations. The purpose of this second term was to down-weight windows in cells with unstable baselines across trials, and it entered the analysis only as a significance criterion; the magnitude threshold defining a sound responder was applied to the trial-averaged baseline-relative response itself. We have revised the Methods to clarify the distinct roles of these two normalization steps.

      To further address this concern, we re-ran the sound-responder classification after removing the second (between-trial) normalization step, so that responder detection depended only on the per-trial baselinerelative response magnitude and its temporal persistence. Across all cells, tones, and sessions (n = 89,504 cell–tone–session classifications), the two procedures agreed on 95.3% of labels. The small fraction of cells whose labels changed were almost exclusively those lying immediately at the detection threshold: 86.6% of changes involved cells moving into or out of the "modulated" category, whereas direct reversals between excitatory and inhibitory classification occurred in only 8 of 89,504 cases (0.009%). Consistent with the between-trial standard deviation being a less stable quantity when few trials are available, label changes were approximately twice as frequent in the three-trial retrieval sessions (5.2%) as in the ten-trial conditioning sessions (2.1%). Overall responder proportions changed only minimally (positive responders +2.2%, negative responders +4.3%), and all population-level findings—including the graded threat-value gradient across tones, its absence in no-shock controls, and its persistence after controlling for freezing in the , as analysis—were unaffected. These analyses demonstrate that our conclusions are robust to this methodological choice.

      (4) It would increase confidence in the tracking of neurons across days if the authors showed some example images of neurons tracked across days.

      We added an example in the Supplement. Additionally, we now provide measures of registration quality (Fig. S3)

      (5) Table S2 is a bit confusing. I take it that Graded A/B were only for CS15, and Graded C/D/E were only for CS3. Also, Common B1-4 were the cells with positive responses to individual stimuli, and Common C1-4 were the cells with negative responses to stimuli. If this is the case, it should be explained in the figure legend (or even better, clusters should be named and numbered consistently in all figures.

      Thank you for pointing this out, we corrected the Table to indicate which test corresponds to which figure and cluster, specifying which ones were positive or negative modulated. Please note old Table 2 is now Table 5

      (6) The term network and subnetworks implies some connectivity between neurons, but in this study, it is used to refer to the ensembles of cells activated in a similar manner. Since connectivity is never assessed, it would be better if the authors stuck to the terms ensembles or populations.

      We changed the wording and now use ensembles or populations

      (7) The legend for Figure S3 has the text 'eded', which seems to be a typo.

      We corrected this typo.

      References

      (1) Kitamura, T., et al., Engrams and circuits crucial for systems consolidation of a memory. Science, 2017. 356(6333): p. 73–78.

      (2) DeNardo, L.A., et al., Temporal evolution of cortical ensembles promoting remote memory retrieval. Nat Neurosci, 2019. 22(3): p. 460–469.

      (3) Tome, D.F., et al., Dynamic and selective engrams emerge with memory consolidation. Nat Neurosci, 2024. 27(3): p. 561–572.

      (4) Mau, W., M.E. Hasselmo, and D.J. Cai, The brain in motion: How ensemble fluidity drives memory-updating and flexibility. Elife, 2020. 9.

      (5) Gallego, J.A., et al., Long-term stability of cortical population dynamics underlying consistent behavior. Nat Neurosci, 2020. 23(2): p. 260–270.

      (6) Zaki, Y. and D.J. Cai, Memory engram stability and flexibility. Neuropsychopharmacology, 2024. 50(1): p. 285–293.

      (7) Lacagnina, A.F., et al., Distinct hippocampal engrams control extinction and relapse of fear memory. Nat Neurosci, 2019. 22(5): p. 753–761.

      (8) Sangha, S., Plasticity of Fear and Safety Neurons of the Amygdala in Response to Fear Extinction. Front Behav Neurosci, 2015. 9: p. 354.

      (9) Bouton, M.E., S. Maren, and G.P. McNally, Behavioral and Neurobiological Mechanisms of Pavlovian and Instrumental Extinction Learning. Physiol Rev, 2021. 101(2): p. 611–681.

      (10) Jenkins, H.M. and R.H. Harrison, Effect of discrimination training on auditory generalization. J Exp Psychol, 1960. 59: p. 246–53.

      (11) Dunsmoor, J.E. and K.S. LaBar, Effects of discrimination training on fear generalization gradients and perceptual classification in humans. Behav Neurosci, 2013. 127(3): p. 350–6.

      (12) Herzog, K., et al., Reducing Generalization of Conditioned Fear: Beneficial Impact of Fear Relevance and Feedback in Discrimination Training. Front Psychol, 2021. 12: p. 665711.

      (13) Lommen, M.J.J., et al., Training discrimination diminishes maladaptive avoidance of innocuous stimuli in a fear conditioning paradigm. PLoS One, 2017. 12(10): p. e0184485.

      (14) Casanova, J.P., et al., Threat-dependent scaling of prelimbic dynamics to enhance fear representation. Neuron, 2024. 112(14): p. 2304–2314 e6.

      (15) Kyriazi, P., D.B. Headley, and D. Pare, Different Multidimensional Representations across the Amygdalo-Prefrontal Network during an Approach-Avoidance Task. Neuron, 2020. 107(4): p. 717–730 e5.

      (16) McKenzie, S. and H. Eichenbaum, Consolidation and reconsolidation: two lives of memories? Neuron, 2011. 71(2): p. 224–33.

      (17) Winocur, G. and M. Moscovitch, Memory transformation and systems consolidation. J Int Neuropsychol Soc, 2011. 17(5): p. 766–80.

      (18) Hockley, A. and M.S. Malmierca, Auditory processing control by the medial prefrontal cortex: A review of the rodent functional organisation. Hear Res, 2024. 443: p. 108954.

      (19) Zikopoulos, B. and H. Barbas, Prefrontal projections to the thalamic reticular nucleus form a unique circuit for attentional mechanisms. J Neurosci, 2006. 26(28): p. 7348–61.

      (20) Phillips, R.G. and J.E. LeDoux, Differential contribution of amygdala and hippocampus to cued and contextual fear conditioning. Behav Neurosci, 1992. 106(2): p. 274–85.

    1. eLife Assessment

      This is a valuable study on the electrophysiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence for decisions. The authors present convincing EEG and behavioural evidence to support their claims. The work will be of interest to cognitive and systems neuroscientists working on decision-making.

    2. Reviewer #1 (Public review):

      Summary:

      This paper characterises the physiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence, with a focus on the centroparietal positivity and motor beta lateralization. The main finding is that the centroparietal positivity builds up during evidence accumulation but falls back to baseline during gaps, while motor beta lateralization maintains a continuous a sustained representation throughout the gap and until response.

      Strengths:

      - Elegant combination of electroencephalography and computational modelling.<br /> - Innovative task design, including parametric manipulation of gap duration.<br /> - The authors describe results of two separate experiments, with very similar results, in effect providing an internal replication.

      Weaknesses:

      - In their response to the reviewers, the authors now include a figure illustrating the relationship between the centroparietal positivity and motor beta lateralisation. However, in the absence of statistical analyses, it remains difficult to draw firm conclusions about this relationship.

      - The paper does not provide an exhaustive characterisation across sensors and frequency bands. However, as the data are publicly available, these questions could be addressed in future work.

    3. Reviewer #2 (Public review):

      Summary:

      This manuscript examines decision-making in a context where the information for the decision is not continuous, but separated by a short temporal gap. The authors use a standard motion direction discrimination task over two discrete dot motion pulses (but unlike previous experiments, fill the gaps in evidence with 0-coherence random dot motion of differently coloured dots). Previous studies using this task (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023) or other discrete sample stimuli (Cheadle et al., 2014; Wyart et al., 2015; Golmohamadian et al., 2025) have shown decision-makers to integrate evidence from multiple samples (although with some flexible weighting on each sample). In this experiment, decision-makers tended not to use the second motion pulse for their decision. This allows the separation of neural signatures of momentary decision-evidence samples from the accumulated decision-evidence. In this context, classic electroencephalography signatures of accumulated decision-evidence (central-parietal positivity) are shown to reflect the momentary decision-evidence samples.

      Strengths:

      The authors present an excellent analysis of the data in support of their findings. In terms of proportion correct, participants show poorer performance than predicted if assuming both evidence samples were integrated perfectly. A regression analysis suggested a weaker weight on the second pulse, and in line with this, the authors show an effect of the order of pulse strength that is reversed compared to previous studies: A stronger second pulse resulted in worse performance than a stronger first pulse (this is in line with the visual condition reported in Golmohamadian et al., 2025). The authors also show smaller changes in electrophysiological signatures of decision-making (central parietal positivity, and lateralised motor beta power) in response to the second pulse. The authors describe these findings with a computational model which allows for early decision-commitment, meaning the second pulse is ignored on the majority of trials. The model-predicted electrophysiological components describe the data well. Some flexible weighting of the second pulse also described the data well (in line with previous studies), but this explanation suffers from additional model complexity. In particular, this analysis of model-predicted electrophysiology is impressive in providing simple and clear predictions for understanding the data.

      Weaknesses:

      Behaviour in this experiment is different from previous experiments which use very similar designs (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023). The authors provide some possible explanations for this in the discussion. Overall performance in this experiment was much worse than previous experiments: Participants achieved ~85% correct following 400 ms of 33 - 45% coherent motion. In previous work, performance was ~90% correct following 240ms of 12.8% coherent motion. A second weakness is that, while bounded model can describe the data in this manuscript, it cannot explain the data from previous experiments showing a stronger weight on the second pulse.

    4. Author response:

      The following is the authors’ response to the previous reviews.

      We have addressed the outstanding points made by the reviewers and provide a detailed description of the additional analyses performed & key results below. We have also updated the manuscript to reflect these additional results, and to contain a more detailed consideration of alternative plausible models.

      We also note that we have corrected one figure panel (Fig.2 panel B, Exp. 2 only), where we identified a small bug in the visualisation code whereby the data of either one or two participants was not correctly plotted in some conditions. This makes no difference to the reported effects.

      Please note that the reviewers acknowledged that your introduction now more broadly refers to the previous work from various groups on motor beta lateralisation (MBL).

      (1) Evaluating the correlation between CPP and MBL, which is key for supporting the claim that CPP is feeding MBL. If, as you are alluding to in your rebuttal, single-trial estimates of CPP are too noisy, trials could be binned based on CPP.

      As requested, we now provide additional analyses binning the data by CPP amplitudes, for the high-low coherence conditions at P1. Full details are provided below. In both experiments we find that, for a given coherence, greater CPP amplitudes at P1 correlate with stronger motor beta lateralisation.

      (2) Examining the possibility of down-weighting (in line with previous studies) compared to your current bounded integration description. Specifically, does the your model predict a bi-modal CPP-P2 distribution that is not evident in the data?

      We have now fit 3 additional models investigating alternative mechanisms that might account for the behavioural results. In particular, we have explored 4 different ways in which flexible weighting of the second pulse might account for both the behavioural and neural data. A full account of the results is provided below, and has been included in the manuscript. In sum, we find that a model which directly and uniformly downweighs evidence from the second pulse (as opposed to indirectly through little or no distance remaining to bound, as in our model) can account for the behavioural data well, but it cannot recapitulate CPP-P2 results unless an accumulation-terminating bound is also included in the model, and the additional complexity of a model with these two free parameters is not supported by model comparison. An alternative model where P2 is downweighed as an inverse function of P1 strength (i.e., stronger downweighing for P1-high coherence pulses) could recapitulate both the behavioural and neural data, but again, this was not favoured by model comparison metrics that account for complexity in the current dataset. We have added a piece on this in the discussion, noting how previous studies mentioned by the reviewer such as Cheadle et al., 2014 and Glickman et al., 2022 find a consistency bias where later evidence is boosted when it agrees with the earlier evidence, opposite to the dampening suggested by the model here, but that a key distinction in our task is that P2 always agreed with P1, so that a dampening might be plausible if subjects tend to withdraw some of their attention from the confirmatory P2 based on the strength of P1.

      Regarding the CPP-P2 distribution, our original bounded model does indeed predict a bimodal CPP-P2 distribution with a peak at 0 arising from the early termination trials, which does not appear in our data (see Fig. S19 and related reply below). However, that EEG noise precludes the detection of any such bimodality in single trial amplitude distributions is demonstrated by the fact a bimodal distribution is strongly predicted for CPP-P1 amplitudes due to the two coherences, most strongly in fact for the unbounded model since there would be nothing to cap the higher-coherence, yet no trace of such bimodality is evident there either, due to EEG noise.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This paper characterises the physiological and computational underpinnings of the accumulation of intermittent glimpses of sensory evidence, with a focus on the centroparietal positivity and motor beta lateralization. The main finding is that the centroparietal positivity builds up during evidence accumulation but falls back to baseline during gaps, while motor beta lateralization maintains a continuous a sustained representation throughout the gap and until response.

      Strengths:

      - Elegant combination of electroencephalography and computational modelling.

      - Innovative task design, including parametric manipulation of gap duration.

      - The authors describe results of two separate experiments, with very similar results, in effect providing an internal replication.

      Weaknesses:

      - A direct characterization of how the centroparietal positivity and motor beta lateralization interact is missing, which limits the novelty. In their reply to reviewers, the authors argue that the signal-to-noise ratio of EEG signals is insufficient for such analyses at the single-trial level. If so, a binned or trial-averaged approach could still be attempted.

      As requested, we have now performed an additional analysis binning trials according to single-trial CPP-P1 amplitudes. To this aim, we sorted trials according to P1 coherence, and median-split them within condition according to the CPP-P1 amplitudes integrated in a time window around the grand-averaged peak [0.4 to 0.6s] after pulse onset, on the same subset of electrodes as in the manuscript. We then plotted motor beta lateralisation (MBL) as the difference in [Contra - Ipsi] hemispheres. Stronger negativities thus indicate stronger lateralisation towards the correct response. In all 4 cases, (both experiments and both coherence levels), higher CPP amplitudes were associated with stronger lateralisation from 0.5s post-pulse onwards (Author response image 1).

      Author response image 1.

      MBL (bottom) traces aligned to P1 onset (time = 0), median split by CPP amplitude [0.4-0.6s] post pulse onset, within P1 coherence condition. Trials with stronger CPPP1 potentials were linked to stronger MBL lateralisation toward the correct response.

      - An exhaustive characterisation of sensors and frequency bands is also missing. In their reply to reviewers, the authors suggest that this would detract from their hypothesis-driven focus. I disagree: the main hypothesis and figures could remain centred on the centroparietal positivity and motor beta lateralization, with a more comprehensive mapping of sensors and frequencies placed in supplementary material. Since the purpose of the paper is to examine EEG-based decision signals in a novel behavioural context, a broader characterisation of the underlying EEG landscape would seem appropriate.

      To broaden our characterisation, we have now included an additional supplementary figure that describes another distinct, relevant EEG signal. Fig. S12 shows the lateralised readiness potential (LRP), a lateralised motor preparation signal that has long been used as an index of relative motor preparation with high temporal resolution (Eimer, 1998; Kelly & O’Connell, 2013; Vidal et al., 2015).The LRP is typically computed as the difference in voltage between [IpsiContra] lateral motor electrodes with respect to eventual response, and it captures the fact that the contralateral motor cortex exhibits more pronounced negative ramps than the ipsilateral one immediately preceding action execution. The EEG landscape characterised in our paper thus comprises four distinct signals that are all functionally relevant to the task, including occipital alpha power, relevant for attention & temporal expectation encoding, which was included both in the main manuscript (Fig. 2) and the supplement (Figs. S9, S11). Given our already extensive supplementary material (18 figures) focused on our main research questions, we feel that a full, hypothesis-free exploration across the dimensions of frequency, space (sensors) and time, considering that there are 60 experimental conditions among which differences may be tested for (2 directions x 3 gaps x 4 coherence pairings in exp 1, plus 2 directions x 4 gaps x 4 coherence pairings in exp 2, plus single-pulse trials), would render the supplemental materials excessive in volume. Again, the data will be shared publicly for future exploration of these many dimensions.

      Reviewer #2 (Public review):

      Summary:

      This manuscript examines decision-making in a context where the information for the decision is not continuous, but separated by a short temporal gap. The authors use a standard motion direction discrimination task over two discrete dot motion pulses (but unlike previous experiments, fill the gaps in evidence with 0-coherence random dot motion of differently coloured dots). Previous studies using this task (Kiani et al., 2013; Tohidi-Moghaddam et al., 2019; Azizi et al., 2021; 2023) or other discrete sample stimuli (Cheadle et al., 2014; Wyart et al., 2015; Golmohamadian et al., 2025) have shown decision-makers to integrate evidence from multiple samples (although with some flexible weighting on each sample). In this experiment, decision-makers tended not to use the second motion pulse for their decision. This allows the separation of neural signatures of momentary decision-evidence samples from the accumulated decision-evidence. In this context, classic electroencephalography signatures of accumulated decision-evidence (central-parietal positivity) are shown to reflect the momentary decision-evidence samples.

      Strengths:

      The authors present an excellent analysis of the data in support of their findings. In terms of proportion correct, participants show poorer performance than predicted if assuming both evidence samples were integrated perfectly. A regression analysis suggested a weaker weight on the second pulse, and in line with this, the authors show an effect of the order of pulse strength that is reversed compared to previous studies: A stronger second pulse resulted in worse performance than a stronger first pulse (this is in line with the visual condition reported in Golmohamadian et al., 2025). The authors also show smaller changes in electrophysiological signatures of decision-making (central parietal positivity, and lateralised motor beta power) in response to the second pulse. The authors describe these findings with a computational model which allows for early decision-commitment, meaning the second pulse is ignored on the majority of trials. The model-predicted electrophysiological components describe the data well. In particular, this analysis of model-predicted electrophysiology is impressive in providing simple and clear predictions for understanding the data.

      Weaknesses:

      Some readers may be left questioning why behaviour in this experiment is so different from previous experiments which use almost exactly the same design (Kiani et al., 2013; TohidiMoghaddam et al., 2019; Azizi et al., 2021; 2023). Overall performance in this experiment was much worse than previous experiments: Participants achieved ~85% correct following 400 ms of 33 - 45% coherent motion. In previous work, performance was ~90% correct following 240ms of 12.8% coherent motion. A second weakness is that, while the authors present a model which describes the data based on pre-mature decision-commitment, they do not examine explanations from the existing literature, that evidence is flexibly weighted, and do not provide any analyses which could be used to compare these descriptions. While their model can describe the data in this manuscript, it cannot explain the data from previous experiments showing a stronger weight on the second pulse.

      The revised version of the manuscript includes a detailed discussion about possible reasons why our stimulus characteristics, task design & experimental protocol may have led to the observed behavioural results (lines 605 onwards). Furthermore, we have now included an extended model comparison as a supplementary note which examines alternative models that could account for the observed data and an additional discussion section that links it to the existing literature.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have responded to each of the comments in the previous review. The manuscript introduction and discussion have been substantially improved, and now more adequately address the previous literature. Limited improvements were made to the analysis, although the authors acknowledged why the suggested improvements from the reviewers were unlikely to be successful, but did not attempt to address the comments using other methods.

      One common theme to both reviews was that, although the model broadly describes the data, it is not fully tested, and alternative descriptions are not fully considered.

      We thank the reviewer for their careful consideration of our data & their reply. We agree it is important to formally test alternative descriptions, most particularly those involving downweighting of the processing of P2, and we have now done so. If participants were simply downweighting P2 by implementing generally smaller drift rates regardless of P1, we would expect the CPP-P2 to exhibit the same classical pattern as in P1, with higher amplitudes following high-coherence P2. The interaction pattern we observe, whereby CPP-P2 amplitudes are systematically lower following high-coherence P1, within each P2 coherence, can only be explained by some dependency between the processes occurring at P1, and those following at P2. What if, as the reviewer suggested originally, “additional evidence from the second pulse was down-weighted according to certainty following the first pulse?” We thus consider this also.

      To formally test this, we have now fit four further models and compare both the fit quality to accuracy data and the predicted EEG results to our original implementation. All models were fit to the grand-averaged accuracy data for all conditions, using 10K simulations, and for both experiments separately (as was done for the bounded model presented in the manuscript). For EEG simulations in unbounded models, we assumed that the CPP signal fell back down to zero upon the dots turning blue, as for the simulations in the main manuscript. Fig. S15 illustrates the fit quality as measured by means of Bayesian Information Criterion (BIC), and we go into more detail on each model in turn below.

      “(1) Unbounded model + P1-independent P2 drift rate reweighting (DR<sub>P2</sub>)

      First, we fit a model with no bound in which the drift rate for P2 was estimated by reweighting (positive or negative) the P1 drift rate via an additional free scaling parameter w. This model thus had the same complexity as our original one (k = 3 free parameters): two drift rate parameters for high and low coherence of P1 (d<sub>high</sub>, d<sub>low</sub>), plus a scaling parameter, w, dictating the strength of P2 relative to same-coherence P1. The drift rate for P2 (d<sub>P2</sub>) was computed as the d*w, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, multiplied by the scaling parameter w, so that if w < 1, P2 drift rates would decrease compared to the same coherence in P1. This model implements the reviewers’ suggestion that systematic downweighting of the second pulse might also account for the results (while also allowing for upweighting (w > 1) for the sake of flexibility).

      This first model (DR<sub>P21</sub>) yielded a similar fit quality to the behavioural data compared to our original bounded model (Bnd), as measured by BIC (Fig. S15). The new model broadly recapitulated the key behavioural results, including order effects, (Fig. S16,A) and could also recapitulate the generally lower CPP-P2 amplitudes (Fig. S16,B). However, this model failed to capture the key coherence-based pattern observed in the CPP-P2 data. Namely, while our EEG results showed that CPP-P2 in trials following P1-low coherence pulses reached overall higher amplitudes than that in trials following P1-high coherence pulses (see manuscript Fig. 3), this model’s simulations predicted that CPP-P2 should scale only with P2 coherence, showing higher amplitudes and steeper build-up rates for P2-high trials, regardless of P1 coherence (Fig. S16C). This is at odds with our empirical results.”

      “(2) Unbounded model + P1-dependent P2 Drift rate reweighting (invDR<sub>P2</sub>)

      Next, we tested a model in which P2 downweighting could depend on P1 strength. That is, we made the P2 drift rate scaling parameter w inversely proportional to P1 coherence so that P2 drift rate d<sub>P2</sub> = d*w/d<sub>P1</sub>, where d equals d<sub>high</sub> or d<sub>low</sub> depending on P2 coherence, and d<sub>P1</sub> indicates the preceding P1 coherence. This implements a kind of certainty weighting, whereby evidence following a strong P1 is more strongly dampened than evidence following a weak P1. This model could recapitulate the key behavioural findings (Fig. S17A), and also qualitatively captured the CPP-P2 effects (i.e. P1-low trials reaching overall higher amplitudes than P1-high trials, Fig. S17C), although the magnitude of this effect was substantially smaller than predicted by the original simple bounded model with no drift rate modulations. However, BICs indicated that this model provided an overall worse fit to the accuracy data compared to the original bounded model (Fig. S15).”

      (3) Bounded models + P2 drift rate reweighting (Bnd + DR<sub>P2</sub>, Bnd+ invDR<sub>P2</sub>)

      Finally, we investigated how well models with both a bound and either of the two P2 drift rate scaling methods we investigated above (uniform downweighting, P1-dependent downweighting) could capture the data, thus effectively testing two extensions of our original implementation that allowed for flexible reweighting of P2.

      The systematic reweighting model with a bound (Bnd + DR<sub>P2</sub>) could recapitulate all key behavioural and EEG findings (Fig. S18), but the additional complexity of the model was not supported by BIC (Fig. S15). Crucially, the reason that this model could recapitulate the CPPP2 results was still the presence of a bound, although we note that the proportion of trials that were predicted to terminate early was reduced in this model compared to the original bounded model presented in the manuscript (c.f. Fig. 4). Yet, this relatively small fraction of trials where accumulation ended early meant that 1) in some trials no accumulation was allowed to occur at all during P2, and 2) where it occurred, the DV was closer to the bound following P1-high coherence trials, thus needing to accumulate less further evidence before reaching a bound and yielding smaller CPP-P2 amplitudes overall in those trials. The inclusion of a bound was thus key to allow a model with systematic P2 downweighting to account for both behavioural and EEG data.

      The model with P1-based scaling of P2 drift rates with a bound could also recapitulate all key behavioural and EEG findings (Fig. S18), but again, the additional complexity of the model was not supported by BICs (Fig. S15).

      “Conclusion

      This extended modelling exercise suggests that 1) a simple bounded model is favoured by model comparison, 2) an alternative model of equal complexity which includes a P1dependent systematic downweighting of P2 rather than a bound can produce qualitatively similar results, the common feature of both viable models being the push-pull relationship between P1 and P2, and 3) more complex models including both a bound and P2 modulations can also account for both the behavioural and EEG results, but the additional complexity is not supported by the current data. While it is possible that some flexible weight modulations occur, these are not sufficiently influential to justify its inclusion in the model. Future work would nevertheless be warranted to explore this possibility in more detail using tailored task paradigms. For the scope of this paper, we have maintained the bounded account in the main manuscript as it is the one supported by the model comparison in the current dataset, but we have also included the alternative P1-based downweighting account as a supplementary figure, along with some additional discussion.”

      In response to my previous comment 3, the authors show their model predicts that there should be no CPP-P2 if the bound is reached before P2, otherwise CPP-P2 is similar to CPPP1 (Figure R2). The argument in the manuscript is that the lower CPP-P2 is because of this bound. The distribution of CPP-P2 amplitudes should therefore have higher variance than CPP-P1 amplitudes, and one might even predict a second mode in the distribution, around 0 amplitude (those trials that terminated before P2). The authors do not show this.

      In Figure R3, it looks like the data have been normalised independently for CPP-P1 and CPPP2 (since the means are approximately the same); normalisation also prevents a comparison of the variance. However, it is apparent that there is no bimodality in the CPP-P2 distribution - were there substantially more trials with 0 CPP-P2 amplitude than CPP-P1? Is the EEG data actually more consistent with a model that systematically downweights P2?

      In the previous Figure R3, data were normalised across CPP-P1 and CPP-P2, not separately. We plot the non-normalised values here (Author response image 2), for comparison, along with median, variance and skewness values for CPP-P1 and CPP-P2. Additionally, we attach the single-participant plots at the bottom of this document (Author response image 3)

      We reanalysed the non-normalised data, excluding outliers (defined as values exceeding the mean +/- 3 times the standard deviation, computed for each pulse & for each participant separately). We found, in both experiments, lower median amplitudes (Exp. 1: t(21) = 2.68, p = 0.013; Exp. 2: t(20) = 4.37, p < 0.001), higher variances (Exp. 1: t(21) = 0.59, p = 0.55; Exp. 2: t(20) = 2.33, p = 0.03) and more positive skewness (Exp. 1: t(21) = 1.79, p = 0.08; skewness: 0.027 vs. 0.69; Exp. 2: t(20) = 3.85, p <0.001; skewness: 0.027 vs. 0.214) in CPP-P2 compared to CPP-P1, although variance and skewness effects were only significant in Exp. 2.

      The reviewer argued in the previous review as well as here that increased variance would be predicted by the model – this is correct, but we believe that the mere presence of higher CPPP2 variances in our empirical data does not, on its own, necessarily support the model. That is because this increased variance could be explained by other factors, such as the increased EEG signal complexity of data at P2 compared to P1. This increased complexity naturally arises from overlapping potentials from CPP-P1 (which can be corrected for, but will increase data noise and thus variability nonetheless), as well as the various gap durations across conditions, which would also affect pre-P2 dynamics. Thus, while our data (partially - in Exp. 2 only) support the reviewer’s interpretation, we would be cautious in using the observation of increased variability as evidence for or against our model given the considerations above.

      Author response image 2.

      A. Empirical CPP–P1 and CPP-P2 amplitude [450-550ms post-pulse] distributions, pooled across coherences. Data were not normalised within-participant. Data were baselined 100ms before pulse onset prior to CPP-P2 amplitude extraction. In both experiments, CPPP2 amplitudes had a lower median (vertical line) amplitude and higher variance than CPP-P1. B. CPP amplitude simulations based on the original bounded model. A high number of trials where CPP-P2 amplitude should equal 0 due to early terminations. C. CPP amplitude simulations based on the inverse weighting model (invDR<sub>P2</sub>). The model did not predict any trials with zero CPP-P2 amplitude because early terminations were not allowed. Rather, the mean distribution shifted towards lower predicted amplitudes, because P2 was downweighted proportionally to P1 coherence.

      The reviewer asks: “Were there substantially more trials with 0 CPP-P2 amplitude than CPPP1?”. Given the noisiness of single-trial EEG data due to high-frequency artifacts, and/or spurious signal drifts (as illustrated by the raw CPP value distributions in Author response image 2A above), it is not be possible to directly detect trials on which CPP = 0. Instead, this must be inferred by other means. If the CPP is in reality at 0 on a larger proportion of trials, then on average, this should manifest as a higher fraction of trials with lower amplitudes, resulting in a more positively skewed distribution of CPP-P2 amplitudes compared to CPP-P1. As reported above, in both experiments we find that skewness is higher in CPP-P2 than in CPP-P1, in line with this hypothesis.

      Regarding the question “Is the EEG data actually more consistent with a model that systematically downweights P2?”, we point to the additional modelling we conducted in response to the comment above. To recapitulate, we find that a model that systematically downweights P2 could account for behavioural findings, but could only recapitulate the key EEG CPP-P2 patterns if an accumulation-ending bound was also included in the model. Instead, a model where P2 is downweighted as a function of P1 strength could qualitatively capture both behavioural and EEG findings, but was not strongly supported by goodness of fit measures in this dataset. We note, however, that the latter model does not predict a bimodal CPP-P2 distribution, but rather a shifted mean and more positive skewness for CPP-P2 trials (Author response image 2C; skewness: CPP-P1 = 1.08; CPP-P2 = 1.24). Thus, in that respect, it does appear to provide a better qualitative recapitulation of the single-trial CPP-P2 data. However, given the extent of EEG noise, the bimodal underlying distribution of our bounded model would also translate to a unimodal, skewed distribution as observed, so this does not provide a strong basis for adjudication. Underscoring this, it is noteworthy that an unbounded model in fact predicts a more separated bimodal distribution for P1 than a bounded model, yet, again, with EEG noise, we are not able to identify any such bimodality in the empirical P1 amplitude distribution.

      Minor:

      In the discussion, the authors write "We also used a narrower range of coherences than the previous studies, which possibly lends itself to calibrating a bound to achieve acceptable accuracy while saving cognitive effort." (Page 22). Perhaps this should be reworded. The range of coherence in this study was ~26-44% in Exp 1 (a difference of 18%) in previous experiments the coherence was 3.2-12.8% (a difference of ~10%). The ranges of performance were similar.

      We agree that the phrasing could be improved. We meant to say that we used only two coherences (high-low) with less than a twofold difference between them, instead of multiple levels of evidence strength in previous studies (e.g. 0,3.2,6.4,12.8) – we have clarified this.

      The authors mention in their rebuttal "However, in contrast to previous studies, we did not include any feedback on a trial-by- trial basis, instead only providing feedback at the end of each block indicating the average accuracy." Actually, Kiani et al., 2013 also only gave feedback at the end of each block. This is also implied in the discussion. I suggest this be removed as the common feedback in Kiani et al., 2013 suggests this cannot explain the difference.

      The methods in Kiani et al. 2013 state “At the end of motion stimulus, a 400–1000 ms delay period (truncated exponential) was imposed before the Go signal, disappearance of the fixation point, was presented. The subject was required to report the net direction of motion within 1 s after the Go signal by pressing a left or right key. Distinctive auditory feedback was delivered for correct and error responses. On trials with 0% coherence, the type of feedback was chosen randomly.“ We understand this means feedback was provided after every trial. If this interpretation is wrong, we would like to kindly ask the reviewer to point us to the relevant methods section so that we can correct the manuscript.

      Author response image 3.

      Individual CPP-P1 (blue) and CPP-P2 (orange), for both experiments (non-z-scored). Vertical lines indicate median CPP amplitudes for each pulse, respectively.

    1. Cover Letter

      I think this is missing a results punchline - TOB is widely used We can observe a bias because of proportions & conversion values And this is important because it does XYZ

    2. the impact of the this work will be significant

      Or,

      "Given over 4,000 published papers have used TOB metrics in the last decade, and it remains an important quantitative outcome measure for perioperative pain research, the potential implications of this bias are significant"

    1. perhaps in a world that istwelve miles long and nine miles wide (the size ofAntigua) twelve years and twelve minutes andtwelve days are all the same

      no thoughts of the "real world" in Antigua + time is warped

    2. he Earthquake; weAntiguans, for I am one, have a great sense ofthings, and the more meaningful the thing, themore meaningless we make it

      library + knowledge

    3. for though you are a tourist on your holiday,what if your heart should miss a few beats? Whatif a blood vessel in your neck should break

      why tourists don't get involved

    4. the thought of what it might be like forsomeone who had to live day in, day out in a placethat suffers constantly from drought, and so has towatch carefully every drop of fresh water used(while at the same time surrounded by a sea andan ocean—the Caribbean Sea on one side, theAtlantic Ocean on the other), must never crossyour mind

      climate change / heat reality + drought + contrast between having the whole sea and needing water

    5. You may be the sort of tourist who wouldwonder why a Prime Minister would want an air-port named after him—why not a school, why not ahospital, why not some great public monument?You are a tourist and you have not yet seen a schoolin Antigua,

      tourists = naive.

    1. L'Inégalité et la Tyrannie du Mérite : Défis pour la Démocratie Contemporaine

      Synthèse Opérationnelle

      Ce document de breffage synthétise les réflexions du philosophe Michael J. Sandel, notamment lors de ses interventions à l'Université de Genève, sur les racines profondes de l'inégalité et la crise actuelle des démocraties.

      Les points saillants sont les suivants :

      • L'inégalité est intrinsèquement morale : Au-delà de l'économie, l'inégalité est une blessure civique liée au manque de reconnaissance et d'estime sociale.

      • La faille de la méritocratie : Bien qu'attrayante en apparence, la méritocratie génère de l'orgueil (hubris) chez les gagnants et du ressentiment chez les perdants, érodant le bien commun.

      • La crise de la reconnaissance : Le mépris ressenti par les classes travailleuses face aux élites diplômées est le principal moteur du retour de bâton populiste.

      • De la justice distributive à la justice contributive : Il est impératif de passer d'une vision de l'individu comme simple consommateur à celle d'un producteur dont la contribution à la société est valorisée, quel que soit son niveau d'études.

      • L'égalité de condition démocratique : La démocratie nécessite des espaces publics communs où les citoyens de toutes origines se côtoient, afin de restaurer la solidarité et le discours public.


      I. Les Deux Sources de l'Inégalité (L'Héritage de Rousseau)

      S'appuyant sur le Discours sur l'origine et les fondements de l'inégalité parmi les hommes de Jean-Jacques Rousseau, l'analyse distingue deux origines à l'inégalité :

      • L'origine économique : Elle naît de l'invention de la propriété privée (« Ceci est à moi »).

      C'est une relation de possession qui engendre crimes, guerres et misères.

      • L'origine morale (la plus profonde) : Elle précède la propriété et naît de la sociabilité.

      Dès que les hommes commencent à s'apprécier mutuellement, ils recherchent l'estime publique.

      La comparaison (qui chante le mieux, qui est le plus fort) engendre la vanité, le mépris, la honte et l'envie.

      Conclusion : La source la plus dévastatrice de l'inégalité n'est pas la disparité de richesse, mais la manière dont nous nous regardons les uns les autres.

      La « politique de l'humiliation » actuelle découle de cette seconde source rousseauiste.


      II. La Tyrannie du Mérite et le Culte du Succès

      Le paradoxe méritocratique

      La méritocratie suggère que, si les chances sont égales, les gagnants méritent leurs gains.

      Or, cette conception a un « côté sombre » :

      • Pour les gagnants : Elle encourage la croyance que leur succès est leur propre fait, oubliant la part de chance et de circonstances favorables.

      Cela mène à une arrogance méritocratique.

      • Pour les perdants : Elle insinue que leur échec est de leur faute.

      Contrairement à une aristocratie où l'on pouvait blâmer le sort, la méritocratie rend l'échec personnel et humiliant.

      L'illusion de la mobilité sociale

      Les données montrent que la mobilité intergénérationnelle est plus faible qu'on ne le pense.

      • Exemple : Il faut 5 générations aux États-Unis ou en Suisse pour qu'un enfant né pauvre atteigne le revenu médian, contre seulement 2 au Danemark.

      • Éducation élitiste : Dans les universités américaines prestigieuses (Harvard, Stanford), il y a plus d'étudiants issus du 1 % le plus riche que de l'ensemble de la moitié inférieure de la population, malgré des politiques d'aide financière généreuses.


      III. Le Backlash Populiste et la Dignité du Travail

      Le soulèvement populiste de 2016 (Brexit, élection américaine) n'était pas seulement une réaction à l'augmentation des inégalités de revenus, mais une révolte contre le mépris des élites.

      • L'insulte du diplôme : En répétant « Ce que vous gagnez dépend de ce que vous apprenez », les élites technocratiques ont envoyé un message insultant à ceux qui n'ont pas de diplôme universitaire : « Si vous luttez, c'est que vous n'avez pas su vous adapter ».

      • Dépréciation sociale : Alors que les revenus des PDG sont passés de 30 fois le salaire moyen dans les années 70 à 300 fois en 2014, l'estime accordée au travail traditionnel a décliné au profit de la gestion financière et des professions technocratiques.

      • Vide moral : La gouvernance technocratique a évacué le jugement moral et politique au profit de l'expertise et des marchés, créant un vide rempli par des formes identitaires dures et un nationalisme intolérant.


      IV. Repenser le Bien Commun et la Justice Contributive

      Vers un changement de paradigme économique

      Il est nécessaire de passer d'une éthique du consommateur à une éthique du producteur.

      | Concept | Vision Consommatrice (Dominante) | Vision Civique / Contributive | | --- | --- | --- | | Bien Commun | Somme des préférences individuelles (PIB). | Réflexion critique sur nos préférences et vie florissante. | | Rôle de l'individu | Satisfaire ses besoins de consommation. | Contribuer à la société et être reconnu pour cela. | | Mesure de la valeur | Le salaire du marché (ex: spéculateur vs infirmier). | La contribution réelle au bien-être de la communauté. |

      La neutralité libérale en question

      Le désir de neutralité (ne pas débattre de la « vie bonne » pour éviter les conflits dans une société pluraliste) a conduit à confier les décisions de valeur aux marchés.

      Cette stratégie d'évitement a échoué.

      Le débat démocratique doit à nouveau aborder des questions morales et spirituelles contestées.


      V. Les Racines Théologiques du Mérite

      Le débat sur le mérite puise ses sources dans la théologie chrétienne :

      • Augustin et Luther : Prônaient le salut par la grâce et non par le mérite (les œuvres), pour préserver l'omnipotence divine contre l'orgueil humain.

      • Calvinisme : Introduit l'idée que le succès dans le travail est un signe de salut (prédestination).

      • Évolution américaine : Ce signe est devenu, au fil du temps, une source de salut.

      Le « self-help » américain transforme la réussite matérielle en preuve de vertu morale, renforçant la tyrannie du mérite.


      VI. Recommandations pour une Égalité de Condition

      Pour restaurer la démocratie, l'accent doit être mis sur l'égalité de condition plutôt que sur la simple égalité des chances.

      • Revaloriser la dignité du travail : Reconnaître que le travail est une source de reconnaissance sociale et de contribution au bien commun.

      • Investir dans l'infrastructure civique : Créer des espaces de mélange des classes (bibliothèques, parcs publics, centres culturels, écoles publiques de qualité).

      • Renforcer les rencontres physiques : La démocratie exige que des personnes issues de milieux différents se croisent dans leur vie quotidienne pour apprendre à négocier leurs différences.

      • Prudence vis-à-vis du Revenu Universel de Base (RUB) : Bien qu'utile comme filet de sécurité, le RUB ne remplace pas le besoin de contribution sociale par le travail.

      Il pourrait même servir d'« achat de paix sociale » par les élites de la Silicon Valley pour éviter de débattre de l'automatisation des emplois.

      • Service National/Civil : Une piste pour favoriser le mélange des classes et cultiver la vertu civique.

      Conclusion finale : La sortie de la tyrannie du mérite nécessite une nouvelle "humilité civique".

      En reconnaissant la part de chance dans notre succès, nous pouvons envisager une vie publique moins rancunière et plus généreuse.

    1. Linear sweep voltammograms (LSV) of a blank solutions (0.25 mol/L NaNO3) in water (grey) and in an alcohol/water mixture (EtOH or MeOH; v/v=3 : 1) both with pH 6.2 and of an alcohol/water mixture (EtOH or MeOH; v/v=3 : 1)

      이그래프 EtOH가 큰이유 보여줌.

    1. Robustesse du Vivant et Sociétés Humaines : Enseignements de la Botanique à l'Ère de l'Anthropocène

      Résumé Exécutif

      Ce document synthétise l'intervention de Mathilde Simon, ethnobotaniste, sur le concept de robustesse par opposition à celui de performance, dans le contexte de l'Anthropocène.

      Alors que nos sociétés modernes privilégient l'optimisation et l'efficience, le monde végétal démontre que la survie à long terme repose sur des stratégies de redondance, de décentralisation et d'acceptation des fluctuations environnementales.

      L'analyse propose un déplacement conceptuel : s'inspirer des mécanismes biologiques pour habiter un monde marqué par la "polycrise" et l'imprévisibilité.

      La pratique de la cueillette sauvage est présentée comme une expérience concrète de cette interdépendance, forçant une confrontation directe avec les limites des ressources et la nécessité d'une éthique de la sobriété.


      1. Le Contexte de l'Anthropocène : Un Monde de Flux et d'Incertitudes

      L'Anthropocène se définit par un état de fluctuation constante et une instabilité découlant des actions humaines.

      Le cadre actuel est marqué par deux concepts clés :

      • La Polycrise : Une multiplication de crises (climatiques, sociales, écologiques) qui ne sont pas indépendantes mais se nourrissent mutuellement.

      Par exemple, une sécheresse entraîne une crise agricole, laquelle génère des tensions sociales.

      L'impact global est supérieur à la somme des crises individuelles.

      • Le Monde des Écarts-types : Citant Olivier Hamant, la conférence souligne la fin de la prévisibilité basée sur des moyennes stables.

      Les données se dispersent, les extrêmes deviennent plus fréquents, et la variabilité même des conditions (températures, précipitations) change, rendant les modèles classiques obsolètes.


      2. Critique de la Performance face à la Robustesse

      Le modèle occidental valorise la performance, définie comme la somme de l'efficacité (atteindre l'objectif) et de l'efficience (avec le moins de moyens possibles).

      Les Écueils de la Performance

      • Fragilité structurelle : L'optimisation maximale élimine les voies alternatives.

      En cas de choc (tempête, maladie, rupture d'approvisionnement), le système s'effondre car il n'a aucune marge de manœuvre.

      • L'exemple de la monoculture : Un système ultra-performant et mécanisé, mais extrêmement vulnérable aux pathogènes et dépendant d'énergies et d'intrants dont l'acheminement est lui-même fragile.

      • Loi de Goodhart : Lorsqu'une mesure devient une cible (ex: se concentrer uniquement sur le rendement), elle cesse d'être une mesure fiable car on occulte la santé globale du système.

      • Invisibilisation des limites : La performance masque les coûts environnementaux et l'épuisement des ressources par l'absence de rétroaction directe.

      Définition de la Robustesse

      À l'inverse, la robustesse est définie comme un "état stable à court terme et viable à long terme malgré les fluctuations". Elle ne cherche pas à contrôler les variables externes, mais à composer avec elles.


      3. Les Piliers Biologiques de la Robustesse Végétale

      Les plantes, organismes fixes présents depuis 400 millions d'années, ont développé des stratégies de persistance plutôt que de maximisation.

      | Pilier de Robustesse | Description et Exemples Biologiques | | --- | --- | | Redondance | Multiplication des organes pour une même fonction. L'arbre a des milliers de feuilles plutôt qu'une seule grande feuille noire (qui serait plus efficiente mais trop fragile). La carotte sauvage produit jusqu'à 800 fleurs par inflorescence. | | Inefficacité / Gaspillage | Le rendement de la photosynthèse est inférieur à 2 %. Le pissenlit produit 1500 graines pour un taux de germination de 0,01 % à 1 %. Ce "gaspillage" nourrit l'écosystème et garantit la survie malgré les prédateurs. | | Coopération Émergente | Les acaraodomacies (structures accueillant des acariens prédateurs sur les feuilles) et les mycorhizes (symbiose racines-champignons) créent des bénéfices mutuels sans intentionnalité, par simple processus évolutif. | | Gouvernance Distribuée | Les réseaux de champignons relient les arbres entre eux sans centre de décision. Les échanges se font par gradients de concentration locaux, rendant le réseau résilient même si une partie disparaît. | | Dominance Relative | Le bourgeon apical (sommet) inhibe les autres, mais si celui-ci est mangé, des bourgeons secondaires déjà formés sont prêts à prendre le relais en quelques jours. |


      4. La Cueillette Sauvage : Une École de la Limite

      La pratique de la cueillette sauvage n'est pas présentée comme une solution alimentaire globale (impossible à l'échelle de la population actuelle), mais comme un cheminement philosophique et politique.

      • Confrontation aux limites : Contrairement aux rayons d'un supermarché toujours réapprovisionnés, le cueilleur voit immédiatement l'impact de son prélèvement.

      Il dépend du cycle naturel qu'il ne peut accélérer.- Décentrement (Non-sencience) : Les plantes ne ressentent pas d'émotions (sencience) mais perçoivent des signaux physiques.

      Apprendre à respecter un vivant si différent de nous favorise l'humilité.

      • Éveil de l'attention : Passer d'un "paysage vert informe" à une distinction fine des espèces (nervures, odeurs, textures) crée une connexion profonde.

      L'étude montre que cette connexion favorise les comportements pro-environnementaux.

      • Engagement politique : L'attachement à un "territoire vécu" via la cueillette incite à s'opposer à son urbanisation et à s'intéresser aux décisions locales.

      5. Perspectives Sociales et Éducatives

      L'analyse conclut sur la nécessité de réintégrer la biologie et les principes de robustesse dans l'organisation humaine :

      • Réhabiliter le "Sauvage" en ville : Accepter des espaces non destinés à l'usage humain (ronces, orties) pour favoriser la biodiversité réelle plutôt que l'ornemental stérile.

      • Éducation à la biologie : Enseigner trois piliers essentiels :

        • Le fonctionnement systémique des écosystèmes et leurs rétroactions.
      • La microbiologie (cellules, virus) pour les enjeux de santé.

      • Les bases de l'alimentation et de l'agriculture.

      • Management et organisations : Questionner le prix réel de l'efficacité (dommages collatéraux humains et écologiques) et réintroduire de la redondance et de la lenteur pour assurer la viabilité à long terme des structures sociales.

      Citation Clé : « Le monde végétal est robuste avant d'être performant. Les systèmes intègrent les fluctuations plutôt qu'ils ne les contrôlent. »

    Annotators

    Annotators

    1. max_num_batched_tokens is the per-iteration token budget: the most tokens the scheduler may place in one forward pass across every request in the batch. It is not a per-request sequence limit. Dividing it by concurrency gives the tokens available per in-flight request, which is the quantity the error tracks.

      add a sentence to explain why this metric is relevant for the current discussion.

    2. Tensor parallelism has to divide both the attention head count and the embedding width, pipeline depth has to divide the layer count, the configuration has to fit inside one node, and GPUs per replica times replica count has to stay within the GPU budget you are willing to provision. Integer arithmetic on the model shape, so effectively free. Removes 3,600 of 11,520, leaving 7,920. One consequence surfaces later: when tensor parallelism exceeds the number of key/value heads, those heads are replicated rather than split, so past that point the KV cache stops shrinking as you add GPUs.

      make it clear at the start that this is just one example of how invalid options are removed.

    3. Figure 4. The seven sieves on the 11,520-configuration grid from the table above. A solid bar means the set shrank at that stage; a dashed outline means the same set with more work done, which is why stages 3 and 5 repeat a count. The cost hierarchy is the point: stages 1 and 2 are arithmetic and run on all 11,520; stage 4’s analytical estimate is cheap enough to run on all 7,840 survivors;

      above diagram has 8 horizontal bars. and 7 sieves. simplify the diagram. it's not very clear.

    4. Instead we filter in stages, cheapest first. The early stages are nearly free and run on every configuration, dropping the ones that can’t work. Only the 20 selected candidates reach the detailed simulation. The early stages are set algebra over model and device facts, so they are orders of magnitude cheaper per candidate than a scheduler replay; we report the candidate counts each stage admits and rejects, not per-stage timings.

      change this part based on changes in the previous para to make it flow

    5. The most detailed simulator replays the serving engine’s scheduler step by step, and even at a few seconds per configuration that would still take hours across the full grid.

      this sentence breaks the flow. doesn't make sense here

    1. eLife Assessment

      This important study measures the development of hearing in the zebra finch and shows that two-day-old hatchlings have no detectable electrophysiological response to 95 dB SPL clicks. Because this stimulus carries energy across a broad range of sound frequencies and is some 60 dB louder than faint, high-pitched parental calls, the findings challenge prior proposals that such calls serve as a channel for heat-dependent prenatal developmental programming. Two reviewers found the evidence on the absence of auditory responses convincing, supported by auditory brainstem recordings using validated methods with appropriate controls and by a separate laser vibrometry experiment ruling out detection through egg vibration. A third reviewer disputed these findings and questioned whether absent physiological markers of hearing definitively establish deafness. On the evidence presented, embryonic hearing is unlikely to mediate the developmental effects reported for heat whistle playback. Whether those effects reflect some other pathway or require re-interrogation of the behavioral phenomenon is now the open question, alongside a need for auditory measurements more sensitive than the auditory brainstem response.

    2. Reviewer #1 (Public review):

      This work by Antonnen et al. was triggered by claims of auditory-mediated effects on altricial avian embryos which were published without any direct evidence that the relevant parental vocalizations were actually heard. I agree with Anttonen et al. that, based on the available evidence about avian auditory development, those claims are highly speculative and therefore necessitate more direct experimental verification.

      Attonen et al. have embarked on a comprehensive series of experiments to

      (1) Better characterize acoustically the relevant parental vocalizations (heat whistles; in a separate preprint, not reviewed here)

      (2) Characterize the auditory sensitivity of zebra finches at various stages of their posthatching development. Despite the long-standing importance of the zebra finch as a songbird model in neuroethology of learned vocalizations, the auditory development of the species had not been studied so far.

      (3) Explore an alternative hypothesis of how the parental vocalizations might be perceived.

      The principal method used here is the non-invasive recording of ABR (auditory brainstem response), a standard neurophysiological method in auditory research. The click-evoked ABR provides a quick and objective assessment of basic hearing sensitivity that does not require animal training. Weaknesses of the technique include its limited frequency specificity and low signal-to-noise ratio. The authors are experienced with ABR measurements and well aware of those issues. ABR responses in zebra finches are shown to gradually appear during the first week posthatching and to mature in subsequent weeks, consistent with the auditory development in other altricial bird species studied previously. When matching the acoustic properties of parental heat whistles and auditory sensitivities, hearing of the parental heat whistles by zebra finch hatchlings was convincingly excluded. Although not directly measured, this also convincingly extrapolates to zebra finch embryos. Finally, the authors tested the hypothesis that parental heat whistles could induce perceptible vibrations of the egg and thus stimulate the embryo via a different modality. The method used here was laser doppler vibrometry, an appropriate, state-of-the-art technique that the authors also have proven experience with. The induced vibrations were shown to be several orders of magnitude below known vibrotactile sensitivities in mammals and birds. Thus, although zebra finch vibrotactile thresholds were not obtained directly, the hypothesis of vibrotactile perception of parental heat whistles by zebra finch embryos could also be rejected convincingly.

      In summary, even when considering some weaknesses of the techniques (which the authors are aware of), the conclusions of the paper are well supported: Auditory and/or vibration perception of parental heat whistles can be excluded as an explanation for previous reports of developmental programming for high ambient temperatures. As a constructive suggestion towards resolving the apparent paradox, the authors recommend to repeat some of the crucial, previous playback experiments at lower sound levels that better match the natural parental vocalizations.

    3. Reviewer #2 (Public review):

      This study by Anttonen, Christensen-Dalsgaard and Elemans describes the development of hearing thresholds in an altricial songbird species, the zebra finch. The results are very clear and along what might have been expected for altricial birds: at hatch (2 days post-hatch), the chicks are functionally deaf. Auditory evoked activity in the form of auditory brainstem responses (ABR) can start to be detected at 4 days post-hatch but only at very loud sound levels. The study also shows that ABR response matures rapidly and reaches adult like properties around 25 days post-hatch. The functional development of the auditory system is also frequency dependent with a low to high frequency time course. All experiments are very well performed. The careful study throughout development and with the use of multiple time-points early in development is important to further ensure that the negative results found right after hatching are not the result of the experimental manipulation. The results themselves could be classified as somewhat descriptive but, as the authors point out, they are particularly relevant and timely. Since 2016, there has been a series of studies published in high profile journals that have presumably showed the importance of prenatal acoustic communication in altricial birds, mostly in zebra finches. This early acoustic communication would serve various adaptive functions. Although acoustic communication between embryos in the egg and parents has been shown in precocial birds (and crocodiles), finding an important function for pre-natal communication in altricial birds came as a surprise. Unfortunately, none of those studies performed a careful assessment of the chicks' hearing abilities. This is done here, and the results are clear: zebra finches at 2 and 6 days post hatch are functionally deaf. Since, it is highly improbably that the hearing in the egg is more developed than at birth, one can only conclude that zebra finches in the egg (or at birth) cannot hear the heat whistles. The paper also ruled out the detection on egg vibrations as an alternative path. The prior literature will have to be corrected, or further studies conducted to solve the discrepancies. For this purpose, the "companion" paper on bioRxiv that studies the bioacoustical properties of heat calls from the same group will be particularly useful. Researchers from different groups will be able to precisely compare their stimuli.

      Beyond the quality of the experiments, I also found that the paper was very well written. The introduction was particularly clear and complete (yet concise).

      Weaknesses:

      My only minor criticism is that you don't discuss potential differences between behavioral audiograms and ABRs. Optimally, one would need to repeat the work of Okanoya and Dooling with your set up and using the same calibration. The ~20dB difference might be real or it might be due to SPL measured with different instruments, at different distances, etc. Either way, you could add a sentence in the discussion that states that even with the 20 dB difference in audiogram heat whistles would not be detected during the early days post-hatch, but that adding a (novel) behavioral assay in young birds could further resolve the issue.

      More Minor Points.

      (1) As mentioned in the main text, the duration of pips (form pips to bursts) affects the effective bandwidth of the stimulus. I believe that you could give an estimate of this effective bandwidth given what is know from bird auditory filters. I think that this estimate could be useful to compare to the effective bandwidth of the heat-call which you can now also estimate.

      (2) Fig 5b. label the green and pink areas as song and heat-call spectrum. Also note that in the legend you say: "Green and red areas display the frequency windows related to the best hearing sensitivity of zebra finches and to heat calls, respectively". I don't think this is what you meant. I agree that 1-4 kHz is the best frequency sensitivity of zebra finches but you probably meant green == "song frequency spectrum" and pink == "heat call spectrum". In either case the figure and the legend need clarification.

      (3) Fig 5c. Here also, I would change the song and heat-call labels to "song spectrum", "heat call spectrum". You don't want readers to think you used song and heat calls in these experiments (maybe next time?). For the same reason, maybe in 5a you could add a cartoon of the oscillogram of a frequency sweep next to your speaker.

      (4) Methods. In your description of the stimulus, you describe "5ms long tone bursts" but these are the tone pips in the main part of the manuscript. Use the same terms.

      Comments on revisions:

      In the latest version of this manuscript, Anntonen et al have diligently addressed all the issues raised by the reviewers. As I mentioned in our discussions among reviewers, it is impossible to "prove" a null hypothesis and both methods and experimental design could always be improved. At this stage however, they provide convincing evidence that the sound intensity of heat calls are below the hearing thresholds of zebra finch chicks.

    4. Reviewer #3 (Public review):

      Summary

      This study aims to contest recent findings that prenatal exposure to natural sounds and anthropogenic noise before hatching affects development and fitness in an altricial songbird. To this aim, it attempts to estimate hearing capacities of zebra finch nestlings. It first uses responses to clicks in nestlings and adults to estimate differences in hearing thresholds but uses experimental parameters that systematically lower nestling responses. It then measures responses to tones but restrict nestling data to a protocol (long tones) that failed in adults; response to tones with a correct protocol (short tones) is repeated in adults only. Thirdly, it includes data on loud airborne sound making eggs vibrate, even though this has no relevance to embryonic vibration perception by direct contact with the incubating parent or in incubators. Lastly, it includes a lengthy discussion on how these results, even though inaccurate or incorrect, would show that zebra finch nestlings and embryos are "functionally deaf". It fails to note that even if the experiment had been performed correctly and still failed to detect an auditory response in 2 day hatchlings - which is unlikely given the above and findings in other songbirds - this would not somehow eliminate the developmental effects of prenatal sounds that have been empirically demonstrated.

      Strength:

      The study is not performed adequately to bring reliable answers, but it addresses the important topic of hearing development in altricial songbirds and the long-held (but untested) assumption that altricial avian embryos cannot hear. More broadly, there is a need to reassess avian auditory perception - and the methodological approaches to measure it - given the accumulating evidence that some bird species respond to high frequency biological sounds beyond their known hearing range.

      Weaknesses:

      The study presents many experimental flaws that specifically compromise response detection in immature animals in the first experiment, and data from the remaining two experiments are invalid. The revision did not fix any of these issues.

      i) Response to clicks, Fig 1: Unlike what the revision claims, the deviations from validated protocols (too rapid stimulation, too few measurements, low temperature, reliance on clicks only) do lead to a potentially large underestimation of nestling hearing sensitivity. Calculation and new data trying to show that these deviations would not matter are wrong or insufficient. Even without strict conventions on ABR methodology, it is striking that every single experimental choice made here i) greatly differs from other avian studies and ii) reduces - rather than maximises - response detection, specifically for low amplitude signals (as in nestlings or near thresholds).

      ii) Response to frequency tones, Fig 2: The experiment on frequency sensitivity (long tone burst) failed in adults (positive control) with 0 to only about half of the adults responding (Fig S1), and sensitivity underestimated by 30dB (Fig2). Measures with a failed positive control are invalid, even if the overall pattern of the few data points obtained in nestlings vaguely resemble (but does not match) expectations. That the experiment was repeated in adults with a correct protocol (short tone pips) - but not in nestlings - is highly misleading.

      iii) Attempts to validate the protocols above by comparing to published adult values are meaningless (response A.3.1; Fig 5). The one experiment being compared (short tones) was only performed in adults in this study. This cannot validate results obtained by different methodologies in nestlings, for responses to clicks or long tones.

      iv) Vibrations: No previous results on the effect of zebra finch heat calls on development rely on the hypothesis or assumption that airborne sound from heat calls makes eggs vibrate. This idea is solely attributable to the authors of the preprint, and is not biologically or experimentally realistic. All speculations from this experiment (Fig3) on vibration perception by embryos or heat call effects are meaningless.

      v) Writing and presentation: The text throughout is highly misinformative for non-specialists or any reader not carefully inspecting figures (including in supplementary material) and methods. The confusion is greatly aggravated in the (6-page long) discussion which i) fails to recognise and account for the study shortcomings, and instead ii) greatly overstates the results and what they mean, and iii) misrepresents current knowledge by excluding highly relevant studies showing evidence of early sound perception in embryos and hatchlings, and introducing many errors in the presentation of published papers. This creates an illusion of a strong mismatch between heat-calls and zebra finch sensory capacities, for which the study actually does not provide any evidence for, and which is extremely unlikely to exist.

      vi) The revision did not address any of the major flaws of the study outlined above (see detailed assessment below). In particular,<br /> ** For point i):<br /> - estimations for the effect of having done too few (400) measurements are wrong - the effect on hearing thresholds cannot be calculated, but would be much greater than 4dB with the expected 37% noise reduction with the standard 1000 sweeps;

      - the new data provided does not measure the impact of high stimulus rate, and measures on adults largely underestimate effects on nestling response;

      - body temperature, now provided in the revision, is at the lowest extreme for the species, which may increase hearing thresholds;

      - tones do elicit a stronger and earlier response than clicks. Whether this is related to stimulus duration does not change the fact that clicks underestimate hearing thresholds, and delay hearing onset by several days.<br /> ** For points ii) and iv):<br /> No new data or any valid explanation was provided in the revision. It is still the case that nestling frequency responses were obtained with an experiment where the positive control failed; and the vibrating egg experiment is irrelevant to vibration perception in embryos and any observed effects of heat calls on development.<br /> ** The revision greatly lengthens the speculations about heat call perception, based on inaccurate or totally incorrect data.

      Conclusion and impact:<br /> i) Overall, the study fails to provide any reliable estimate of zebra finch nestling hearing capacities: it underestimates hearing onsets and thresholds (clicks) and gives no information on nestling frequency sensitivity. It is extremely likely that a better designed ABR experiment, or a more sensitive methodology (e.g. electrophysiology), would have detected a response in 2 day-old nestlings, as in other songbirds. The conclusion of "deafness" in hatchlings (or embryos, which were not tested) is clearly unsupported.<br /> If zebra finch hatchlings were clearly deaf, a well-designed study would have shown this a lot more convincingly.

      ii) The data on adults in not new (3 prior studies) - although this study generally underestimates zebra finch high frequency sensitivity. The presentation in relation to heat-calls is flawed: true values of heat-call frequency range and sound levels (>6kHz at 45dB) fall within the adult hearing range.

      iii) Without any new evidence, this study does not progress our understanding of heat-call and noise impact on development. Exactly the same issue remains: that heat-call and noise effects on development contradict the general VIEW on hearing ontogeny in altricial birds. But we have learnt nothing from this study on hearing ontogeny or frequency sensitivity. The study provides no reliable neuroscience data to advance the debate.

      iv) The largest section of the preprint is a highly speculative discussion, based on erroneous data and wrong interpretations, as well as a misrepresentation of what is known. That these issues would not be recognised is - in my view - of serious concern.

      Detailed assessment:

      The summary by the authors of my major concerns (R3.A0) is incorrect and misreport many of my statements, without giving any meaningful answers. I will not engage in such discussion.

      My actual three major concerns do remain:

      (1) Results on Fig 1 overestimate hearing onset and thresholds: All experimental parameters chosen to test responses to clicks are known to lower detection. From their cumulative impact, there is absolutely no doubt that the data in Fig 1 underestimate hearing capacities, and disproportionately so in nestlings compared to adults. Therefore, no quantitative estimate of nestling hearing threshold, absolute or relative (i.e. the estimated "54dB difference"), can be taken from this study. The age of hearing onset based on clicks (4 day old) is also wrong.

      (2) All results on Fig 2 are false: 40 to 100% of adults (positive control) failed to respond within their normal hearing range (Fig S1). Data on nestling frequency sensitivity using this methodology (Fig 2) are clearly invalid.

      (3) All results on Fig 3 are irrelevant: they assume zebra finch parents are hoovering in front of their nest while heat-calling rather than incubating their eggs, which is nonsense. Any speculation on the role of bone-conduction or vibrotactile stimulation for embryonic heat-call detection is simply unfounded.

      The authors' responses to these comments are largely mistaken (see below), and do not change the 3 facts stated above. Overall, the data produced are unreliable, and so are the interpretations and conclusions.<br /> The conclusions that i) heat-calls fall outside of adult hearing range and ii) young nestlings (and embryos) are deaf, are both incorrect. The study provides no actual estimate of nestling hearing and how far off their hearing range heat-calls fall.

      (1) Responses to clicks underestimate hearing abilities, specifically in nestlings.

      (i) DISPROPORTIONATE EFFECT OF PROTOCOL ON NESTLINGS: As reported in other avian studies (e.g. Brittan Powell et al 2004), because of their immature neural system, nestlings show low amplitude waveforms compared to adults. This occurs even with sound much louder than their hearing threshold. For example here, 8 day old nestlings show waveforms of only low amplitude at 95dB, even though they respond to sound 40dB softer, at 55dB (Fig 1H).

      This characteristic of nestling waveforms means that:<br /> - nestling responses are harder to detect,<br /> - nestling responses are more easily attenuated below the detection criteria (>2 S/N),<br /> - a low amplitude response in nestlings at a given sound level does not predict that softer sounds will not be perceived by the animal.

      Any deviation in protocol that attenuates waveform amplitude will therefore disproportionately affect nestling thresholds (and artificially lead to the "54dB" estimate).<br /> Even without universal standards (resp R3.A1.1), knowing this, the protocol should be adjusted to maximise response detection. This study does exactly the opposite.

      (ii) INSUFFICIENT MEASUREMENTS: using only 400 sweeps, rather than the typical 1000 sweeps, reduces response detection, especially in nestlings.

      - The claim that using 1000 sweeps would only decrease the hearing threshold by 4dB (response A3.2; ms L530) is false:<br /> - Based on the square root relationship mentioned by the authors, using 1000 averages instead of 400 would improve the signal-to-noise ratio by 37%, which would allow detecting many small amplitude waves which are currently hidden in the abnormally high noise.<br /> - Lowering the noise floor level by 4dB does not mean that the hearing threshold would only decrease by 4dB. The improvement in hearing threshold would be much greater, especially in nestlings.

      (iii) UNSUITABLY HIGH STIMULATION RATE: the new data on pairs of clicks greatly underestimate the attenuation by high stimulation rates, and so do adult measurements compared to nestlings'.<br /> - the click rate used in this study (25 clicks per second) is 5 to 25 times faster than most previous studies in young birds (e.g. Saunders et al 1973, 1974: 1 stim/sec or less; Katayama 1985: 3.3 clicks/sec; Brittan-Powell et al 2004: 4 stim/sec), and 6 times faster than zebra finch heat-calls.<br /> - responses to pairs of clicks (new data in Fig S3; L 543-552, response A3.3) does not measure the response dampening caused by high stimulation rates. Response attenuation (adaptation) after a single click is much weaker than after a train of 400 consecutive clicks (or even just 30). This paired-click measure ignores the cumulative attenuation observed with many consecutive stimuli. Paired-click paradigm is used to measure immediate refractoriness for other purposes (e.g. temporal resolution, diagnosis tool) but does not replicate high stimulation rates.<br /> It is unclear why the authors chose to use this weak approximation rather than simply replicating the measurements at the same slow rate as in other avian studies. This would have given a straightforward answer, comparable to previous studies that have demonstrated the detrimental impact of fast rate on response strength by directly comparing responses to different stimulation rates (e.g. Saunders et al 1973; Brittan-Powell et al 2004).<br /> - Adults are less sensitive to high stimulation rates than immature individuals (e.g. Saunders et al 1973; Khayutin 1985; Brittan-Powell et al 2004 and 6 references cited therein). Effects of fast rate in adults (new data in Fig S3) therefore largely underestimate effects on nestlings (not measured), and the high stimulation rate used in this study increases the relative difference between nestlings and adults.<br /> - Given the 2 points above, it is incorrect to conclude from this new data that the high stimulation rate used had "no effect on ABR amplitude or thresholds" (responses A3.3 L551). Instead, given that 40ms (i.e. interval for 25 clicks/ sec) is at the limit of what adults can handle after a single click, this new data does confirm that the stimulation rate is indeed too high and underestimates hearing in nestlings (and adults to a lesser extent).

      (iv) LOW BODY TEMPERATURE can reduce ABR response.<br /> - The average body temperature (39.5C) now provided in the revised manuscript (response A2) is at the lowest extreme of the range for zebra finches. The normal average body temperature for zebra finches is 41C at low ambient temperature, rising to 43-44C at high ambient temperatures (when heat-calls would be produced). This value of 41C is consistent across many studies in wild and domestic zebra finches (e.g. Wojciechowski et al 2020 [avr 41C at 23C]; Udino and Mariette 2022 [avg 41, min=40C at 32C]; Bech and Midtgard 1981 [avr ~41C]; Pessato et al 2022 avg=41C at 27C]). The only study reporting a body temperature average as low as 39C (Cooper et al 2020) was indeed under hypothermic conditions, obtained in adult zebra finches under extreme fasting conditions (17hrs of food deprivation), at ambient temperature below thermoneutrality, during the night (i.e. during "nocturnal hypothermia", when bird body temperature is normally lower).<br /> Therefore 39.5C does corresponds to hypothermic conditions for most individuals. While this body temperature remains considerably higher than that used experimentally to supress hearing response, using a below-normal body temperature, added to other factors, may lower wave amplitude and therefore increase hearing threshold estimates.

      (v) CLICKS UNDERESTIMATE HEARING ONSETS<br /> - My statement that, compared to tones, clicks elicit a smaller response, at a later age, is correct (Saunders et al 1973; Brittan-Powell et al 2004).<br /> - It does not matter whether this is due to differences in stimulation duration (response R3.A2.1), since clicks are always much shorter than tones. What matters is that hearing onset based on responses to clicks has been found to underestimate the earliest age at which an auditory response can be obtained by several days (Saunders et al 1973).<br /> - Without any valid nestling data for tones (see below), this study, based on clicks only, does not provide the true age of hearing onset.<br /> - The suggestion of using 170dB clicks (response R3.A2.1) is very odd. The correct approach would have been to use a correct protocol for tones in nestlings, rather than solely correcting it for adults (see below).

      (vi) CONFUSING INTERPRETATION OF WAVE AMPLITUDE AND LATENCY<br /> - the text implies throughout that nestlings lacking an adult-like amplitude and latency is a sign of poor hearing (e.g. L122 "ABR wave I gradually reaches maturity 25 days after hatching").<br /> - It should be clarified that this is a characteristic of immature systems but does mean these nestlings do not hear. For example, this characteristic persists even in 10d old nestlings which have similar hearing thresholds to adults.

      (2). Results on nestling frequency sensitivity are invalid. Using long (25ms) tones with subdermal electrodes is incorrect, and not used by anyone other than the authors. The failure to repeat the experiment with a correct protocol (5ms tones) in nestlings is misleading:

      (i) When 60% of 4-day old chicks respond to clicks (Fig 1), they are described as "functionally deaf" (L78, L261). By contrast, when 0% to 60% of adults respond to tones in the core of their hearing range (Fig S1 for data shown in Fig 2), this is merely presented as a methodological limitation (text added L186-197), and the authors still proceed to presenting results on nestlings with a method that failed in adults (positive control).

      (ii) When the positive control fails, the experiment is failed. This principle applies in Neuroscience as in any scientific field.

      (iii) That the experiment repeated in adults with the standard and correct short tone protocol (5ms) gave normal results, is not a validation. Instead, it proves the point that 25ms is unsuitable (regardless of whether that is due to rising time, L399). This correction does NOTHING to fix results in nestlings, which were ONLY done with the unsuitable 25ms tones.

      (iv) The mixture of adult data with 5 and 25ms tones, instead, creates confusion, because differences in protocols between nestlings and adults are blurred in the text. The abnormally small proportion of adults responding (0 to 60%) with 25ms tones is only shown in supplementary material, and attenuated in the main text (L186: "about half").

      (v) The claim (response R3.A1.2) that long 25ms tones is commonly used with the ABR methodology used here is false:<br /> ** ALL (but 1) avian studies using tones of 20ms or longer (cited by myself or by the authors in their reply R3.A1.2) used a different ABR methodology, with electrodes implanted through the skull. The only exception, using 25ms tones with subdermal electrodes (as here), is a study by the authors themselves.<br /> ** All other avian studies with subdermal electrodes used short tones, of 5ms or less (e.g. Brittan-Powell et al 2002, 2004; Henry & Lucas 2008), as in the corrected adult experiment.<br /> ** My initial comment already specified this difference in electrode placement.<br /> ** The results here (failing in adults) unquestionably show that the method of long tones with subdermal electrodes does not work, and the literature shows that no one else uses it.

      (vi) When the aim of the study is to contest effects of high frequency sounds (L39-45), it is odd to choose a protocol specifically directed at testing responses to very low frequencies (responses A3.1, R3A.0; L398-403) and that compromises all results.

      (vii) That the shape of the few data points obtained in nestlings with long tones would broadly resemble expectations (responses A3.1, R3A.0) is not a validation: See ii.<br /> - The shape is not even correct:<br /> ** Why aren't 4 and 6 day old nestlings responding to any frequency when they were responding to clicks (in spite of poor detection conditions for clicks)?<br /> *** Why are they not responding, when 2 day old flycatchers respond to tones from 1 to 4kHz at 45dB, and 0.5 to 5kHz at 60dB (Aleksandrov and Dmitrieva 1992, Korneeva et al 2006)?

      (viii) That only relative measures between nestlings and adults matter (resp R3.A4; L404) is incorrect. Measures that are qualitatively (see vii) and quantitatively (see i) wrong cannot be compared.<br /> If these data with long tones were indeed acceptable, why repeat the experiment in adults with a correct protocol, even before this flaw was pointed to during the review?

      (ix) The unusually low number of measurements (400 instead of 1000; L530) will have reduced detectability of low amplitude responses, in nestlings and near thresholds, here as in the click experiment (see above). The claim that this would only increase threshold by 4dB (responses A3.2, L530) is wrong (see above).

      (x) There is no explanation as to why only half of the 6-day old nestlings (n=5-6) were tested with tones. Would having 10 in this group shown a response at 6-day old?

      (xi) Even the correct 5ms tone protocol in adults overestimated thresholds at higher frequencies (>3kHz) by up to 30dB compared to other published estimates (Fig 5). If inter-population differences among domestic zebra finches (L 393; Fig 5 legend; resp R3.A4) were enough to cause 30dB differences, results on domestic zebra finches here should not be extrapolated to those on wild-derived Australian zebra finches documenting developmental effects of heat calls.

      The authors give no other elements than the above in their responses that could demonstrate the validity of their nestling frequency data.

      Based on other studies, we can expect zebra finch hearing to develop sensitivity to high frequencies after that to middle frequencies. But pretending that this study provides any evidence towards this is wrong. There is no valid data.

      (3) The experiment using loud sounds to make eggs vibrate is biologically and experimentally meaningless. The claim that these measurements would rule out vibration perception in embryos is totally unfounded. The preprint is creating the illusion of having ruled-out a mechanism that they have not tested.

      (i) zebra finches are not hovering in front of their nest when heat-calling. They are in physical contact with the eggs while incubating, with vibrations expected to travel directly through solids from adults to eggs, without attenuation.

      (ii) the claim that previous studies on heat-call developmental impact rely on the assumption of airborne sound making egg vibrate (response A3.5, R3.A8) is incorrect. This is conceptually wrong and there is no indication of this in any paper.

      (iii) vibration perception in embryos in birds and other taxa (e.g. amphibian, reptiles and insects) rely on direct contact with the source, not on loud airborne sounds shaking eggs. Beyond any consideration on heat-calls, claiming that this experiment could give any indication of vibration perception in embryos is nonsense.

      The authors give no information in their responses that could demonstrate the validity of this experiment.

      (4) Interpretation and discussion

      All other songbird studies, individually and collectively, show much greater hearing sensitivity in nestlings that what this study is trying to suggest (e.g Khayutin 1985; Korneeva et al 2006; Aleksandrov and Dmitrieva 1992; Rivera et al 2018; Platzen and Magrath 2004; Haff & Magrath 2012). The only way the authors can reach their conclusion is by:<br /> - presenting data that greatly underestimate hearing sensitivity in zebra finches (see sections 1 and 2) and overstating them, and;<br /> - misrepresenting current knowledge by excluding relevant papers and being unclear about what the literature shows.<br /> This study tries to impose the idea that zebra finch do not detect any sound before day 4-6 post-hatch, and have extremely rudimental hearing until day 8-10 post-hatch. By contrast, other studies show very young hatchlings respond to sound 1 to 3 days after hatching (first age tested) and show quite sophisticated, and totally functional, responses to relevant sounds at 5 days old.

      This is not to say that hearing does not continue to improve post-hatch in birds, or that sensitivity to mid frequencies post-hatch would not precede that to high frequencies. But statements throughout this study are so exaggerated and/or wrong that nothing can be learnt. It is undeniable that this study fails to provide the useful and balanced assessment needed to establish were true heat-calls actually sit relative to zebra finch adult, nestling and embryonic hearing range.

      It is literally impossible to correct every wrong statement in the discussion and responses to reviewers. I focus here on some examples related the claims of "nestling deafness" or of heat-calls being outside of adult zebra finch hearing range, as well as inaccuracies leading to an apparent match with current literature.

      (i) MISALIGNMENT WITH OTHER STUDIES AND MISREPORTING

      *** Neurological evidence

      - Highly relevant evidence on response to sound in zebra finch embryos (Rivera et al<br /> 2018) is totally excluded.<br /> Excluding this study on the basis that it used a different methodology that does not directly quantifying auditory sensitivity (resp R3.A6a) is a poor justification. Evidence, even indirect or imperfect, should be brought to the attention of the readers.

      - This applies to many other studies cited in my first review and arbitrarily excluded here. If one wants to conclude on "deafness", all evidence, even indirect, should be considered. Excluding non-ABR studies means the conclusion cannot be extended beyond flat ABR traces.

      - Other studies show much greater sensitivity that what the text describes:

      * Flycatcher hatchlings at 2-3d post hatch (first age tested), respond across a wide range of frequencies (0.3 to 5kHz), at low to moderate sound levels (45-65dB)<br /> (Aleksandrov and Dmitrieva 1992, Korneeva et al 2006).

      * Stating that these studies in flycatchers "likely yield lower thresholds" (L364) is an astonishing understatement. Thresholds in Aleksandrov and Dmitrieva (1992) at 2-3 day old were 35 dB lower than those here at 4 day old with clicks, and 60 to 80dB lower than those with tones of 1-2kHz at 8 day old.

      * Claims that "sensitivity improves rapidly postnatally, with ~40 dB threshold decreases (L365)" is also misreporting these studies' findings. Thresholds decreased by 25dB consistently at 9 out of 11 frequencies tested from 0.3 to 8kHz (Aleksandrov and Dmitrieva, 1992). A difference of 40dB was only found at 5-6kHz (Aleksandrov and Dmitrieva, 1992). Likewise, improvement also varied from 25 to 40dB in Korneeva 2006. These do not average to "~40 dB".

      * Even birds developing 4 times slower than songbirds (budgerigars) show a response at 5 day old. My statement that Brittan-Powell 2004 shows an auditory response at 5d old is correct. Re-response R3.A2.2: Fig 1B shows one example for ONE individual. All other figures based on multiple individuals shows at least some individuals responded at 5-6 days old at frequencies less than 4Khz (Fig 1C, 4 and 5).

      - Many inaccuracies on avian hearing remain uncorrected in the revision.<br /> * e.g. L 329: extrapolating high frequency hearing (>6khZ) from precocial species is incorrect because even adult chicken and ducks are not sensitive to high frequencies.<br /> The authors imply elsewhere that species of songbirds cannot be compared (resp A4, R3.A6), but make extrapolations from species that are far more remote, phylogenetically, developmentally and ecologically than other songbirds.

      *** Behavioural evidence

      - contradicts this study findings:

      * When correctly cited and described, the literature, does not support the authors' statement that nestling behavioural response "typically emerges between ~5-10 days post-hatch" (L358). It emerges earlier, at an unknown age, including potentially from hatch (present at 1.5day) for innate responses (see below).

      * Results in other songbirds are not consistent with this study finding that a "response to loud click stimuli is first detectable at 4-8 days post hatch" (L230). Instead they show nestling hearing capacities described here are abnormally poor.

      * Therefore, the conclusion that "The timeline of behavioural studies closely matches the onset and maturation of ABR responses observed here in zebra finches" (L360) is wrong.

      - shows a very early response (1-2day post hatch), with no known onset:

      * songbird behavioural response to sound, with parental alarm call suppressing begging, has been demonstrated in nestlings as young as 1.5 or 3 days old (Khayutin 1985, Korneeva et al 2006, Aleksandrov and Dmitrieva 1992). None of these results are mentioned in the revision when discussing behavioural evidence (L357-360).

      * Instead, they exclusively mention studies that started testing nestlings at 5-6 days old, or much later (e.g. 17 day old: Suzuki 2011; 14 day old: Barati & McDonald 2017).

      * ALL of these studies (cited or not) demonstrated a significant response of nestlings to calls on the first age tested. NONE tested the onset of this response.

      * the one study looking at progression across 3 ages, at 5, 8 and 11 day old shows parental alarm calls suppress nestling calling at day 5 as much as later on, with no effect of age (Platzen and Magrath 2004). The authors failed to acknowledge this in their response (resp A4) or revision (L359).

      * the claim that Haff & Magrath (2012) showed nestlings did not respond at 5-6 day old but did at 10-11 days (resp A4) is wrong. At 5 day old, they responded to their own species alarm calls, as well as to another similar sounding species and to the sound of predators themselves (Table2 in Haff & Magrath 2012; as in Platzen and Magrath 2004).

      - learning, not just hearing, improves nestling response with age:

      * Haff & Magrath (2012) showed that by 10-11 day old, nestlings had learnt to also respond to heterospecifics, demonstrating that learning improves nestling response with age.

      * the intensity of the response to low frequency sound improved more with age than that to high frequency calls. If improvement were related to hearing limitations, the opposite would be expected (Haff & Magrath (2012).

      - nestlings discriminate complex calls, including at high frequency

      * By 5-6 day old, nestlings can already discriminate several different sounds indicative of danger, among the complex natural acoustic background (Haff & Magrath 2012).

      * nestling do not respond indiscriminately to any calls, but only to relevant sounds that specifically present a threat to them (Haff and Magrath 2012; Magrath et al 2006).

      * the idea that nestlings only distinguish "low frequency broadband cues" (L362) is inaccurate (resp R3.A6). Nestlings respond to scrubwren chip calls and fairywren alarm calls that are narrowband calls with the fundamental at 8 and 10 kHz respectively (Platzen and Magrath 2004; Haff and Magrath 2012).

      * These studies indeed "do not imply mature auditory sensitivity" (L362), they show that auditory maturity is not needed to show a perfectly functional response to biologically meaningful sounds in a natural context.

      - Overall, every other songbird species tested shows greater hearing sensitivity than that proclaimed here for zebra finches. It is very unlikely that zebra finches would be such an outlier.

      (ii) INCORRECT CONCLUSION ON DEAFNESS

      Deafness of young zebra finch nestlings cannot be demonstrated because:

      - the study has no valid response to tones in nestlings to establish the age of hearing onset.

      - responses to clicks are greatly underestimated (see 1). Had correct, more sensitive, protocol parameters been used, it is very likely that:<br /> * Most 2 day old nestlings would have responded to clicks, instead of 60% of 4 day old nestlings.<br /> * Thresholds would be lower, especially in nestlings.<br /> The >54dB difference between nestlings and adults based on an assumed 95dB threshold in 2d old nestlings is wrong.

      - based on data on other songbird nestlings, a difference of 25dB would be more realistic (at frequencies of 1-2 kHz, comparable to clicks, Aleksandrov and Dmitrieva 1992, Korneeva et al 2006). This greatly contrasts with the >54dB estimate here. While the authors qualify their estimate as "conservative" (legend Fig4, L293), it is actually greatly overestimated.

      - Even assuming the data in this study were correct, the interpretation is erroneous. Concluding "deafness of young nestlings" is incorrect when 60% of 4 day old nestlings respond to clicks at 80dB. This is especially wrong given that the ABR method systematically overestimates thresholds by 20-40dB.

      - That ABR is used in humans to diagnose deafness (L263) does not make ABR the most suitable method for birds, when evidence shows that other methods (e.g. electrophysiology or behaviour) are more accurate. Hearing screening in humans can only use non-invasive methods, and has additional criteria than accuracy (price, ease, etc).

      - The discussion fails to acknowledge the implications of the limitations of the study (detailed above). For example, talking in broad terms of differences in protocols (L392-395), does not tell readers that this study, because of the parameters chosen, led to an underestimation of zebra finch hearing capacities.<br /> The discussion instead greatly overstates what the study shows (e.g. L229-232, 261, 277-278, 288, 332, 340, etc).

      (iii) INCORRECT CONCLUSION THAT HEAT-CALL FALL OUTSIDE THE ADULT HEARING RANGE.

      - heat-calls fall outside the adult hearing range solely because of the authors' decision to:<br /> * restrict heat-call frequency range to the authors' own estimation (7-10kHz) on 4 birds instead of using published values that caused developmental impact (e.g. Katsis et al: 6-10kHz), or were produced in vitro (5.9kHz; Anttonnen et al 2025);

      * lowering heat call sound level (34dB at 10 cm) to measurements on 4 isolated birds under unknown temperature conditions for an unknown amount of time, when the same paper gives values of 43.4 dB at 1 m (range: 30.7-53.2) in standard in vitro conditions.

      * As soon as EITHER of these two values is corrected, the statement that heat calls are outside zebra finch adult hearing range is false.

      * In addition, assuming constant heat-call sound level across contexts is unreasonable. That zebra finch can produce inspiratory syllable very similar to heat calls and during inspiration at 65dB during song (Goller & Dalley 2001) argues against the assumption that heat-calls are always soft.

      - Other species also show perception, not just production (resp R3.A10a), of calls above their known hearing range (10-20kHz; e.g. Duque et al 2020). The zebra finch is not the only case in birds of mismatch between signals perceived and known hearing range.

    5. Reviewer #4 (Public review):

      Reviewing Editor:

      There is a virtuous circle between behavior and neuroscience. Sometimes neuroscience identifies a signal that no behavioral study alone could discern: place cells in rats, replayed song sequences in sleeping birds, compass-like signals in the central complex of the fruit fly. Sometimes behavioral observation is what inspires neuroscience: precise measurements of short and long latency reflexes implicate different neural pathways, and directed escape responses in fish are so fast that they require specialized and lateralized escape circuits. And sometimes a behavioral claim requires an animal's sensory system to possess a capacity nobody had documented. A prime example is bat echolocation: Spallanzani inferred in 1793 that bats navigate by hearing, but the claim was dismissed for over a century because the signal he proposed could not be detected. It was vindicated only when Griffin and Galambos measured the ultrasonic emissions directly in 1940, and the loop closed fully when Suga and colleagues found cortical neurons in the mustached bat tuned to precisely the echo delays and Doppler shifts the behavior required.

      In this manuscript we are faced with a fascinating and contemporary instance of this important dialogue between behavior and neurophysiology. Several high profile papers have reported that incubating zebra finch parents produce high frequency "heat calls" when ambient temperatures rise, and that playback of these calls to eggs during late incubation alters offspring growth, begging behavior, thermal preference, and reproductive success in adulthood. That function requires that the embryo be able to sense, presumably to hear, the heat calls. This is textbook ecology and often cited as the prime example of adult behavior affecting the development of their unborn (or unhatched) offspring.

      Yet the auditory capacity of very young zebra finches had never been carefully examined. This paper does exactly that, using auditory brainstem responses, a method that detects synchronous volleys of afferent input through low levels of the auditory system. ABRs are not perfect and may miss very subtle or sparsely represented signals, but they are a time-tested way to assess hearing capacity and the development of a system.

      Here, the authors show, convincingly in my view and in the views of Reviewers 1 and 2, that 2 DPH hatchlings have no detectable ABR to a 95 dB SPL broadband click, and that sensitivity then rises progressively across the first two postnatal weeks. This is the most consequential result, since it bears directly on whether an auditory route to heat call perception is feasible at all. A second finding is that ABR wave I amplitude continues to mature until 20 to 25 days post hatch, coinciding with the onset of sensory song learning, just as young finches need to form a template of an adult male song to copy.

      The key question relevant to mediating this review process is: If a two-day-old hatchling shows no detectable auditory brainstem response, over 400 averaged sweeps, to a 95 dB SPL broadband click, in 14/14 animals tested, then how could that same animal, two days earlier and inside an egg, have been sensitive to a far weaker stimulus of roughly 33 dB SPL at 6.8 kHz?

      Below I set out the considerations that influenced my judgment as I handled these reviews.

      Strengths:

      The developmental series is the strength of the design. Seven ages, with multiple early time points, means the negative result at 2 DPH is not a lone flat trace but sits at one end of a graded and internally consistent trajectory. Body temperature was monitored and held within {plus minus}0.5{degree sign}C, and differed by no more than 0.6{degree sign}C across age groups. Every stimulus parameter was applied identically at every age, so the developmental comparison is a within-method one. The vibrometry experiment addresses the most obvious alternative modality with an appropriate and state-of-the-art technique.

      The key findings related to the development of the ABR response and, presumably, the hearing capacity. At two days post hatch, the authors presented broadband clicks of 20 µs duration at 95 dB SPL peak equivalent, a stimulus carrying energy across the full spectrum including the 6.8 kHz of the heat whistle. Clicks were delivered at 25 Hz with a 40 ms inter-stimulus interval, 400 presentations per condition with alternating polarity, and responses were scored both by an automated signal-to-noise criterion and by independent visual inspection. No response was observed in any of the 14 animals tested. Two days later, under the same protocol, eight of thirteen animals responded, with a mean threshold of 80.6 {plus minus} 1.6 dB SPL and wave I latencies of roughly 4.0 to 4.7 ms against 1.8 to 2.5 ms in adults. These responses were small, broad, and slow, the signature of an immature auditory system. By 6 DPH nine of ten animals responded, and from 8 DPH onward every animal did at every age tested. Adult thresholds settle at 40.6 {plus minus} 4.0 dB SPL, more than 54 dB below the 2 DPH value. The 4 DPH data show the preparation resolves weak, desynchronized responses under exactly the parameters used at 2 DPH, and that sensitivity emerges on a trajectory consistent with auditory development in other birds and in mammals.

      On Reviewer 3's methodological objections:

      Reviewer 3 raised a series of objections to the recording parameters: the stimulus rate is faster than in comparable avian developmental studies, 400 sweeps is fewer than the conventional 1000, body temperature sits at the low end of the reported range for the species, and clicks lag tones in developmental onset. Each of these describes a mechanism that would reduce the amplitude of an evoked response relative to some ideal stimulus. But in aggregate my take was that none of these would completely abolish a response that is detectable two days later. That distinction is the crux of my reading and others may disagree. A response reduced by 70 percent is still a response, and the question at 2 DPH is not why the trace was small but why there was no trace at all over 400 averages in all 14 animals tested.

      Reviewer 3's concerns are difficult to translate into threshold shifts, and on the narrow point that amplitude reductions do not map cleanly onto decibels, the difficulty that the reviewer faced in converting stimulus concerns to impact on dB threshold is well taken. But the figures supplied are themselves bounded: a 37 percent noise reduction from additional sweeps, a 70 percent amplitude reduction from stimulus rate in the youngest birds, a four day shift in apparent onset from using clicks rather than tones. These are large effects. They are not unbounded ones. And because the decibel is a logarithmic unit, in which every 20 dB corresponds to a tenfold change in sound pressure, the gap they are being asked to explain, exceeding 54 dB, amounts to a difference of more than 500-fold.

      The 4 DPH data bear directly on this, because every parameter at issue was applied identically at that age. The same 25 Hz rate, the same 400 sweeps, the same temperature, the same click stimulus resolved small, dispersed, long latency responses in 8 of 13 animals two days later. These are precisely the weak and desynchronized responses an immature auditory system is expected to produce, and precisely what the objections predict should have been lost. Whatever theoretically possible deficits in stimulus design, they did not prevent detection of a marginal response in animals two days older, and no mechanism has been proposed by which their cost would fall from total suppression at 2 DPH to negligible at 4 DPH.

      A 95 dB SPL broadband click carries energy at every frequency including 6.8 kHz and exceeds the level at which heat whistles arrive at 10 cm by more than 60 dB, before any attenuation by shell or by an incubating parent. Even a substantial underestimate of hatchling sensitivity leaves that gap open, and at 6.8 kHz it is a broadband stimulus being compared against a narrowband signal in the frequency region that matures last, in this dataset and in every developmental series reported.

      Sensory systems as matched filters:

      A second consideration: sensory systems as matched filters (survival-critical sensory responses are typically associated with expanded sensory sensitivity for them). Failure to detect an auditory response is not proof of functional deafness. An evoked potential measures synchronized population activity, so sparse discharge below the threshold for ABR sensitivity (a population response) can never be excluded. This is the crux of Reviewer 3's concerns.

      But this second consideration makes the difficulty of finding any evidence of hearing or vibrotactile responsiveness in these animals more compelling. Rüdiger Wehner, working on desert ant navigation, noted that sensory systems are matched filters. They are tuned to the narrow slice of the world that matters for survival and reproduction and discard the rest. From the tuning of an ant's polarization channel you can infer what it navigates by, and from the work of Capranica we know that the tuning of a frog's inner ear relates to the sound of conspecific frogs. If acoustic or vibrotactile processing of heat calls were genuinely critical to the embryo, natural selection would have built the outsized sensitivity to receive it. Instead there is a greater than 54 dB deficit and a frequency range where nothing is detectable at all. I interpret this as evidence that evolution did not build auditory sensitivity into two-day-old hatchlings, which suggests functional deafness in the embryo.

      One might invert this argument and object that the filter in the embryo is precisely for the heat whistle itself, so that neither a broadband click nor a 25 ms tone burst, the only two stimuli presented to 2 DPH animals, is the right probe. Feature detectors that respond to a specific signal while ignoring stimuli of far greater energy are common: cricket AN2 neurons fire to bat-like pulse intervals and stay quiet to loud broadband noise, and anuran midbrain neurons are interval-tuned and silent to spectrally matched noise exceeding the call in level. Embryonic sensory systems also carry transient specializations that later disappear, and the heat whistle is unusually specifiable, narrowband near 6.8 kHz and delivered in rhythmic trains. A detector tuned to that rhythm would not have been engaged by any stimulus used here, since the actual call was never played to any animal at any age.

      But feature selectivity is typically computed centrally, downstream of peripheral transduction, and ABR wave I reflects auditory nerve output, below any circuit that could implement pattern selectivity. A central detector still requires afferent input to operate on, and it is that input which seems to be absent here. It's conceivable that the cochlea itself exhibits some very special tuning to the heat whistle and implements some kind of yet-to-be discovered surround suppression at the periphery, such that a click would not be able to activate the hair cells in the heat whistle's frequency range. But 2 DPH animals were also tested directly with narrowband 6 and 8 kHz stimuli at 95 dB SPL, bracketing the heat whistle frequency, also without response. The surviving version of the objection therefore requires a pathway with too few fibers to generate a far-field response, tuned to a temporal structure nobody tested, in an animal whose auditory nerve shows no signal at some 60 dB above the natural signal level.

      Caveats:

      The paper comes with some caveats, which the authors note in their discussion. First, no embryo was tested. The embryonic claim is an extrapolation from hatchlings. I regard it as a reasonable one, since sound must additionally traverse the shell and since auditory systems gain rather than lose function as development proceeds, but it is an extrapolation nonetheless. Second, frequency-specific thresholds in nestlings were not obtained with a full stimulus set. The low to high developmental sequence therefore rests on a conservative stimulus. Pip data in nestlings would be a valuable addition to the record. Finally, the study offers no pre-neural measure, such as cochlear microphonics or otoacoustic emissions. Such a measure would distinguish a cochlea that is not transducing from one that transduces without synchronized output, and would address the strongest form of the objection raised in review.

      Summary:

      Absence of evidence is not evidence of absence, and some future study could in principle identify single neuron responses to heat calls in an embryo. The evidence in this paper makes that outcome unlikely in my assessment. It is worth noting what the precedent actually supports. Prenatal hearing is real and well documented in precocial birds: mallard embryos deprived of exposure to their own calls fail to recognize the maternal assembly call after hatching, and chickens show evoked responses well before hatching. But those animals hatch at a developmental stage altricial songbirds do not reach until days later, and the calls involved carry their energy below 3 kHz, the frequency region that comes online first. The claim at issue here requires sensitivity at 6.8 kHz, the region that matures last, in an animal at a far earlier stage. Something may still be happening to the embryo during heat calls, but whatever the mechanism, it is probably not the embryo listening. Behavioral claims that lack strong precedent, such as hearing through an eggshell, need to sit in a virtuous circle with neurophysiological investigation, and this study is what that looks like from the physiological side.

    6. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work by Antonnen et al. was triggered by claims of auditory-mediated effects on altricial avian embryos, which were published without any direct evidence that the relevant parental vocalizations were actually heard. I agree with Anttonen et al. that, based on the available evidence about avian auditory development, those claims are highly speculative and therefore necessitate more direct experimental verification.

      Attonen et al. have embarked on a comprehensive series of experiments to:

      (1) Better characterize acoustically the relevant parental vocalizations (heat whistles; in a separate preprint, not reviewed here)

      (2) Characterize the auditory sensitivity of zebra finches at various stages of their posthatching development. Despite the long-standing importance of the zebra finch as a songbird model in neuroethology of learned vocalizations, the auditory development of the species has not been studied so far.

      (3) Explore an alternative hypothesis of how the parental vocalizations might be perceived.

      The principal method used here is the non-invasive recording of ABR (auditory brainstem response), a standard neurophysiological method in auditory research. The click-evoked ABR provides a quick and objective assessment of basic hearing sensitivity that does not require animal training. Weaknesses of the technique include its limited frequency specificity and low signal-to-noise ratio. The authors are experienced with ABR measurements and well aware of those issues. ABR responses in zebra finches are shown to gradually appear during the first week posthatching and to mature in subsequent weeks, consistent with the auditory development in other altricial bird species studied previously. When matching the acoustic properties of parental heat whistles and auditory sensitivities, hearing of the parental heat whistles by zebra finch hatchlings was convincingly excluded. Although not directly measured, this also convincingly extrapolates to zebra finch embryos. Finally, the authors tested the hypothesis that parental heat whistles could induce perceptible vibrations of the egg and thus stimulate the embryo via a different modality. The method used here was laser doppler vibrometry, an appropriate, state-of-the-art technique that the authors also have proven experience with. The induced vibrations were shown to be several orders of magnitude below known vibrotactile sensitivities in mammals and birds. Thus, although zebra finch vibrotactile thresholds were not obtained directly, the hypothesis of vibrotactile perception of parental heat whistles by zebra finch embryos could also be rejected convincingly.

      In summary, even when considering some weaknesses of the techniques (which the authors are aware of), the conclusions of the paper are well supported: Auditory and/or vibration perception of parental heat whistles can be excluded as an explanation for previous reports of developmental programming for high ambient temperatures. As a constructive suggestion towards resolving the apparent paradox, the authors recommend repeating some of the crucial, previous playback experiments at lower sound levels that better match the natural parental vocalizations.

      (R1. A1) We thank the reviewer for their time and effort to thoroughly review our paper and for the positive comments on our manuscript. In the revised manuscript we have addressed the concerns that you have raised in the Joint recommendations above (Pages 1-4).

      Reviewer #2 (Public review):

      This study by Anttonen, Christensen-Dalsgaard, and Elemans describes the development of hearing thresholds in an altricial songbird species, the zebra finch. The results are very clear and along what might have been expected for altricial birds: at hatch (2 days post-hatch), the chicks are functionally deaf. Auditory evoked activity in the form of auditory brainstem responses (ABR) can start to be detected at 4 days post-hatch, but only at very loud sound levels. The study also shows that ABR response matures rapidly and reaches adult-like properties around 25 days post-hatch. The functional development of the auditory system is also frequency dependent, with a low-to-high frequency time course. All experiments are very well performed. The careful study throughout development and with the use of multiple time-points early in development is important to further ensure that the negative results found right after hatching are not the result of the experimental manipulation. The results themselves could be classified as somewhat descriptive, but, as the authors point out, they are particularly relevant and timely. Since 2016, there have been a series of studies published in high-profile journals that have presumably shown the importance of prenatal acoustic communication in altricial birds, mostly in zebra finches. This early acoustic communication would serve various adaptive functions. Although acoustic communication between embryos in the egg and parents has been shown in precocial birds (and crocodiles), finding an important function for prenatal communication in altricial birds came as a surprise. Unfortunately, none of those studies performed a careful assessment of the chicks' hearing abilities. This is done here, and the results are clear: zebra finches at 2 and 6 days post-hatch are functionally deaf. Since it is highly improbable that the hearing in the egg is more developed than at birth, one can only conclude that zebra finches in the egg (or at birth) cannot hear the heat whistles. The paper also ruled out the detection on egg vibrations as an alternative path. The prior literature will have to be corrected, or further studies conducted to solve the discrepancies. For this purpose, the "companion" paper on bioRxiv that studies the bioacoustical properties of heat calls from the same group will be particularly useful. Researchers from different groups will be able to precisely compare their stimuli.

      Beyond the quality of the experiments, I also found that the paper was very well written. The introduction was particularly clear and complete (yet concise).

      Weaknesses:

      My only minor criticism is that the authors do not discuss potential differences between behavioral audiograms and ABRs. Optimally, one would need to repeat the work of Okanoya and Dooling with your setup and using the same calibration. The ~20dB difference might be real, or it might be due to SPL measured with different instruments, at different distances, etc. Either way, you could add a sentence in the discussion that states that even with the 20 dB difference in audiogram heat whistles would not be detected during the early days post-hatch. But adding a (novel) behavioral assay in young birds could further resolve the issue.

      (R2. A0) We thank the reviewer for their time and effort to thoroughly review our paper, and for the positive comments on our manuscript.

      In our revision, we have added a new figure (Fig 5) and three new paragraphs (Lines 387-422) in the discussion to compare all published ABR and behavioral audiograms and the differences between and among these datasets. Our adult data is consistent with the reported findings from four other labs despite differences in stimulus design, setups and genetic background of the animals, providing strong support for our findings.

      Furthermore we have more clearly presented our argument why we think juvenile and embryos cannot detect heat whistles. For clarification, we have added a new figure (Fig 4) and four new paragraphs (Lines 271-355) in the discussion.

      We agree with the reviewer that more data is needed on the development of hearing in songbirds and zebra finches especially; both anatomical data and functional, such as innervation. We emphasize this need in our discussion (Lines 352-355 and Lines 405-411).

      More Minor Points:

      (1) As mentioned in the main text, the duration of pips (from pips to bursts) affects the effective bandwidth of the stimulus. I believe that the authors could give an estimate of this effective bandwidth, given what is known from bird auditory filters. I think that this estimate could be useful to compare to the effective bandwidth of the heat-call, which can now also be estimated.

      (R2. A1) Please see answer A3.4 under Joint Recommendations.

      (2) Figure 5b. Label the green and pink areas as song and heat-call spectrum. Also note that in the legend the authors say: "Green and red areas display the frequency windows related to the best hearing sensitivity of zebra finches and to heat calls, respectively". I don't think this is what they meant. I agree that 1-4 kHz is the best frequency sensitivity of zebra finches, but they probably meant green == "song frequency spectrum" and pink == "heat call spectrum". In either case, the figure and the legend need clarification.

      (R2. A2) We thank the reviewer for pointing out these issues. We have changed the figure and legend accordingly. In the meantime, we published a paper measuring the in vivo source levels of the heat whistles (Anttonen et al., Current Biology 2025), and we have carefully gone through this manuscript to incorporate those findings and adjust the text accordingly. We have therefore adjusted the analysis in Fig 3B to correct for the narrower frequency distribution of the heat whistles.

      (3) Figure 5c. Here also, I would change the song and heat-call labels to "song spectrum", "heat call spectrum". The authors would not want readers to think that they used song and heat calls in these experiments (maybe next time?). For the same reason, maybe in 5a you could add a cartoon of the oscillogram of a frequency sweep next to your speaker.

      (R2. A3) We thank the reviewer for pointing out these issues We mention the frequency sweep in the legend for panel A, but decided against including this in the figure to prevent too much clutter. We have changed the figure = legend to (new text underlined):

      Legend Fig 3. “A Setup used to measure sound-induced vibrations of eggs. A 94 dB, 0.25-to-10 kHz frequency sweep was played at the eggs to determine the vibration transfer function.”

      (4) Methods. In the description of the stimulus, the authors describe "5ms long tone bursts", but these are the tone pips in the main part of the manuscript. Use the same terms.

      (R2. A4) Thank you for catching this, we have changed this into “5ms long tone pips".

      Reviewer #3 (Public review):

      Summary

      Following recent findings that exposure to natural sounds and anthropogenic noise before hatching affects development and fitness in an altricial songbird, this study attempts to estimate the hearing capacities of zebra finch nestlings and the perception of high frequencies in that species. It also tries to estimate whether airborne sound can make zebra finch eggs vibrate, although this is not relevant to the question.

      Strength

      That prenatal sounds can affect the development of altricial birds clearly challenges the long-held assumption that altricial avian embryos cannot hear. However, there is currently no data to support that expectation. Investigating the development of hearing in songbirds is therefore important, even though technically challenging. More broadly, there is accumulating evidence that some bird species use sounds beyond their known hearing range (especially towards high frequencies), which also calls for a reassessment of avian auditory perception.

      Weaknesses

      Rather than following validated protocols, the study presents many experimental flaws and two major methodological mistakes (see below), which invalidate all results on responses to frequencyspecific tones in nestlings and those on vibration transmission to eggs, as well as largely underestimating hearing sensitivity. Accordingly, the study fails to detect a response in the majority of individuals tested with tones, including adults, and the results are overall inconsistent with previous studies in songbirds. The text throughout the preprint is also highly inaccurate, often presenting only part of the evidence or misrepresenting previous findings (both qualitatively and quantitatively; some examples are given below), which alters the conclusions.

      Conclusion and impact

      The conclusion from this study is not supported by the evidence. Even if the experiment had been performed correctly, there are well-recognised limitations and challenges of the method that likely explain the lack of response. The preprint fails to acknowledge that the method is well-known for largely underestimating hearing threshold (by 20-40dB in animals) and that it may not be suitable for a 1-gram hatchling. Unlike what is claimed throughout, including in the title, the failure to detect hearing sensitivity in this study does not invalidate all previous findings documenting the impacts of prenatal sound and noise on songbird development. The limitations of the approach and of this study are a much more parsimonious explanation. The incorrect results and interpretations, and the flawed representation of current knowledge, mean that this preprint regrettably creates more confusion than it advances the field.

      (R3. A0) We thank the reviewer for their detailed and critical assessment. We agree that establishing auditory sensitivity in very young altricial birds is technically challenging and that careful interpretation of ABR data is essential. We also appreciate the reviewer’s recognition of the importance of obtaining direct physiological data on auditory development.

      However, we respectfully disagree with the reviewer’s central claim that our methodology is flawed or that our conclusions are unsupported. Many of the concerns raised reflect misunderstandings of ABR methodology, selective interpretation of the literature, or assumptions that are not supported by empirical evidence. Below, we address the main points in turn.

      As also stated in our Provisional response, the reviewer’s critique can be distilled into four main arguments:

      (1) ABR cannot be reliably measured in very small animals.

      (2) Our stimulus design (especially 25 ms tone bursts) invalidates frequency-specific results.

      (3) ABR thresholds should be corrected to behavioral thresholds, which would alter conclusions.

      (4) Our findings are inconsistent with prior studies in songbirds.

      We address each of these below before responding point-by-point.

      (1) Suitability of ABR in small animals.

      Reviewer claim: ABR may not be suitable for very small hatchlings.

      This claim is not supported by existing evidence. ABR measures summed neural activity, and signal amplitude depends in part on the distance between neural tissue and recording electrodes. In smaller animals, this distance is reduced, which can increase signal amplitude and improve signal-to-noise ratio.

      Consistent with this, ABR has been successfully recorded in animals substantially smaller than zebra finch hatchlings, including zebrafish (Jørgensen et al., 2012), 10 mm froglets (Goutte et al., 2017) and 5 mm salamanders (Capshaw et al., 2020). It is in fact much more surprising the technique still provides robust signals even in extremely large animals such as Minke whales, where the distance between electrodes and brain is on the decimeter scale (Houser et al., 2024). We have extensive experience of recording ABRs in such small systems.

      Thus, there is no principled reason why ABR would be an invalid method to study auditory sensitivity in zebra finch hatchlings.

      (2) Stimulus design and tone duration

      Reviewer claim: Use of 25 ms tone bursts invalidates frequency-specific results.

      We agree that stimulus duration affects frequency specificity and ABR detectability. However, the reviewer’s assertion that there is a single “correct protocol” (≤5 ms) is inaccurate. In avian ABR studies, stimulus duration varies depending on experimental goals.

      Our choice of 25 ms tone bursts was intentional and necessary to accurately represent low frequencies (down to 250 Hz), ensuring sufficient cycles per stimulus in the plateau segment of 15 ms and minimizing spectral splatter (see auditory brainstem response design considerations discussed in Lauridsen et al., 2021) and our responses below.

      Key clarifications:

      a) Click-evoked ABRs form the basis of our conclusions about onset of hearing, not tone bursts.

      b) Tone bursts were used primarily to assess frequency-dependent maturation, not detect earliest sensitivity.

      c) We explicitly demonstrate that:

      - 25 ms bursts yield higher thresholds (lower sensitivity)

      - 5 ms pips yield lower thresholds and align with published ABR audiograms

      We have now:

      - Further clarified the rationale of stimulus design in the methods (Line 520-530) and added a section in the discussion (Lines 397-411).

      - Included additional comparison between burst and pip datasets (Lines 387-395).

      - Clarified that conclusions about early hearing do not depend on tone-burst data (Fig 4 and Lines 271-294).

      (3) ABR vs behavioral thresholds

      Reviewer claim: Failure to correct ABR thresholds (20–40 dB) invalidates conclusions.

      We agree that ABR thresholds typically overestimate behavioral thresholds. However, we disagree that this invalidates our conclusions.

      Importantly:

      a) We do not replace measured ABR data with corrected values, as this would be methodologically inappropriate.

      b) Instead, we:

      - Present measured ABR thresholds transparently (Fig 1-3)

      - Compare them directly to published behavioral audiograms (Fig 5)

      - Explicitly discuss the expected offset (Lines 413-422)

      In the revised manuscript we:

      a) Add a new figure (Fig. 5) compiling all published ABR and behavioral audiograms

      b) Show that:

      - ABR and behavioral audiograms have similar shapes (Fig 5, new discussion Lines 387-422)

      - Offsets are typically ~20 dB (Line 413-422)

      Crucially, even under conservative corrections:

      - Early hatchlings remain far less sensitive than adults (>54 dB SPL) to clicks.

      - Heat whistle levels remain at or below detection limits even in adults (new figure Fig 4)

      - The developmental gap (>50 dB between adults and 2 DPH hatchlings) remains decisive.

      Thus, incorporating ABR–behavioral differences does not change the central conclusion.

      (4) Consistency with prior literature in developing songbirds.

      Reviewer claim: Results contradict previous studies in developing songbirds.

      We respectfully disagree. The cited studies fall into three categories:

      (1) Behavioral studies (e.g., alarm-call responses)

      (2) Gene expression studies (e.g., ZENK activation)

      (3) Different species with different developmental trajectories

      None of these directly measure auditory sensitivity thresholds in zebra finch embryos or hatchlings.

      We emphasize:

      - Behavioral responses do not provide threshold measurements

      - Neural activation (e.g., ZENK) does not demonstrate functional perception thresholds

      - Cross-species comparisons must consider differences in developmental timing.

      We have now expanded the Discussion to explicitly address these studies and clarify how they relate to our findings (Lines 357-367). Our data are consistent with what is known about the physiology of auditory development in all birds studied so far.

      (5) Final statement

      We have revised the manuscript extensively to:

      - Clarify methodology and experimental design

      - Expand discussion of ABR limitations

      - Incorporate additional literature and comparisons

      - Correct inconsistencies in reporting

      We maintain that our central conclusions—that early zebra finch hatchlings lack detectable auditory brainstem responses and are unlikely to perceive parental heat calls at natural levels—is robust and supported by the data. We go into more detail in our point-by-point rebuttal below.

      Detailed assessment

      For brevity, only some references are included below as examples, using, when possible, those cited in the preprint (DOI is provided otherwise). A full review of all the studies supporting the points below is beyond the scope of this assessment.

      (A) Hearing experiment

      The study uses the Auditory Brainstem Response (ABR), which measures minute electrical signals transmitted to the surface of the skull from the auditory nerve and nuclei in the brainstem. ABR is widely used, especially in humans, because it is non-invasive. However, ABR is also a lot less sensitive than other methods, and requires very specific experimental precautions to reliably detect a response, especially in extremely small animals and with high-frequency sounds, as here.

      (1) Results on nestling frequency sensitivity are invalid, for failing to follow correct protocols:

      (R3. A1.1) We disagree that our protocol is invalid. There is no universal ABR protocol standard in birds. Our approach is consistent with established principles of stimulus design and is validated by:

      - Robust click-evoked responses

      - Consistent developmental trajectories

      - Agreement between pip-based ABR and behavioral audiograms

      We now clarify this explicitly. See response R3. A0 (Stimulus design and tone duration).

      The results on frequency testing in nestlings are invalid, since what might serve as a positive control did not work: in adults, no response was detected in a majority of individuals, at the core of their hearing range, with loud 95dB sounds (Figure S1), when testing frequency sensitivity with "tone burst".

      This is mostly because the study used a stimulation duration 5 times larger than the norm. It used 25ms tone bursts, when all published avian studies (in altricial or precocial birds) used stimulation of 5ms or less (when using subdermal electrodes as here; e.g., cited: Brittan-Powell et al 2004; not cited: Brittan-Powell et al 2002 (doi: 10.1121/1.1494807), Henry & Lucas 2008 (doi: 10.1016/j.anbehav.2008.08.003)). Long stimulations do not make sense and are indeed known to interfere with the detection of an ABR response, especially at high frequencies, as, for example, explicitly tested and stated in Lauridsen et al 2021 (cited).

      (R3. A1.2) ABR with long-duration stimuli were shown previously to work perfectly well in birds and do not interfere with the detection of an ABR response. Longer stimuli have been used, e.g. in the following bird papers:

      - Amin et al., J Neurophysiol 2007 (cited) on zebra finch ABR: 20 ms tone bursts

      - Korneeva et al. (2006), evoked responses from field L in flycatcher (cited by the reviewer): 20 ms

      - Saunders et al. (1973) (cited): 60 ms tone bursts.

      - Larsen ON, Wahlberg M, Christensen-Dalsgaard (2020) Amphibious hearing in a diving bird, the great cormorant (Phalacrocorax carbo sinensis), J Exp Biol, doi:10.1242/jeb.217265: 25 ms tone bursts

      Human ABR has been measured even with long-duration speech signals (duration 40 ms and longer). See for example Binkhamis et al., Ear and Hearing 40: 659-670, 2019.

      Furthermore, the reviewer unfortunately misunderstood some aspects in Lauridsen et al., which tested a specifical method exploiting neural phase locking, and showed that this method has a low-frequency bias because neural phase locking decreases at high frequencies.

      Also, Lauridsen et al clearly show the reason for using 25 ms bursts: because we aimed to measure frequencies down to 250 Hz, we need to have a sufficient number of cycles (3) in the plateau segment to represent the frequency adequately (Lauridsen et al. fig 1). A 15 ms plateau contains 3 cycles, plus 5 ms rise/fall time (1 cycle) equals 25 ms. We have clarified this in our methods (L520-525) and added a section in the discussion (L397-411).

      Thus, long-duration stimulations make sense and have been used before successfully. See also response R3.A0 (Stimulus design and tone duration).

      Adult response was then re-tested with a correct 5ms tone duration ("tone-pip"), which showed that, for the few individuals that responded to 25ms tones, thresholds were abnormally high (c.a. by 30dB; Figure 2C).

      Yet, no nestlings were retested with a correct protocol. There is therefore no valid data to support any conclusion on nestling frequency hearing. Under these circumstances, the fact that some nestlings showed a response to 25ms tones from day 8 would argue against them having very low sensitivity to sound.

      (R3. A1.3) Please see answer A3.1 under Joint Recommendations and R3.A0 (Stimulus design and tone duration).

      (2) Responses to clicks underestimate hearing onset by several days:

      Without any valid nestling responses to tones (see # 1), establishing the onset of hearing is not possible based on responses to clicks only, since responses to clicks occur at least 4 days after responses to tones during development (Saunders et al, 1973). Here, 60% of 4-day-old individuals responding to clicks means most would have responded to tones at and before 2 days post-hatch, had the experiment been done correctly.

      (R3. A2.1) We disagree that clicks necessarily underestimate onset.

      Clicks are broadband stimuli that:

      - Recruit large neural populations

      - Are commonly used to detect early auditory responses

      The cited delay between tone and click responses reflects stimulus energy differences, not an inherent limitation of clicks. We have clarified this in the revision and softened language to refer to “no detectable ABR response” rather than absolute deafness.

      The report that Saunders could only see responses to clicks later than to tones only reflects that he used click amplitudes that were insufficiently high. The ABR responses reported were to extremely intense tones (110 dB SPL) of long duration (60 ms).

      If Saunders had used clicks (duration 60 µs) with comparable sound energy, they would have had been very difficult to produce. He should have used clicks with an amplitude 1000 times (60 dB) higher than the tones to produce the same sound energy. This would have been clicks at 170 dB SPL, equivalent to the sound at the mouth of a medium-sized military cannon. Applying this pressure would not be a recommendable method for hearing assessment, but instead lead to irreversible hearing damage.

      In budgerigars, hearing onset occurs before 5 days post hatch, since responses to both clicks and tones were detectable at the first age tested at 5dph (Brittan-Powell et al, 2004).

      (R3. A2.2) This is not how we interpret the cited paper. They state that ‘Responses were first obtained from 1-week-old at high stimulation, and their click responses (Fig. 1) show no wave 1 peak at 6 days post-hatch. Also, their conclusion (p 3101) states that ‘budgerigars probably cannot hear at hatching’.

      (3) Experimental parameters chosen lower ABR detectability, specifically in younger birds: Very fast stimulus repetition rate inhibits the ABR response, especially in young:

      (a) The stimulus presentation rate (25 stim/ sec) is 6 times faster than zebra finch heat-calls, and 5 to 25 times faster than most previous studies in young birds (e.g., cited: Saunders et al 1973, 1974: 1 stim/sec or less; Katayama 1985: 3.3 clicks/sec; Brittan-Powell et al 2004: 4 stim/sec).

      Faster rates saturate the neurons and accordingly are known to decrease ABR amplitude and increase ABR latency, especially in younger animals with an immature nervous system.

      In birds, this occurs especially in the range from 5 to 30 stim/sec (e.g., cited: Saunder et al 1973, Brittan-Powell et al 2004). Values here with 25 rather than 1-4 stim/min are therefore underestimating true sensitivity.

      (R3. A3a) Please see answer A3.3 under Joint Recommendations.

      (b) Averaging over only 400 measures is insufficient to reliably detect weak ABR signals: The study uses 2 to 3 times fewer measures per stimulation type than the recommended value of 1,000 (e.g., Brittan-Powell et al 2002, 2024; Henry & Lucas 2008). This specifically affects the detection of weak signals, as in small hatchlings with tiny brains (adult zebra finches are 12-14g).

      (R3. A3b) Please see answer A3.2 under Joint Recommendations.

      (c) Body temperature is not specified and strongly affects the ABR:

      Controlling the body temperature of hatchlings of 1-4 grams (with a temperature probe under a 5mm-wide wing) would be very challenging. Low body temperature entirely eliminates the ABR, and even slight deviance from optimal temperature strongly increases wave latency and decreases wave amplitude (e.g., cited: Katayama 1985).

      (R3. A3c) Please see answer A2 under Joint Recommendations.

      (d) Other essential information is missing on parameters known to affect the ABR: This includes i) the weight of the animals,

      (R3. A3d-i) These important details have now been added in Table S6.

      (ii) whether and how the response signal was amplified and filtered,

      (R3. A3d-ii) Signal was amplified 500 times (74 dB). We have included these important details to the methods (Line 508).

      (iii) how the automatised S/N>2 criteria compared to visual assessment for wave detection,

      (R3. A3d-iii) There is no universally accepted/fitting protocol for performing ABR recordings in various animals, we decided to perform both visual and automated criteria detection of thresholds. The automated criterion in our experience is more strict approach than visual detection of thresholds, because using an automated criteria for threshold detection will remove potential experimenter bias from the results. We added this in our methods (Lines 583-585).

      (iv) what measures were taken to allow the correct placement of electrodes on hatchlings less than 5 grams.

      (R3. A3d-iv) We have placed electrodes in much smaller animals than 5 grams, and the common landmarks (ear opening, midline of skull) could easily be identified in the hatchlings.

      (4) Results in adults largely underestimate sensitivity at high frequencies, and are not the correct reference point:

      (a) Thresholds measured here at high frequencies for adults (using the correct stimulus duration, only done on adults) are 10-30dB higher than in all 3 other published ABR studies in adult zebra finches (cited: Zevin et al 2004; Amin et al 2007; not cited: Noirot et al 2011 (10.1121/1.3578452)), for both 4 and 6 kHz tone pips.

      (b) The underlying assumption used throughout the preprint that hearing must be adult-like to be functional in nestlings does not make sense. Slower and smaller neural responses are characteristic of immature systems, but it does not mean signals are not being perceived.

      (R3. A4) We acknowledge variation across studies and now include a comprehensive comparison of all zebra finch audiograms (new Fig. 5 and discussion Lines 387-422).

      Importantly:

      - Our pip-based audiogram aligns with previous ABR studies

      - Differences in high-frequency sensitivity likely reflect methodological variation or population differences.

      Our conclusions rely on relative developmental changes, not absolute thresholds.

      (5) Failure to account for ABR underestimation leads to false conclusions:

      (a) Whether the ABR method is suitable to assess hearing in very small hatchlings is unknown. No previous avian study has used ABR before 5 days post-hatch, and all have used larger bird species than the zebra finch.

      (R3. A5a) As stated above (R3.A0), small animals should give better signals, and we have been able to measure ABR in much smaller animals previously.

      (b) Even when performed correctly on large enough animals, the ABR systematically underestimates actual auditory sensitivity by 20-40 dB, especially at high frequencies, compared to behavioural responses (e.g., none cited: Brittan-Powell et al 2002, Henry & Lucas 2008, Noirot et al 2011). Against common practice, the preprint fails to account for this, leading to wrong interpretations.

      (R3. A5a) See our answer to R3.A0 above.

      For example, in Figure 1G (comparing to heat call levels), actual hearing thresholds would be 3040dB below those displayed. In addition, the "heat whistle" level displayed here (from the same authors) is 15dB lower than their second measure that they do not mention, and than measures obtained by others (unpublished data). When these two corrections are made - or even just the first one - the conclusion that heat-call sound levels are below the zebra finch hearing threshold does not hold.

      (R3. A5a) Our conclusion that heat whistles are unlikely to be perceived does not rely on a single dataset or method, but on the convergence of three independent constraints: (i) signal amplitude, (ii) adult auditory sensitivity, and (iii) developmental immaturity of the auditory system.

      First, heat whistles are low-amplitude signals. Our in vivo measurements show levels of ~33 dB re 20 µPa at 10 cm and ~14 dB at 1 m (Anttonen et al., 2025, Curr Biol). Even allowing for uncertainty in near-field estimation, a conservative upper bound at very close range (<5 cm) is ~40 dB SPL.

      Second, the most sensitive available measure—behavioral audiograms—places adult zebra finch thresholds at ~40 dB SPL at ~6 kHz (Okanoya and Dooling, 1987, J Comp Psychol), increasing steeply toward higher frequencies. Thus, even under optimal conditions, heat whistles fall at or below the detection threshold of adults, and only potentially at very close range.

      Third, auditory sensitivity in early development is substantially reduced. Our ABR data show a ≥40–60 dB decrease in sensitivity in hatchlings relative to adults for click stimuli, which provide the most favorable conditions for eliciting responses. Because frequency-specific sensitivity develops later, thresholds at 6–8 kHz are expected to be even higher in hatchlings and embryos.

      Taken together, these constraints define a narrow and unfavorable detection window: a low-amplitude, high-frequency signal positioned at the edge of adult sensitivity, combined with a large developmental decrease in auditory sensitivity. Under these conditions, it is unlikely that heat whistles are detectable by hatchlings or embryos.

      Importantly, this conclusion does not depend on precise correction factors between ABR and behavioral thresholds. Even when considering the most sensitive behavioral data and conservative estimates of sound level, the signal remains at or below the limits of detection in adults, and far below expected sensitivity in early developmental stages.

      We have included this argument more clearly in our discussion (Line 271-367), illustrated by new Fig. 5.

      (c) Rather than making appropriate corrections, the preprint uses a reference in humans (L180), where ABR is measured using a much more powerful method (multi-array EEG) than in animals, and from a larger brain. The shift of "10-20dB" obtained in humans is not applicable to animals.

      (R3. A5c) Again our conclusions rely on relative developmental changes, not absolute thresholds. The clinical practice in humans to measure ABR is with 4 electrodes, not a multi-electrode EEG array.

      Animal studies where ABR audiograms have been compared directly to psychophysical audiograms show differences of around 20 dB. For example in Brittain powell et al 2002, audiogram comparisons between behavioral and ABR in budgerigars were made within the same lab, same animal population and by the same people, leading to 20 dB difference. Our discussion includes a new paragraph on this topic (Lines 413-422).

      (6) Results are inconsistent with previous findings in developing songbirds:

      (R3. A6) We now explicitly discuss all cited studies. Key points:

      - Early behavioral responses do not imply high-frequency sensitivity

      - Studies in other species do not directly translate to zebra finches

      - None of the cited work provides direct measures of auditory thresholds in embryos

      As expected from all of the above, results and conclusions in the preprint are inconsistent with findings in other songbirds, which, using other methods, show for example, auditory sensitivity in: a) zebra finch embryos, in response to song vs silence (not cited: Rivera et al 2018, doi: 10.1097/WNR.0000000000001187)

      (R3. A6a) We thank the reviewer for pointing out this study. We agree that the question of auditory responsiveness in embryos is important, and that a range of approaches have been used to address it. However, the study cited (Rivera et al., 2018) does not directly measure auditory sensitivity, but instead infers auditory processing from differences between treatment groups exposed to different acoustic conditions. As such, it is not directly comparable to physiological measures of hearing sensitivity, such as ABR or behavioral thresholds.

      In addition, interpretation of these results is complicated by limited characterization of the acoustic environment and differences in experimental handling between groups, which may introduce confounding factors unrelated to auditory perception. Given these considerations, and because our study focuses specifically on quantifying auditory sensitivity using established physiological methods, we have chosen not to include a detailed discussion of this work.

      (b) flycatcher hatchlings at 2-3d post hatch (first age tested), across a wide range of frequencies (0.3 to 5kHz), at low to moderate sound levels (45-65dB) (cited: Aleksandrov and Dmitrieva 1992, not cited: Korneeva et al 2006 (10.1134/S0022093006060056)).

      (R3. A6b) Korneeva et al. (2006) and Aleksandrov and Dmitrieva (1992) report evoked responses in very young flycatcher hatchlings across a broad frequency range. Notably, these measurements were obtained using more invasive recording approaches (e.g., implanted electrodes in Field L) in unanesthetized birds, which are known to yield lower thresholds compared to far-field ABR recordings under anesthesia. These methodological differences likely account for part of the higher sensitivity reported.

      Importantly, even in flycatchers, auditory sensitivity shows substantial postnatal improvement: thresholds decrease by up to ~40 dB over the first days after hatching, and the upper frequency limit expands from ~4 to ~7 kHz. Thus, while absolute sensitivity may differ across species and methods, the overall developmental trajectory—gradual improvement in sensitivity and progressive extension toward higher frequencies—is consistent with our findings and with broader patterns reported in songbirds.

      We have added this paper in our discussion (Lines 362-365).

      (c) songbird nestlings at 2-6d post hatch, which discriminate and behaviourally respond to relevant parental calls or even complex songs. This level of discrimination requires good hearing across frequencies (e.g., not cited: Korneeva et al 2006; Schroeder & Podos 2023 (doi: 10.1016/j.anbehav.2023.06.015)).

      (R3. A6c) The species mentioned are different species from our study species. In the Pied flycatchers (Korneeva et al. 2006) experimental conditions were different: recordings were made from unanesthetized nestlings with implanted electrodes directly in the brain (field L), so likely with better SNR. The audiograms show a 10 dB SPL threshold after day 11, so the species may be considerably more sensitive than the zebra finch. The swamp sparrows in the Schroeder and Podos (2023) behavioral study were exposed for 4 days starting at 4-7 days post-hatch, so the study does not address embryonal hearing.

      (d) zebra finch nestlings at 13d post-hatch, which show adult-like processing of songs in the auditory cortex (CNM) (Schroeder & Remage-Healey 2021, doi: 10.1002/dneu.22802).

      (R3. A6d) This study does not conflict with our data. Even though sensitivity is lower at 10 days than in adults, cortical processing could still be ‘adult-like’.

      (e) zebra finch juveniles, which are able to perceive and learn song syllables at 5-7kHz (fundamental frequency) with very similar acoustic properties to heat calls, and also produced during inspiration (Goller & Daley 2001, doi: 10.1098/rspb.2001.1805).

      (R3. A6e) This result is not in conflict with our data. First, the onset of song learning occurs earliest at 20 DPH as discussed in the paper and our work demonstrates that click-evoked ABR thresholds are adult-like at 20 DPH. In the cited paper, the tutoring experiments were initiated at 35 DPH so the auditory system of the studied juveniles is mature.

      Second, even though Goller & Daley 2001 do not report the source level of the specific syllable or the playback sound pressure levels, the source level of the inspiratory notes is comparable to other syllables, and thus around ~67 dB and ~34 dB louder than heat whistles.

      NONE of these results - which contradict results and claims in the preprint - are mentioned.

      Instead, the preprint focuses on very slow-developing species (parrots and owls), which take 2-4 times longer than songbirds to fledge (cited: Brittan-Powell et al 2004; Köppl & Nickel 2007; Kraemer et al 2017).

      (R3. A6f) We have included papers in our discussion that reflect the known data (to our best knowledge) on the developmental neurophysiology and neuroanatomy of the auditory system and not proxies thereof.

      (7) Results in figures are misreported in the text, and conclusions in the abstract and headers are not supported by the data:

      For example:

      (a) The data on Figure 1E shows that at 4 days old, 8 out of 13 nestlings (60%) responded to clicks, but the text says only 5/13 responded (L89).

      (R3. A7a1) We apologize for this typo. Corrected.

      When 60% (4dph) and 90% (6dph) of individuals responded, the correct term would be that "most animals", rather than "some animals" responded (L89).

      (R3. A7a2) We have rephrased this sentence into “observable in most animals during” as suggested.

      Saying that ABR to loud sound appeared "in the majority only after one week" (L93) is also incorrect, given the data.

      (R3. A7a3) We have rephrased this sentence into: “Thus, sounds at loud, yet physiologically relevant SPLs do not evoke ABRs in the first days after hatching, but do so in all animals at 8 DPH.” (Lines 97-99).

      It follows that the title of the paragraph is also erroneous.

      (R3. A7a4) The paragraph title supports our conclusions and we will keep it.

      (b) The hearing threshold is underestimated by 40dB at 6 and 8Kz on Fig 2C, not by "10-20dB" as reported in the text (L178).

      (R3. A7b) We have changed the title of this section and moved the last sentence to the discussion to remove the focus on heat whistles. We added a paragraph in the discussion to specifically address the difference between ABR and behaviorally measured audiograms (Lines 413-422).

      (B) Egg vibration experiment

      (8) Using airborne sound to vibrate eggs is biologically irrelevant:

      (R3. A8.1) We agree that parental contact could influence vibration transmission.

      However, (1) prior studies assume airborne sound transmission, and (2) our experiment tests this assumption directly. We now clarify this scope (Lines 328-342 and Figure 4) and discuss contact-based transmission as a potential future direction.

      The measurement of airborne sound levels to vibrate eggs misunderstands bone conduction hearing and is not biologically meaningful: zebra finch parents are in direct contact with the eggs when producing heat calls during incubation, not hovering in front of the nest. This misunderstanding affects all extrapolations from this study to findings in studies on prenatal communication.

      (R3. A8.2) The definition of bone conduction is the response to sound that is not mediated by a functional middle ear, but through the skull. In the earlier study, the eggs were stimulated by sound from a headphone, so that is the reason for using the same stimulation here. See also joint response A3.5 above.

      (C) Misrepresentation of current knowledge

      (9) Values from published papers are misreported, which reverses the conclusions:

      (R3. A9) We thank the reviewer for identifying inconsistencies and have:

      - Corrected heat whistle frequency ranges consistently through our paper

      - Added a comprehensive comparison figure gathering all available audiograms (Fig 5)

      - Expanded discussion of high-frequency hearing.

      These revisions do not alter our conclusions.

      Most critical examples:

      (a) Preprint: "Zebra finch most sensitive hearing range of 1-to-4 kHz (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)" (L173).

      Actual values in the studies cited are:

      1-to-7kHz, in Amin et al 2007 (threshold [=50dB with ABR] is the same at 7kHz and 1KHz).

      1-to-6 kHz, in Okanoya and Dooling (the threshold [=30dB with behaviour] is actually lower at 6kHz than at 1KHz).

      1-to-7kHz, in Yeh et al (threshold [=35-38dB with behaviour] is the same at 7kHz and 1KHz).

      (R3. A9a.1) In this sentence presenting our results (“sensitive hearing range of 1-to-4 kHz”) we originally wrote that “these are consistent with the following papers (Amin et al., 2007; Okanoya and Dooling, 1987; Yeh et al., 2023)". This latter part was left out during the writing process. This explains the different numbers. We apologize for this mistake.

      To avoid confusion, in our revision we have placed all ABR curves together into new Fig 5 and have included a new paragraph to discuss the differences (Lines 387-422).

      Note that zebra finch nestlings' begging calls peaking at 6kHz (Elie & Theunissen 2015, doi: 10.1007/s10071-015-0933-6), would fall 2kHz above the parents' best hearing range if it were only up to 4kHz.

      (R3. A9a.2) Of course that is possible. However this representation is incorrect because begging calls are harmonic sounds with a fundamental frequency around 500 Hz and formant at 6 kHz. Begging calls thus contain lots of energy at frequencies below 6 kHz, while the heat whistles do not. The peak frequency of heat whistles is also their lowest frequency component.

      (b) The preprint incorrectly states throughout (e.g., L139, L163, L248) that heat-calls are 7-10kHz, when the actual value is 6-10kHz in the paper cited (Katsis et al, 2018).

      (R3. A9b) The authors in Katsis et al. 2018 provided a range of 6-10 kHz estimated from the spectrogram without any further specification of methods. In another manuscript, we have quantified the heat whistle frequency (Anttonen et al Curr Biol https://doi.org/10.1016/j.cub.2025.08.054) to be 6.8 ± 0.6 kHz. We have changed this accordingly throughout our manuscript.

      (c) Using the correct values from these studies, and heat-calls at 45 dB SPL (as measured by others (unpublished data), or as measured by the authors themselves, but which is not reported here (Anttonen et al 2025), the correct conclusion is that heat calls fall within the known zebra finch hearing range.

      (R3. A9c) Please see our answer R3.A5a. We have included this argument more clearly in our discussion (Lines 271-355), illustrated by new Fig. 4.

      (10) Published evidence towards high-frequency hearing, including in early development, is systematically omitted:

      (a) Other studies showing birds use high frequencies above the known avian hearing range are ignored. This includes oilbirds (7-23kHz; Brinklov et al 2017; by 1 of the preprint authors, doi: 10.1098/rsos.170255) and hummingbirds (10-20kHz; Duque et al 2020, doi: 10.1126/sciadv.abb9393), and in a lesser extreme, zebra finches' inspiratory song syllables at 57kHz (Goller & Dalley, 2001).

      (R3. A10a) We agree that some bird species produce or use acoustic signals extending into high frequencies. However, signal production is not evidence of perceptual sensitivity. Many animals, including birds and mammals, produce signals that contain harmonic or broadband components extending beyond their most sensitive hearing range without implying functional detection at those frequencies.

      The cited examples (oilbirds, hummingbirds, inspiratory song syllables in zebra finches) concern signal production or ecological specializations in different species, not measured auditory sensitivity in zebra finches, and particularly not during early development. As such, they do not provide evidence that zebra finches—adults or embryos—can detect low-amplitude, narrowband signals in the 6–7 kHz range.

      Our study explicitly addresses auditory sensitivity using physiological measurements, which is the relevant metric for evaluating detectability.

      (b) The discussion of anatomical development (L228-241) completely omits the well-known fact that the avian basilar papilla develops from high to low frequencies (i.e., base to apex), which - as many have pointed out - is opposite to the low-to-high development of sensitivity (e.g., cited: Cohen & Fermin 1978; Caus Capdevila et al 2021).

      (R3. A10b) We agree that the avian basilar papilla develops from base to apex (high to low frequency). We have now added a sentence in the Discussion to acknowledge this (Lines 406411).

      Importantly, morphological development does not directly translate to functional sensitivity. Functional hearing depends critically on factors such as hair cell innervation, synaptic maturation, and central auditory processing, which are known to develop over time.

      Our data show a low-to-high frequency progression in functional sensitivity, consistent with previous physiological studies. This apparent mismatch between anatomical gradients and functional onset has been noted in other systems and likely reflects the later maturation of neural encoding rather than hair cell differentiation per se. We now clarify this distinction in the revised manuscript (Lines 406-411).

      (c) High frequency hearing in songbirds at hatching is several orders of magnitude better than in chickens and ducks at the same age, even though songbirds are altricial (e.g., at 4kHz, flycatcher: 47dB, chicken-duck: 90dB; at 5kHz, flycatcher: 65dB, chicken-duck: 115dB; Korneeva et al 2006, Saunders et al 1974). That is because Galliformes are low-frequency specialists, according to both anatomical and ecological evidence, with calls peaking at 0.8 to 1.2kHz rather than 2-6kHz in songbirds. It is incorrect to conclude that altricial embryos cannot perceive high frequencies because low-frequency specialist precocial birds do not (L250;261).

      (R3. A10c) We agree that species differ in their auditory ecology and frequency specialization, and we do not claim that all altricial birds share identical developmental trajectories.

      However, the cited comparisons involve different species, methodologies, and developmental timelines, which limits their direct comparability. In particular:

      Developmental staging is not directly comparable across species using days post-hatch alone.

      - Different methods (e.g., invasive recordings vs. ABR vs behavioural assays) yield systematically different thresholds.

      - Ecological specialization (e.g., low-frequency vs. broadband species) influences adult audiograms and likely developmental trajectories.

      We have revised the Discussion to explicitly acknowledge these limitations and to avoid overgeneralization across species. Importantly, our conclusions are based on within-species comparisons (adult vs. hatchling zebra finches) combined with measured signal levels of heat whistles. These constraints are sufficient to evaluate detectability without relying on cross-species extrapolation.

      (11) Incorrect statements do not reflect findings from the references cited For example:

      (a) "in altricial bird species hearing typically starts after hatching" (L12, in abstract), "with little to no functional hearing during embryonic stages (Woolley, 2017)." (L33).

      There is no evidence, in any species, to support these statements. This is only a - commonly repeated - assumption, not actually based on any data. On the contrary, the extremely limited evidence to date shows the opposite, with zebra finch embryos showing ZENK activation in the auditory cortex in response to song playback (Rivera et al, 2018, not cited).

      The book chapter cited (Woolley 2017) acknowledges this lack of evidence, and, in the context of song learning, provides as only references (prior to 2018), 2 studies showing that songbirds do not develop a normal song if the song tutor is removed before 10d post-hatch. That nestlings cannot memorise (to later reproduce) complex signals heard before d10 does not mean that they are deaf to any sound before day 10.

      Studies showing hearing in young songbird nestlings (see point 6 above) also contradict these statements.

      (R3. A11a) We agree that the precise onset of hearing in altricial embryos is not well established. We have therefore revised the wording in the Abstract and Introduction to avoid categorical statements and instead reflect the limited available evidence (Lines 13-16 and 33-37).

      Our data provide direct physiological measurements showing extremely low sensitivity immediately after hatching, which constrains the likelihood of functional hearing in earlier embryonic stages.

      Regarding the cited ZENK study, we note that immediate early gene expression indicates neural activation but does not provide a measure of auditory sensitivity or detection thresholds. As such, it cannot be directly compared to physiological or behavioral measures of hearing.

      (b) "Zebra finch embryos supposedly are epigenetically guided to adapt to high temperatures by their parents high-frequency "heat calls" " (L36 and L135).

      This is an extremely vague and meaningless description of these results, which cannot be assessed by readers, even though these results are presented as a major justification for the present study. Rather than giving an interpretation of what "supposedly" may occur, it would be appropriate to simply synthesize the empirical evidence provided in these papers. They showed that embryonic exposure to heat-calls, as opposed to control contact calls, alters a suite of physiological and behavioural traits in nestlings, including how growth and cellular physiology respond to high temperatures. This also leads to carry-over effects on song learning and reproductive fitness in adulthood.

      (R3. A11b) We thank the reviewer for raising this point. In the revised manuscript, we have replaced the previous phrasing with a more precise and neutral summary of what these studies report, namely that embryonic exposure to heat-call playbacks has been associated with differences in physiological and behavioral traits.

      Our study, however, addresses a distinct question—whether such acoustic signals are detectable by embryos given known constraints on signal amplitude and auditory sensitivity. The cited studies do not directly quantify auditory perception or the physical sound environment experienced by embryos. As a result, they do not provide a direct test of the sensory mechanism required for acoustic communication. A detailed evaluation of experimental design and interpretation in those studies is beyond the scope of the present manuscript, and we therefore limit our discussion to assessing the biophysical and physiological plausibility of the proposed mechanism.

      (c) "The acoustic communication in precocial mallard ducks depends specifically on the lowfrequency auditory sensitivity of the embryo (Gottlieb, 1975)" (L253)

      The study cited (Gottlieb, 1975) demonstrates exactly the opposite of this statement: it shows that duckling embryos, not only perceive high frequency sounds (relative to the species frequency range), but also NEED this exposure to display normal audition and behaviour post-hatch. Specifically, it shows that duckling embryos deprived of exposure to their own high-frequency calls (at 2 kHz), failed to identify maternal calls post-hatch because of their abnormal insensitivity to higher frequencies, which was later confirmed by directly testing their auditory perception of tones (Dimitrieva & Gottlieb, 1994).

      (R3. A11c) We thank the reviewer for this clarification and have revised the relevant text. Our intention was to highlight that embryonic auditory experience can shape postnatal behavior, not to imply strict low-frequency limitation. Therefore we already included the actual frequency in the original sentence. We have removed the non-descriptive term “low-frequency” (Lines 330-332).

      (12) Considering all of the mistakes and distortions highlighted above, it would be very premature to conclude, based on these results and statements, that altricial avian embryos are not sensitive to sound. This study provides no actual scientific ground to support this conclusion.

      (R3. A12) We respectfully disagree with the reviewer’s conclusion.

      Our study does not make a general claim that altricial embryos are incapable of perceiving sound. Rather, we evaluate a specific hypothesis: whether zebra finch embryos and hatchlings can detect sound and parental heat whistles.

      Our conclusions are based on the convergence of:

      (1) Measured low sound pressure levels of heat whistles,

      (2) Established adult auditory thresholds (behavioral data),

      (3) A large developmental decrease in auditory sensitivity demonstrated by our ABR measurements.

      Even under conservative assumptions, these constraints place heat whistles at or below adult detection thresholds and far below expected sensitivity in hatchlings and embryos.

      Thus, our conclusion is not based on absence of evidence, but on quantitative constraints that make detection unlikely under biologically realistic conditions.

      Recommendations for the authors:

      Joint recommendations:

      In response to the joint recommendations, we have:

      - Expanded methodological transparency (temperature, electrode setup, stimulus parameters),

      - Added new data (Fig S3) and figures (Fig 4 and 5),

      - Clarified ABR limitations and interpretation,

      - Strengthened the separation between measured results and interpretation,

      - Reframed conclusions to avoid overstatement.

      These revisions leave the central two conclusions unchanged: 1) zebra finch hatchlings and embryos are functionally deaf, and 2) under biologically realistic conditions, heat whistles are unlikely to be detectable by zebra finch hatchlings or embryos.

      (A) Reviewers 1 and 2:

      Much of the reviewer discourse revolved around providing clarifications of methodology for measuring the ABR and caveats for interpretation. There was near consensus with reviewers 1 and 2 on issues related to the ABR, which should be addressed.

      We appreciate the reviewers’ consensus that the main conclusions are supported, while requesting clarification of methodological details and interpretation of ABR measurements.

      (1) Please address all of the issues raised by reviewers 1 and 2 above.

      (A1) All points raised by Reviewers 1 and 2 have been addressed in detail in our point-by-point rebuttal below. In addition, we have revised the manuscript to improve clarity, added new figures (Fig. 4, 5), and substantially expanded the Discussion with eight new paragraphs to better contextualize our findings.

      (2) Please also

      - clarify all aspects of experimental details of the ABR that were missing, including temperature control (estimate body and ambient temperatures during ABR recordings,

      - please address the possibility of hypothermia of hatchlings that could have reduced ABR responses,

      - and potential local head cooling due to surgical exposure and its likely effect on highfrequency response depression).

      (A2) In our revision, we have expanded the Methods section (L479-485) and added the following new data:

      Body and ambient temperature/hypothermia

      We have now included the body temperatures during ABR recordings in new table S6. These data show that:

      (1) Body temperature was stable throughout recordings,

      (2) Temperatures were within the physiological range,

      (3) Conditions were consistent across all age groups.

      Importantly, even the youngest hatchlings maintained stable temperatures and showed no indication of hypothermia. Therefore, differences in ABR responses cannot be attributed to temperature effects.

      Potential cooling due to surgical exposure

      This concern does not apply to our experiments. We used subdermal needle electrodes, which do not require surgical exposure. Therefore, no local cooling of the head occurred, and no tissue exposure could affect high-frequency sensitivity. We have added a clarifying sentence in the Methods section to explicitly state this (Line 501-503).

      (B) Reviewer 3 also had additional requests for clarification that should also be addressed:

      (3.1) Stimulus duration too long: The study used 25 ms tone bursts instead of the standard {less than or equal to} 5 ms "pips." Could this prevent reliable ABR detection, especially at high frequencies?

      (A3.1) We agree that stimulus duration affects ABR characteristics and now clarify our rationale in the manuscript.

      - The 25 ms tone bursts were deliberately chosen to ensure sufficient cycle representation at low frequencies (down to 250 Hz) and to avoid frequency splatter.

      - Using a constant duration across frequencies ensures comparable stimulus energy.

      Importantly:

      - The 25 ms data yield audiogram shapes consistent with both click responses (Fig 2C) and published behavioral data (new Fig 5).

      - To address potential high-frequency limitations, we included a dataset using 5 ms tone pips, which produced thresholds consistent with published ABR studies (new Fig 5).

      Thus, both stimulus types support the same conclusion: a gradual maturation of hearing sensitivity from low to high frequencies. We have expanded the Discussion with three paragraphs to clarify these methodological trade-offs (Lines 387-422).

      (3.2) Were 400 sweeps enough averaging? Might a signal appear at 1000 or more?

      (A3.2) Signal-to-noise ratio improves with the square root of the number of averages. Increasing from 400 to 1000 sweeps would therefore reduce thresholds by at most ~4 dB. This magnitude is small relative to the >54 dB developmental differences observed, and the large gap between signal levels and detection thresholds. Thus, increasing sweep number would not alter the conclusions. We now clarify this explicitly in the Methods (L530).

      (3.3) Was the repetition rate too high? How does the stimulus presentation affect the ABR? Might a signal have emerged with 1-4 per second?

      (A3.3) We have clarified stimulus presentation rates in the revised manuscript:

      - Clicks were presented at 25 Hz. Control measurements (now included as Supplementary Fig. S3) show no effect of this rate on ABR amplitude or threshold.

      - Tone bursts and pips were presented at ~3 Hz, consistent with commonly used rates that avoid neural adaptation. We apologize for leaving this out in our original submission.

      We now explicitly describe these parameters and their rationale in the Methods (Lines 543-552).

      (3.4) If possible, provide an estimate of the effective bandwidth of the tone pips and compare it with the bandwidth of the parental heat-whistles.

      (A3.4) We agree that stimulus bandwidth differs between tone pips and heat whistles, and that broader signals may stimulate multiple auditory filters. Shorter stimuli (e.g., 5 ms pips) have broader bandwidth and may stimulate multiple filters—particularly at low frequencies—potentially lowering thresholds, whereas longer stimuli (25 ms bursts) are more frequency-specific and may yield higher thresholds. At higher frequencies (including the heat whistle range), this effect is expected to be smaller.

      However, quantitative correction is currently not possible due to a lack of species-specific data on auditory tuning curves in zebra finches. The only available avian data (budgerigar; Saunders et al., 1979) suggest auditory filter bandwidths (Q10 dB, i.e., the bandwidth 10 dB below the peak divided by peak frequency) of ~1.4 at low frequencies and ~1000 Hz at higher frequencies, but how multifilter stimulation affects thresholds is unknown and likely species- and frequency-dependent.

      Given these uncertainties, direct comparison between tone stimuli and heat whistles requires strong assumptions. We therefore suggest that future studies should measure responses to natural heat whistles directly.

      (3.5) Egg Vibration Experiment. Address the possibility that if a parent were physically lying on top of an egg and generated a heat call, parental body vibration could significantly communicate some perceptual vibrotactile signal to the egg. Reviewer 3 raised the possibility that the experiments in this paper tested the extent to which an auditory input can vibrate the egg - what if a vocalizing bird was on the egg?

      (A3.5) We agree that embryos may receive multiple types of sensory input from parents, including direct mechanical cues.

      However, our experiment specifically tests the hypothesis proposed in prior work: that airborne sound (heat whistles) induces egg vibrations sufficient for perception. Our findings show that airborne sound-induced vibrations are orders of magnitude below known vibrotactile sensitivity thresholds.

      Regarding parental contact:

      - Heat whistles are produced by an aerodynamic whistle mechanism, not tissue vibration (Anttonen et al., Curr Biol 2025), meaning most respiratory energy is radiated as sound rather than dissipated as heat/vibration in the body.

      - A parent sitting on the egg would attenuate airborne sound transmission, not amplify it.

      We now clarify in the Discussion that other cues (e.g., respiration, direct contact, temperature) may exist and need to be included in new experiments (Lines 349-352). Even so, these are distinct mechanisms and were not the hypothesis tested in prior playback studies.

      In our paper, we will not add a detailed discussion of these prior papers as this is outside the scope of this paper. Instead, we added a paragraph what would be a constructive way forward (Lines 349-355). Hopefully somebody in the community will have the good fortune to secure research funding to continue this benchmarking work.

      (4) Finally, all reviewers agreed that some more context on the ABR and its relationship to functional hearing could be provided, with less direct focus on the heat-call experiments.

      (A4) We agree and in the revision discussion have compiled all published ABR and behavioral audiograms (Fig. 5) and added new paragraphs on functional hearing (Lines 261-269), and ABR vs behavioural audiograms (Lines 413-422).

      Furthermore, to remove focus on the heat whistles, we have moved all heat-whistle-specific interpretation out of the Results into a single, focused Discussion section (Line 271-355).

      The behavioral studies mentioned below show auditory responses (e.g., begging suppression). However, these behaviors are tested between ~5–10 days post-hatch (consistent with our findings), in different species, and do not provide quantitative sensitivity thresholds, nor do they address detectability of low-amplitude, high-frequency signals like heat whistles.

      In our revision we added a new discussion paragraph including these behavioral studies (Lines 357-367).

      For example, there are ample cases in the literature of altricial birds exhibiting behavioral evidence of auditory sensitivity by reducing begging calls in response to parental alarm calls:

      Platzen & Magrath (2004) - Playback of parental alarm calls nearly abolished nestling non-begging calls and reduced begging in scrubwrens. Proc. R. Soc. B 271:1271-1276.

      Different species: scrubwrens. Playback age: 5-, 8- and 11-DPH nestlings.

      Magrath, Haff, Horn & Leonard (2006) - Review and experiments on the developmental shift to silence/freeze after aerial alarm calls as chicks become fledglings; documents nestling quieting to alarms. Proc. R. Soc. B 273:2335-2341.

      Different species: scrubwrens. Playback age: 7-9 DPH nestlings, and 2- 4 days after fledging.

      Magrath, Pitcher & Dalziell (2007) - Nestlings respond to the sound of a predator's footsteps and parental food/alarm calls; includes begging suppression following predator sounds. Anim. Behav. 74:1117-1129.

      Different species: scrubwrens. Playback age: 8 DPH nestlings.

      Haff & Magrath (2012) - Nestlings suppress calling after heterospecific alarm calls (when acoustically similar to conspecific alarms), indicating generalized auditory danger recognition. Anim. Behav. 84:e.g., 495-505 (article).

      Different species: scrubwrens. Playback age: 5-6 and 10-11DPH nestlings. They show that 10-11 days old suppress calling while 5-6 days old do not.

      Barati & McDonald (2017) - Noisy miner nestlings suppress begging after conspecific alarm calls and some heterospecific cues; stronger/longer suppression for terrestrial-predator alarms. Sci. Rep. 7:9563.

      Different species: Noisy miner (Manorina melanocephala ). Playback age: 14 DPH. Nestlings started to vocalise at 5 DPH.

      Suzuki (2011) - In Paridae, parental alarm calls encode predator type; prior work (cited within) shows young of altricial species suppress vocalizations to alarms. Curr. Biol. 21:15-20.

      Different species: great tits. Playback age: 17 DPH.

      Can you please contextualize the present results about the timing of auditory development with the above body of work with respect to the timing of alarm call-induced begging call suppression?

      In our revision we have added a new paragraph in the discussion on these papers (Lines 357-367), and highlight the need for comparative work on hearing development in different species (Lines 352-355 and Lines 405-411).

      (5) Strictly speaking, a flat ABR does not equal deafness - at the extreme, an average of 10,000 trials may pull out a minuscule signal. Thus, the more rigorous path would be, in the results section, to ensure that statements summarize the data as they are, representing an absence of a neural signal.

      Save the interpretation of what this may mean for the discussion, and provide alongside this interpretation the necessary caveats related to temperature, rendition rate, averaging, etc.

      Clarify the conditions where a flat ABR demonstrates or fails to demonstrate immature deafness.

      Expand clarification for how the known 20-40 dB difference between ABR and behavioral thresholds can exist if a flat ABR can be interpreted as deafness.

      Consider refraining from concluding deafness from a flat ABR. Discuss that behavioral, single-unit, or alternative physiological assays might detect responses below the ABR threshold. If such cases exist, cite.

      (A5) We thought about this considerably before starting our measurements. What constitutes the absence of a signal? Even with intracellular recordings of all but one of the auditory neurons, the last one could still contain a signal and theoretically transmit information to the nervous system. We agree with the reviewers that absence of an ABR response should not be equated with absolute deafness. We have revised the manuscript accordingly and removed all statements implying “deafness” from the Results. The Results now strictly report presence or absence of detectable ABR responses.

      However, in both clinical and comparative contexts, absence of ABR responses at high SPLs (e.g., 90–95 dB) is widely interpreted as functionally non-responsive hearing. The developmental shift we observe (>54 dB) is far larger than typical ABR–behavioral offsets (20–40 dB). In the Discussion, we have added a new paragraph arguing that we think that the term functional deafness is reasonable here (Lines 261-269).

  2. www.researchsquare.com www.researchsquare.com