Reviewer #1 (Public review):
Summary:
This study combines representational similarity analysis (RSA) with 7T layer-specific fMRI and EEG to examine how neural representations in specific cortical layers of EVC and LOC correspond to the temporal dynamics of visual processing. The authors interpret these correspondences as reflecting feedforward and feedback processes, based on their relative timing and their similarity to representations in different layers of a deep neural network (DNN).
Strengths:
The combination of RSA with laminar fMRI is a promising approach for dissociating the functional roles and dynamics of different cortical layers within the same functional region, and it holds considerable potential for elucidating computational mechanisms both within and between levels of the visual hierarchy. However, several issues should be addressed before the authors' conclusions can be fully supported.
Weaknesses:
(1) The authors report that the representation in the LOC superficial layer resembles EEG-derived neural representations at ~400 ms post-stimulus, and that this similarity is best explained by representations in the higher layers of the DNN. From these two observations, they conclude that activity in the LOC superficial layer is driven by feedback signals. However, neither line of evidence directly dissociates feedforward from feedback contributions.
Specifically, late-stage representations in LOC could instead reflect the outcome of local recurrent computation, given that the superficial layer also serves as an output layer of the local cortical circuit. Moreover, the correlation with the DNN peaks at higher layers rather than being dominated by them, and feature tuning in higher DNN layers does not necessarily map onto higher-order cortical regions such as PFC.
While a feedback contribution to the LOC superficial layer is consistent with theoretical predictions and known cortical anatomy, the current evidence is indirect. I would recommend that the authors either tone down this conclusion or, at a minimum, explicitly clarify the strength and limitations of the evidence in the Discussion.
(2) I could not find information regarding the fMRI slice orientation or whether temporal regions beyond LOC were covered. The reported FOV (192 × 192 mm) seems quite large if only EVC and LOC were targeted. Did the authors acquire data from other object-selective regions in the temporal cortex, and if so, did they analyze these?
It would strengthen the feedback interpretation considerably if the RDM of the LOC superficial layer could be shown to resemble RDMs from more anterior temporal regions, which would be consistent with feedback originating from higher-order object-processing areas.
(3) Related to the previous point, LOC is a relatively large region, and based on the figures, it appears that the LOC ROI may contain two subregions. It would be helpful for the authors to show the location and extent of the LOC ROI in example participants.
If the ROI does indeed span two subregions, do these subregions share the same laminar profile and temporal dynamics?
(4) The authors report no feedback-related information in EVC, which contrasts with a number of prior fMRI studies that have demonstrated object-related feedback signals in EVC. One plausible explanation for this discrepancy is task relevance: in the present study, participants performed only a fixation color-change task, whereas in previous work they were required to attend to object features or identity (e.g., Morgan et al., 2019, J Neurosci; Kok et al., 2016, Curr Biol; Mohsenzadeh et al., 2018, eLife; Hou et al., 2026, eLife). Task demands on object processing may substantially modulate the strength of feedback signals to EVC, and this possibility warrants discussion.
(5) A substantial body of work has used specialized paradigms to dissociate feedforward and feedback signals in EVC (e.g., Williams et al., 2008, Nat Neurosci; Fan et al., 2016, PNAS; Hou et al., 2026, eLife). These studies are directly relevant to the current work but are not cited.
(6) Multidimensional scaling (MDS) visualizations of the RDMs (as in, e.g., Mohsenzadeh et al., 2018) are not included in the manuscript. These visualizations are important for interpreting the representational format across different layers of LOC and EVC, and I would encourage the authors to include them.