Reviewer #1 (Public Review):
Summary:<br /> In this study, the authors investigate a very interesting but often overlooked aspect of abstract vs. concrete processing in language. Specifically, they study if the differences in processing of abstract vs. concrete concepts in the brain are static or dependent on the (visual) context in which the words occur. This study takes a two-step approach to investigate how context might affect the perception of concepts. First, the authors analyze if concrete concepts, expectedly, activate more sensory systems while abstract concepts activate higher-order processing regions. Second, they measure the contextual situatedness vs. displacement of each word with respect to the visual scenes it is spoken in and then evaluate if this contextual measure correlates with more activation in the sensory vs. higher-order regions respectively.
Strengths:<br /> This study raises a pertinent and understudied question in language neuroscience. It also combines both computational and meta-analytic approaches.
Weaknesses:<br /> Overall, the study had many intermediary steps that required manual subsection / random sampling and variable choices (like the time lag of analysis) with almost no visualization and interpretation of how these choices affect the observed results. The approach was also roundabout.
Peaks and Valleys Analysis:<br /> 1. Doesn't this method assume that the features used to describe each word, like valence or arousal, will be linearly different for the peaks and valleys? What about non-linear interactions between the features and how they might modulate the response?<br /> 2. Doesn't it also assume that the response to a word is infinitesimal and not spread across time? How does the chosen time window of analysis interact with the HRF? From the main figures and Figures S2-S3 there seem to be differences based on the timelag.<br /> 3. Were the group-averaged responses used for this analysis?<br /> 4. Why don't the other terms identified in Figure 5 show any correspondence to the expected categories? What does this mean? Can the authors also situate their results with respect to prior findings as well as visualize how stable these results are at the individual voxel or participant level? It would also be useful to visualize example time courses that demonstrate the peaks and valleys.
Estimating contextual situatedness:<br /> 1. Doesn't this limit the analyses to "visual" contexts only? And more so, frequently recognized visual objects?<br /> 2. The measure of situatedness is the cosine similarity of GloVE vectors that depend on word co-occurrence while the vectors themselves represent objects isolated by the visual recognition models. Expectedly, "science" and the label "book" or "animal" and the label "dog" will be close. But can the authors provide examples of context displacement? I wonder if this just picks up on instances where the identified object in the scene is unrelated to the word. How do the authors ensure that it is a displacement of context as opposed to the two words just being unrelated? This also has a consequence on deciding the temporal cutoff for consideration (2 seconds).<br /> 3. While the introduction motivated the problem of context situatedness purely linguistically, the actual methods look at the relationship between recognized objects in the visual scene and the words. Can word surprisal or another language-based metric be used in place of the visual labeling? Also, it is not clear how the process identified in (2) above would come up with a high situatedness score for abstract concepts like "truth".<br /> 4. It is a bit hard to see the overlapping regions in Figures 6A-C. Would it be possible to show pairs instead of triples? Like "abstract across context" vs. "abstract displaced"? Without that, and given (2) above, the results are not yet clear. Moreover, what happens in the "overlapping" regions of Figure 3?
Miscellaneous comments:<br /> 1. In Figure 3, it is surprising that the "concrete-only" regions dominate the angular gyrus and we see an overrepresentation of this category over "abstract-only". Can the authors place their findings in the context of other studies?<br /> 2. The following line (Pg 21) regarding the necessary differences in time for the two categories was not clear. How does this fall out from the analysis method?<br /> 3. Both categories overlap **(though necessarily at different time points)** in regions typically associated with word processing.