Reviewer #2 (Public review):
In this manuscript, Ritter et al. propose a model of working memory (WM) that combines feedforward and rotational dynamics. The model is discovered by optimizing a linear RNN using a loss function that encourages maximization of signal-to-noise ratio (SNR) and minimization of activation magnitude. The authors argue that the optimized model outperforms other WM models in terms of SNR and energetic efficiency, while also better replicating key features of neural responses recorded in monkey pre-frontal cortex (PFC) during a WM task. The authors also draw connections to state space models (SSM) used for other machine learning applications.
My main issue with this manuscript is that it does not appear to convincingly demonstrate that rotational dynamics offer any advantage over purely feedforward dynamics. The authors adopt three criteria according to which they compare models:<br /> (1) SNR.<br /> (2) Energy efficiency.<br /> (3) Similarity to neural data.
In terms of SNR, purely feedforward models seem to perform similarly to the optimized models (Figure 1). Figure 1 does seem to show that the optimized network produces responses of smaller magnitude when the number of units is large, but the authors do not explain why adding rotational dynamics would produce such a relationship. In fact, the responses that are plotted for the feedforward network in Figures 1B, 2C, and 5E look similar, if not smaller in magnitude than those of the optimized model. Lastly, while the authors claim in the body of the text that the optimized model replicates key features of monkey PFC responses better than the purely feedforward model, this is not apparent to me from the comparisons plotted in Figure 5E-J. The authors thus do not show strong evidence that the model they propose beats what they claim is an established baseline on any of the three criteria.
Another weakness of the manuscript is that the comparison to attractor and feedforward models seems somewhat unfair. In Figure 1, the rotational model is optimized, while the parameters for the attractor and feedforward models seem to have been at least partially chosen by hand. Figure 5C again shows the three models side by side, but the fact that it compares the same network at different stages during training complicates the comparison. Instead, one should compare the rotational solution to the optimal attractor and feedforward models, respectively (obtained by constrained optimization). From looking at the flow-fields, it seems that a feedforward network with an optimized level of amplification may work just as well. On a mechanistic level, it is unclear what computational advantage rotations offer over feedforward dynamics in the WM context.
The choice of baseline models to compare against might be questionable. The simple line attractor model by Seung et al. (1996) was initially designed to explain oculomotor integration. It is true that a line attractor has been suggested as a mechanism for working memory, e.g., in the seminal work by Machens et al (2005). However, it seems fair to say that most studies employing non-linear networks have focused on point attractors as mechanisms of working memory (e.g., Wong & Wang, 2006; Driscoll, Shenoy, Sussillo, 2024). A point attractor arguably does not suffer the SNR issues of a line attractor, because it does not lead to integration of the noise over time. However, non-trivial point attractors cannot be implemented in linear networks of the kind studied by the authors of the present study.
The authors should expand their discussion to include other, potentially closely related work proposing rotation-like dynamics in artificial neural networks during working memory. In particular, the manuscript does not discuss Sharma, Proca, et al, ICML 2026, which describes a rotational solution to a similar WM task obtained by optimizing linear RNNs (Sharma et al., 2026, Fig. 6). Notably, Sharma et al. arrive at a similar rotational (and likely also non-normal) mechanism without using either noisy inputs or a constraint on energy efficiency. The authors of the present manuscript should discuss to what extent this finding contradicts their claim that "normative pressures on noise-robustness and energetic cost shape the complex dynamics of WM circuits." (present manuscript, Introduction). Given the obvious parallels between the two studies, a comparison between the present work and Sharma et al. (2026) would add necessary context to the Discussion.
The authors should also clarify the significance of the "novel method for optimization of continuous-time RNNs driven by noisy inputs" (see Discussion) that the authors propose. This method is mentioned in the first line of the Discussion section but is barely discussed, let alone sufficiently explained, in the previous Sections. The only time a comparison to BPTT with a simple MSE loss is mentioned, it is stated that the two procedures produce the same results. The novel method appears to consist of a loss with two terms, the second of which is a well-known L2-penalty on unit activations (Sussillo et al., 2015). It is not clear that the method is either novel or necessary to obtain the reported results.
Except for the fact that higher-dimensional networks also converge on rotational solutions, Figure 3 does not add much to the reader's understanding of the optimized model (except for panel F). I find the comparison to SSMs too superficial to provide real insight.
Figure 4 claims to show that the optimized model recapitulates "a range of properties observed in prefrontal cortex and other brain areas during WM tasks" (p. 7) but does not show neural data for comparison.
