2 Matching Annotations
  1. Last 7 days
    1. Signed review contributed through PaperStack

      This review was contributed through PaperStack and is reproduced with the reviewer’s permission.


      1. Benchmarking methodology and expert consensus

      The benchmarking methodology would benefit from additional detail. It is unclear how many experts contributed to the “human expert consensus” used in the curation-agreement benchmarks (Fig. 4b and Tables 1–2). The number of reviewers is specified only for the separate time-efficiency experiment (Fig. 5d–e), and it is not clear whether the same reviewers generated the consensus labels.

      The manuscript also does not explain how disagreements were resolved—for example, by majority vote—or whether the few-shot examples provided to the VLM came from recordings independent of the test data, which is important for excluding information leakage.

      Because inter-expert agreement is itself only 77.8–83.3% (Fig. 5e), the expert consensus should be treated as an uncertain reference standard rather than biological ground truth. The evaluation would be strengthened by reporting the number of experts labeling each unit, the consensus procedure, individual expert–model agreement, and the sensitivity of performance to the choice of expert reference.

      Reporting precision, recall, sensitivity, specificity, and confusion matrices in addition to overall percentage agreement would also clarify performance, particularly given the apparent Good/Noise class imbalance.

      2. Breadth of validation

      Validation is currently limited to two short recordings from two animals: a one-minute Neuropixels segment and a two-minute flexible-probe segment, each obtained from a single mouse.

      Testing on held-out recordings from additional animals, laboratories, brain regions, recording durations, and signal-quality conditions would better establish whether the VLM’s curation criteria generalize beyond the recordings used to construct its prompts and few-shot examples.

      Validation against experimentally derived ground truth, where available, or comparison with established automated quality-control tools would also help distinguish agreement with human judgment from unit-classification accuracy. This distinction is important because the model and human reviewers could share the same systematic biases.

      3. Scope of the “fully autonomous end-to-end” framing

      The “fully autonomous end-to-end” framing may require qualification. For the LLM-backend benchmark in Fig. 3, the Methods state that SpikeInterface operations such as spike sorting and waveform extraction were replaced by loading presaved results and waveforms.

      This is a reasonable design for isolating the models’ code-generation and tool-chaining reliability, but the resulting completion-rate and token-usage measurements do not evaluate a complete raw-data-to-curated-output execution.

      The manuscript would be clearer if it explicitly described Fig. 3 as a workflow-orchestration benchmark rather than a full end-to-end benchmark. Alternatively, a separate evaluation could run the complete pipeline on raw data and assess both successful execution and the validity of the resulting sorted and curated units.

      Clarifying this distinction would help readers calibrate expectations about what “autonomous” performance has actually been demonstrated versus what remains to be tested.


      This review is signed, but the reviewer’s name is omitted because bioRxiv’s community guidelines do not permit reviewer names in comments. To contact the reviewer, email support@paperstack.pub; PaperStack will forward the request.

      PaperStack · The trust layer for preprints

  2. Jul 2026
    1. An external discussion of this preprint is available on PubPeer: https://pubpeer.com/publications/8DA4B0AA40D8625B196C1014E32ED3

      The discussion focuses on the study’s cross-validation design, including whether random rather than contiguous splits could allow temporal dependencies to influence the reported encoding results. It also considers how the language features were defined, how the findings compare with concept-cell coding, and how strongly the results support interpretations involving pattern separation and completion. Readers may wish to consult the original discussion for the full context.

      paperstack.pub · The trust layer for preprints