39 Matching Annotations
  1. Aug 2026
      1. Beyond cell annotation and spatial niche characterization, recovering unmeasured456proteins represents another important downstream application in spatial proteomics.

      How exactly would the authors go about tackling this given the approach presented in the manuscript?

      1. vely, these analyses identify a reproducible collagen-rich ECM-remodeling449niche that can be robustly transferred across independent cohorts and is preferentially450enriched in recurrent colorectal cancer

      Important application

      1. The resulting embedding space exhibited well-organized continuous manifolds,

      Surprising to see a continuous manifold.

    1. Beyond cell annotation and spatial niche characterization, recovering unmeasured456proteins represents another important downstream application in spatial proteomics.

      How exactly would the authors go about tackling this given the approach presented in the manuscript?

    2. vely, these analyses identify a reproducible collagen-rich ECM-remodeling449niche that can be robustly transferred across independent cohorts and is preferentially450enriched in recurrent colorectal cancer

      Important application

      1. dding variants, colored by cancer type. b, Image-to-DNAmretrieval across GBM, LGG, HNSC, and BRCA using matched slide and methylation embeddings,shown as Recall@1, Recall@5, and Recall@10. These analyses use pre-trained embeddings beforesupervised downstream task adaptation, providing a direct

      In figure 2(a) why are the Gastric points so disparate/spread out for HistoMethylGigaPath? Same for Lung ADC?

      I think either this is a t-SNE artifact or implies something abt the model.

      Then again for 2(a) one could argue with the four panels => four separate times get embeddings => four independent optimizations with their own initialization, seed, and convergence. I wouldn't expect there to be a shared coordinate system. Even the same embedding re-run at a different seed will redistribute cluster compactness.... If so though the figure doesn't imply much at all / is lowkey useless.

      1. toMethyl overview. a, Cancer cohorts from TCGA [14] used for HistoMethyl pre-training.b, Pre-training inputs, including WSIs divided into tiles and paired DNAm beta values correspondingto each case. c, Pre-training with contrastive alignment. WSI and DNAm embeddings are extractedby their respective encoders, projected by MLPs, and aligned with a contrastive loss. d, Downstreamtask adaptation. HistoMethyl uses task-specific supervised models to adapt aligned slide embeddingsto downstream tasks, including gene mutation, morphology, survival analysis, and DNAm pr

      Methodologically its really smart to split the WSI branch from the DNAm branch and then use a contrastive alignment on the embeddings. What does the joint space look like? Any structure that falls out? Have the authors investigated any patterns in the joint embedding space? I would expect some rather clear boundaries given the inputs.

    1. dding variants, colored by cancer type. b, Image-to-DNAmretrieval across GBM, LGG, HNSC, and BRCA using matched slide and methylation embeddings,shown as Recall@1, Recall@5, and Recall@10. These analyses use pre-trained embeddings beforesupervised downstream task adaptation, providing a direct

      In figure 2(a) why are the Gastric points so disparate/spread out for HistoMethylGigaPath? Same for Lung ADC?

      I think either this is a t-SNE artifact or implies something abt the model.

      Then again for 2(a) one could argue with the four panels => four separate times get embeddings => four independent optimizations with their own initialization, seed, and convergence. I wouldn't expect there to be a shared coordinate system. Even the same embedding re-run at a different seed will redistribute cluster compactness.... If so though the figure doesn't imply much at all / is lowkey useless.

    2. istoMethyl demonstrates that DNA methylation can be used not only as a molecularendpoint, but also as a biologically structured supervision signal for learning clini-cally useful whole-slide representations.

      Important contribution!

    3. toMethyl overview. a, Cancer cohorts from TCGA [14] used for HistoMethyl pre-training.b, Pre-training inputs, including WSIs divided into tiles and paired DNAm beta values correspondingto each case. c, Pre-training with contrastive alignment. WSI and DNAm embeddings are extractedby their respective encoders, projected by MLPs, and aligned with a contrastive loss. d, Downstreamtask adaptation. HistoMethyl uses task-specific supervised models to adapt aligned slide embeddingsto downstream tasks, including gene mutation, morphology, survival analysis, and DNAm pr

      Methodologically its really smart to split the WSI branch from the DNAm branch and then use a contrastive alignment on the embeddings. What does the joint space look like? Any structure that falls out? Have the authors investigated any patterns in the joint embedding space? I would expect some rather clear boundaries given the inputs.

      1. As the taxonomy progresses from broad, coarse-grained catalytic classes (EC .-.-.-)down to highly specific, four-digit reaction profiles (EC x.x.x.x) and targeted families, the challenge of filtering out false positives among closely related sequences scales exponentially. Throughout this entire taxonomic gradient, Boltz2ESI consistently exhibited superior discriminative resolution compared to baseline approaches. Notably, the sustained precision at the deep EC x.x.x.x level indicates that by capturing the adaptive biophysical microenvironment within the active site, our framework effectively untangles tight sub-family specificity that remains hidden to one-dimensional sequencemetrics or rigid-scaffold modeling

      Performance claims here are uninterpretable without the evaluation design. How are negatives sampled at each EC depth?

      1. first localizes the active-sitepocket using Multiple sequence alignment (MSA) guided co-folding, and subsequentlyre-folds the active site and substrate in an MSA-free regime to capture ligand-inducedside-chain and backbone adaptations

      Stage 2 being MSA-free is presumably to avoid the consensus/apo bias that deep MSAs impose on side-chain placement. But does stage 2 condition on stage-1 coordinates? If so, evolutionary information persists as a geometric prior, and the ablation that matters is MSA-free stage 2 without the stage-1 pose, otherwise you can't attribute the induced-fit gains to the MSA-free regime.

      1. From this simulated complex, the framework extracts interaction repre-sentations and a predicted structure to capture local biophysical constraints. Toreconcile this local structure with global biological context, Boltz2ESI fuses these geo-metric descriptors with residue-level evolutionary context derived from ESM3 [28],alongside geometry-aware molecular embeddings [29] and topological fingerprints

      Four feature streams, no ablation. Do the topological fingerprints add anything over the molecular embeddings; both substrate-side, presumably overlapping Id guess? And do the geometric descriptors actually survive/lead to signal gain once ESM3 features are present? If ESM3's structure track is used, those two streams overlap by construction, since both descend from the stage-1 co-folded pose; that's also a leakage concern. Separately: stage 1 is MSA-guided, so the pipeline draws on evolutionary information twice. What does the PLM contribute then?

    1. . From this simulated complex, the framework extracts interaction repre-sentations and a predicted structure to capture local biophysical constraints. Toreconcile this local structure with global biological context, Boltz2ESI fuses these geo-metric descriptors with residue-level evolutionary context derived from ESM3 [28],alongside geometry-aware molecular embeddings [29] and topological fingerprints

      lol, four feature streams, no ablation. Do the topological fingerprints add anything over the molecular embeddings; both substrate-side, presumably overlapping Id guess? And do the geometric descriptors actually survive/lead to signal gain once ESM3 features are present? If ESM3's structure track is used, those two streams overlap by construction, since both descend from the stage-1 co-folded pose; that's also a leakage concern. Separately: stage 1 is MSA-guided, so the pipeline draws on evolutionary information twice. What does the PLM contribute then?

    2. As the taxonomy progresses from broad, coarse-grained catalytic classes (EC .-.-.-)down to highly specific, four-digit reaction profiles (EC x.x.x.x) and targeted families,the challenge of filtering out false positives among closely related sequences scales expo-nentially. Throughout this entire taxonomic gradient, Boltz2ESI consistently exhibitedsuperior discriminative resolution compared to baseline approaches. Notably, the sus-tained precision at the deep EC x.x.x.x level indicates that by capturing the adaptivebiophysical microenvironment within the active site, our framework effectively untan-gles tight sub-family specificity that remains hidden to one-dimensional sequencemetrics or rigid-scaffold modeling

      Performance claims here are uninterpretable without the evaluation design. How are negatives sampled at each EC depth?

    3. Fig. 1 Overview of the Boltz2ESI framework. (a) Given an enzyme sequence and a substrateSMILES string, Boltz-2 co-folding first crops the putative active-site pocket, then the active siteand substrate are re-folded without MSA, and the active-site-level single representation, pair repre-sentation, and distogram of the predicted complex structure are extracted. (b) ESM3 embeddingscapture evolutionary context for enzyme residues, while Uni-Mol2 embeddings and Morgan finger-prints encode substrate molecular features. These priors are fused with the single representation toform a unified representation. (c) The interaction prediction module enriches the pair representa-tion with the fused single representation and the distogram, processes the result through a 4-blockPairformer stack, and outputs an interaction probability via mean pooling and an MLP head. Theresulting scores enable two complementary application modes: ranking candidate substrates for agiven enzyme, and ranking candidate enzymes for a given substrate.

      Beautiful figure!

    4. first localizes the active-sitepocket using Multiple sequence alignment (MSA) guided co-folding, and subsequentlyre-folds the active site and substrate in an MSA-free regime to capture ligand-inducedside-chain and backbone adaptations

      Stage 2 being MSA-free is presumably to avoid the consensus/apo bias that deep MSAs impose on side-chain placement. But does stage 2 condition on stage-1 coordinates? If so, evolutionary information persists as a geometric prior, and the ablation that matters is MSA-free stage 2 without the stage-1 pose, otherwise you can't attribute the induced-fit gains to the MSA-free regime.

  2. Apr 2026
    1. ur free, flow-based, and steered MD simulations not only substantiated previous experimental findings but also revealed previously unrecognized mechanisms of VWF mechanomodulation, including dynamic interactions between the N′AIM and C′AIM regions and the A1 domain.

      Well supported. Good job!

    2. Insights from our flow simulations (Movie S4), which recapitulate the flow-induced unfurling of VWF and the uncoiling of N’AIM and C’AIM to expose the A1 domain (Fig. 3A), revealed that while O-linked glycans enhance steric shielding of A1 from GPIbα, they also modulate the stability of AIM–A1 interactions. Specifically, glycan-induced steric hindrance shortened the lifetimes of both N’AIM–A1 and C’AIM–A1 interactions (Fig. 3B), leading to earlier uncoiling events compared to the unglycosylated system (Fig. 3C). Importantly, the key residues mediating these interactions were conserved regardless of glycosylation status (Fig. 3D), indicating that the observed differences arise primarily from sterics 20.

      With what certainty/confidence? I see the blue/red shadows in 3B but no numerical bound.

    3. However, at sites of vascular injury, elevated shear stress acts as a mechanical cue that triggers VWF to unfurl into an extended, conformation exposing cryptic binding sites for the platelet surface receptor glycoprotein Ibα (GPIbα)5. Remarkably, the spatial organization of VWF is highly context dependent. Within the trans-Golgi network, VWF monomers assemble via head-to-head interactions through the D’D3 domains and tail-to-tail associations via their C-terminal regions, forming higher-order multimers with a characteristic bouquet-like architecture

      Good background tbh!

    1. This dual-masking formulation drives the model to learn robust representations by predicting masked values from complementary perspectives: r

      bro. In Dataset 1, they have 16 panels but the leave-one-panel-out drop is <8%. That's the better evidence for robustness.

      Then you claim pretraining drives "robust representations despite marker inconsistency" based on the KO task, where Dataset 2 has a completely consistent panel.

      Those two claims aren't using the same evidence base and shouldn't be merged into one conclusion.

    2. Notably, large-scale pretraining yielded considerable performance gains in small-data settings, attributable to robust cellular representations that recover biological signals despite marker inconsistency.

      If you're concluding this based on per-class AUC and markers are inconsistent isn't this claim sketchy at best?

    3. Dataset 1: Longitudinal mouse immunophenotype datasetAs part of a long-running mutagenesis project to investigate novel genetic causes of immune dysfunction [18], flow cytometry phenotypes for over forty thousand C57BL/6 mice were obtained at the Australian Phenomics Facility between 1995 and 2015. This data is comprised of predominantly eight-colour experiments with varying marker/antibody/fluorophore combinations, yet most samples include a backbone of six common markers (IgM, IgD, B220, CD44, CD4, CD3) (Supplementary Table 10 and 8).In the present analysis, we have chosen a subset of 14,014 flow cytometry samples (6,978 female, 7,036 male) with a consistent gender metadata label and mostly pan-leukocyte marker panels. Sexual dimorphism rarely produces landmark cell populations readily detectable by manual analysis of flow cytometry data. However, this has proven a tractable problem with application of neural networks [19], with discriminative signals usually subtle and dispersed across multiple cell populations.Dataset 2: Knockout Mouse Project immunophenotype datasetThe Knockout Mouse Project (KOMP) [20] generated mouse strains harbouring gene knockouts for the majority of genes in the mouse genome, accompanied by phenotype data including flow cytometry information for a subset of mutant mouse lines. For our purposes, we focus on a subset of samples subjected to flow cytometry assay of a T cell immunophenotyping panel [21] (Supplementary Table 3). Despite containing nearly 7000 samples, this dataset poses a classic lack-of-data problem, as each knockout (KO) is represented by only 10 to 20 samples. As most knockouts in this dataset were found to lack discernible cellular phenotypes [21], we selected just 5 knock-out lines with clear mutant phenotypes characterised by the original study. This yields 72 samples (Supplementary Table 9) for a 5-class KO classification task.

      So you have 14k samples for dataset 1, a slight imbalance in male/female, but only 72 samples for dataset 2 because of selecting only 5 knockout lines? Also "most knockouts in this dataset were found to lack discernible cellular phenotypes"? Is that not concerning if you want to claim general ability/can build on for flow cytometry?

      Pre-training distribution has a significant impact on downstream utility.

    4. We evaluated the impact of cross-dataset pretraining on the model generalisation scenario using two configurations. The first model, the D1 encoder (Experiment A and B), was trained exclusively on Dataset 1. The second, the generic encoder (Experiments C), was pretrained on combined training data from Datasets 1 and 2 before downstream training on Dataset 1 only. Results in Fig. 2b (1) demonstrate that including even a small fraction of Dataset 2 in the pretraining phase significantly improved downstream generalisation to Dataset 2 testing samples.

      What exactly do the pre-training distributions look like? Whats the exact mix? Is dataset 1 sufficiently different from dataset 2, specifically as it relates to sample quality and number of samples?

    5. In this regard, GPCT can be interpreted through the attention mechanism used by the decoder: during inference, each attention head in the multi-head attention layer assigns a weight to every cell, representing its relative contribution to the decision-making process. These weights serve as a quantitative measure of per-cell “importance”, and while they are typically averaged across heads per layer for visualisation, each layer may capture distinct patterns that reflect the model’s internal processing steps.

      Interesting concept to make them cell level. Why not clusters of cells?

    1. Within-method exact agreement on normalized relevance labels was modest (Figure 3; Table 2). The best agreement was between Claude Code runs 2 and 3 (54/73 orthogroups; 0.740), while the lowest was between Claude Code runs 1 and 3 (25/73; 0.342). Mean within-method agreement was in the same range for all three configurations (0.516–0.562), so no configuration was dramatically more reproducible than the others at the tier-label level. These results argue against relying on a single stochastic agent run for final biological claims, even when the input files and prompt are identical.

      Is within-method exact agreement really the best metric? Recommending to not run against a single stochastic agent is fine but what is the delta? Running many costs more for what benefit?

    2. lthough coverage was complete, calibration differed strongly across runs (Figure 2; Table 1). Claude App run 2 was highly conservative, assigning 67 of 73 orthogroups to a low or background tier and only one high call. Claude App runs 1 and 3 were less conservative, with 11 and 8 high calls, respectively. Claude Code with scientific skills produced fewer high calls overall (1, 3, and 2), but shifted substantially between low and watchlist labels across runs. Codex App with scientific skills showed the widest high-call range, from no high calls in run 2 to 12 high calls in run 3.

      How does temperature/nucleus sampling/effort affect these results? Did you control for potential variation in these parameters?

    3. Here we use a controlled, repeated-run comparison to evaluate three agent configurations as they were used on the same orthogroup annotation prompt. The goal is not to rank proprietary foundation models in general. Instead, we ask a practical question relevant to bioinformatics groups: when agents are asked to retrieve, integrate, and interpret a large set of complex protein annotations, where do they help, where do they fail, how consistent are repeated runs, and how should their outputs be merged into a defensible final annotation table?

      Great experiment! I wonder what metrics are reported and how representative/relevant those metrics are given real life tasks.

    4. ombining these evidence streams is routine, but the final biological interpretation is still difficult because many protein families are multidomain, repetitive, lineage-specific, or only indirectly connected to the process of interest.

      Absolutely!! A challenge of great importance.

  3. Mar 2026
    1. To reduce class imbalance, we exclude these longer sequences from the training dataset.

      Downsampling/excluding a minority class is an interesting decision. Why not include some of them, use some flavor of stratified sampling/curriculum learning/train a specialized subnetwork/set of heads on the larger sequences with a reasonable split? How does the model generalize/perform on larger sequences?

    2. In our earlier work, we introduced DisPredict3.0, the most recent iteration of the DisPredict series, which integrates evolutionary representations derived from protein language models to improve the prediction of intrinsically disordered regions (IDRs) [5]. This approach achieved the top ranking on the Disorder NOX dataset in CAID2. Building on this foundation, we now present ESMDisPred, a structure-aware disordered protein predictor that incorporates embeddings from the Evolutionary Scale Modeling-2 (ESM2) language model [3]. ESM2 is considered the SOTA language model and has demonstrated exemplary performance in protein structure prediction (ESMFold)

      This is interesting. Evolutionary context can be really informative?

    1. First, generating a textual analysis of a binding site based on a protein sequence.Second, predicting a plausible binding site conformation given a specific ligand.Third, synthesizing a functional description by integrating the protein sequence, the predicted conformation, and the ligand information.

      Nice breakdown tbh!

    2. Output: A natural language answer describing the protein’s function, activity, or binding mode under the specified ligand conditions.

      To what extent is this accurate / aligned with biological reality? Does generating natural language answers introduce a source of error/confounding? What happens as answers become shorter vs longer vs less/more complex?

    3. A central challenge is learning aligned and effective representations across different data types, such as learning effective binary descriptors that can maintain group fairness [27].

      To what extent has this been solved by better molecular representations? Proteins and ligands are still molecules, and wouldn't atom-level representations ensure consistency across these data types? Boltz/BoltzGen does leverage atom-level information...

    4. SE(3)-invariant encoder combined with a temporal-aware VQ-VAE style quantization module. This allows us to convert diverse binding pocket conformations (e.g., apo, holo, or intermediate states) into discrete tokens, effectively capturing their dynamic variations. Furthermore, we integrate standard SMILES string tokenization for small molecules, alongside specialized amino acid tokens and the native Llama3 text tokenizer, expanding the LLM’s vocabulary to encompass these crucial biological entities.

      Why the Llama3 tokenizer among all other choices? Seems odd methodologically? Why not something designed for this kind of purpose? https://arxiv.org/html/2409.15370v1

    1. Discrete diffusion objective: we experiment with two different masking techniques, the first is the standard discrete diffusion objective where the masking fraction is sampled from a uniform distribution over (0, 1), in the second we sample the masking fraction 80% of the time from a β(3, 9) distribution and 20% of the time from a uniform distribution over (0, 1). This approach, adapted from (Hayes et al., 2024), aims to balance representation and generation capabilities. It allows the model to observe masking fractions across (0, 1), with an average <img class="highwire-embed" alt="Embedded Image" src="https://www.biorxiv.org/sites/default/files/highwire/biorxiv/early/2025/04/08/2025.04.02.646805/embed/inline-graphic-1.gif"/>. Both these objectives improve the effectiveness for iterative denoising during sequence generation with respect to standard MLM

      Other objectives could have been chosen beyond ease of implementation, why in particular this objective? Why not a hybrid objective? what abt retrieved neighbor training?

    2. Here, we propose a method to make them homology-aware. We introduce RAG-ESM, a retrieval-augmented framework that allows to condition pretrained ESM2 protein language models on homologous sequences, using a minimal number of additional cross-attention parameters and minimal computational cost

      This is an interesting idea. I wonder what the scaling looks like and what the efficacy of the augmentation is with respect to context window size and the quality of retrieval.

    3. We introduce RAG-ESM, a retrieval-augmented framework that allows to condition pretrained ESM2 protein language models on homologous sequences, using a minimal number of additional cross-attention parameters and minimal computational cost.

      This is an interesting idea. I wonder what the scaling looks like and what the efficacy of the augmentation is with respect to context window size and the quality of retrieval.