Functional Profiling of Thousands of Sequence-DiverseProtease Homologs with GROQ-seq
The MSA match columns used for the tree and identity statistics were selected with a <50% gap threshold. This seems largely driven by the ~1,450 sequences at >=80% identity to TEV S219V, as opposed to the divergent tail. Is there a concern that this skews the identities and N_eff values in Table 1 and Supplementary Table 1? Also, Supplementary Table 1 lists 11,729 sequences at >=10% identity compared to the library having 11,722 unique sequences. Where do the extra seven come from?