At 30–40 patients: random forest, gradient boosting, and frozen pretrained audio embeddings with a linear probe become viable as parallel arms.
these can capture non-linear interactions between features that logistic regression would miss. For example, maybe the relationship between damping and resonance shift only matters when Q factor is below a certain threshold. Trees find that automatically. The reason we can't do this at 15–20 is that these models have more free parameters and will fit noise in ways that are harder to detect and interpret.