Clustering
Clustering creates two-dimensional embeddings (scatter plots) which are designed to cluster genetically related samples next to one another. This uses the stochastic cluster embeddding / mandrake algorithm which we have ported to rust. This is an exploratory tool for finding genetically related clusters in microbial population genome data, and for creating visualisations.
General considerations
- Clustering can be run from:
- Sequence alignments (FASTA / gz).
- Accessory presence/absence matrices (Rtab / tsv / gz).
- Sketch databases (from sketchlib.rust). Databases must be >v0.4
The first two methods use pairsnp to calculate pairwise distances. Sketches you can use either a Jaccard distance at a single k value, or core distance (when >3 k-mers are available in the dataabase).
Note that for large numbers of samples:
- You may need to increase the maximum updates to achieve convergence.
- The distance step may become slow. We don’t support caching distances, so for many reruns you may prefer the CLI app.
Sample labels:
- Can be provided, so they are coloured in based on these categories in the final embedding.
- Can be generated by running HDBSCAN on the final embedding, if the box is ticked.
Parameters
The values in brackets are the default ones:
- Sparsification [k-nearest neighbours]: knn retains only the nearest k distances [default: 15] to each sample; threshold retains all distances below this threshold.
- Preplexity [30]: Perplexity roughly balances global versus local attention, see this explainer.
- Maximum updates [10^6]: The number of stochastic steps applied. Increase for sample numbers >5000, decrease for <1000.
- Repulsion samples [5]: The number of sample repulsions to apply for each attraction.
- Learning rate [false]: The rate to decrease attraction/replusion forces over the optimisation.
- Use initial exaggeration [false]: Make selected samples 4x more attracted to each other in the first 10% of updates.
Example
The following data can be used to try out gene calling:
- Alignment: 5000 HIV genomes.
- Accessory matrix: 1973 S. pneumoniae.
- Sketches: 616 S. pneumoniae .skd and .skm metadata.