Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Clustering

Clustering creates two-dimensional embeddings (scatter plots) which are designed to cluster genetically related samples next to one another. This uses the stochastic cluster embeddding / mandrake algorithm which we have ported to rust. This is an exploratory tool for finding genetically related clusters in microbial population genome data, and for creating visualisations.

General considerations

  • Clustering can be run from:
    • Sequence alignments (FASTA / gz).
    • Accessory presence/absence matrices (Rtab / tsv / gz).
    • Sketch databases (from sketchlib.rust). Databases must be >v0.4

The first two methods use pairsnp to calculate pairwise distances. Sketches you can use either a Jaccard distance at a single k value, or core distance (when >3 k-mers are available in the dataabase).

Note that for large numbers of samples:

  • You may need to increase the maximum updates to achieve convergence.
  • The distance step may become slow. We don’t support caching distances, so for many reruns you may prefer the CLI app.

Sample labels:

  • Can be provided, so they are coloured in based on these categories in the final embedding.
  • Can be generated by running HDBSCAN on the final embedding, if the box is ticked.

Parameters

The values in brackets are the default ones:

  • Sparsification [k-nearest neighbours]: knn retains only the nearest k distances [default: 15] to each sample; threshold retains all distances below this threshold.
  • Preplexity [30]: Perplexity roughly balances global versus local attention, see this explainer.
  • Maximum updates [10^6]: The number of stochastic steps applied. Increase for sample numbers >5000, decrease for <1000.
  • Repulsion samples [5]: The number of sample repulsions to apply for each attraction.
  • Learning rate [false]: The rate to decrease attraction/replusion forces over the optimisation.
  • Use initial exaggeration [false]: Make selected samples 4x more attracted to each other in the first 10% of updates.

Example

The following data can be used to try out gene calling: