Supplementary repository for paper Pareto optimization of masked superstrings improves compression of pan-genome k-mer sets.
This repository contains the implementation of Pareto optimization of masked superstrings for k-mer sets and computation of lower bound for the number of runs of ones in the mask (lower bound for the number of matchtigs).
It also contains Snakemake pipelines reproducing the results from the paper, and uses conda to manage software versions and most of the dependencies.
We plan to also add the exact scripts creating the figures in the paper; however, this is currently still a TODO.
In case of any questions, feel free to ask by an email.
To install dependencies from conda and build localy provided software, use:
conda env create --name ms-pareto-optimization
conda activate ms-pareto-optimization
./prepare.shNext time, you only need to activate the conda environment with:
conda activate ms-pareto-optimizationTo run an experiment, move to the corresponding directory (ex1-..., ex2-..., ...) and run:
rm -rf results/<dataset for which you want to recompute the results>
snakemake -j <number_of_threads_to_use> <any_optional_parameters_of_snakemake>Experiments produce .tsv files with results (and, in the future, plots).
If you used git clone to obtain the repository, you can compare them to the original results using git diff results.
Example snakemake configuration: snakemake -j 10 --rerun-incomplete --resources mem_mb=26000
To tweak an experimental setup, modify the Snakefile in the corresponding directory.
Constants defined on top of Snakefiles define which datasets,
values of k, run penalties, and computation or compression methods are used.
Locally installed programs and snakefiles are stored in tools directory.
Other directory names can be modified in tools/project_structure.smk if needed.
Default names are:
- Datasets are downloaded into
datadirectory. - Computed superstring representations are stored in
computeddirectory (it is recommended to turn off file search indexing for this directory). - Results are stored in respective numbered directories (
ex1-...,ex2-..., ...) in aresultssubdirectories.
If you want to use the pipeline with custom datasets, you can modify datasets.txt.
Add the dataset name and url, the pipeline downloads datasets automatically
and can handle xz-compressed and uncompressed FASTA files.
In case you need to support other compression formats,
you can modify the download_data rule in tools/download.smk,
but probably the simplest way is to just download and extract the file by yourself,
in which case the downloading part of the pipeline is not used at all.
Details about the implementation are linked in separate README.
The work was implemented inside a fork of KmerCamel🐫 and then copied over to another fork, which resulted in a weird commit history for a while. Now it is settled like this:
- the first commit contains the history of original KmerCamel🐫, which has been forked.
- the relevant changes (Pareto optimization implementation) are in the second commit made by @Jajopi (d86ccf1a6b8007a6cb8680bb46f47a91ef058beb).
The main concept (Pareto optimization) was internally called Joint optimization for a long time. All occurences in the files should be up to date; however, you may encounter the old name in the history.