A snakemake pipeline to analyse experimental evolution data.
ExpEvoAnalyzer uses shovill for assembly, bakta for genome annotation, SKA2/bwa-mem + samtools/bcftools for variant calling, and SnpEff for variant annotation.
ExpEvoAnalyzer can either be run on assemblies or paired-end reads, and works with gzipped files.
ExpEvoAnalyzer filters variants caused by assembly errors by mapping reference reads back to the original assembly, and filtering any shared variants.
- python>=3.9
- bakta
- ska2
- spades
- pandas
- biopython
- snakemake>=9.6
- snpeff
- shovill
- pyvcf
- gzip
- bwa
- samtools
- bcftools
Install the required packages using conda/mamba:
git clone https://github.com/samhorsfield96/ExpEvoAnalyzer.git
cd ExpEvoAnalyzer
mamba env create -f environment.yml
mamba activate expevoanalyzer
ExpEvoAnalyzer expects all reads to be trimmed (adapters removed) if provided. ExpEvoAnalyzer can also be run with assemblies.
Update config.yaml to specify workflow and directory paths.
output_dir: path to output directory. Does not need to exist prior to running.input_dir: path to directory containing paired-end FASTQ files (gzipped or uncompressed).reference_reads_R1: path to FASTQ of read 1 for the reference isolate. Can be provided in absence ofreference_assembly. If provided with a differentreference_assembly, will be treated as and "ancestral" sequence, whereby any variants shared with it in the populaiton will be ignored.reference_reads_R2: path to FASTQ of read 2 for the reference isolate. Can be provided in absence ofreference_assembly. If provided with a differentreference_assembly, will be treated as and "ancestral" sequence, whereby any variants shared with it in the populaiton will be ignored.reference_assembly: path to FASTA of reference assembly (if already generated). Can be provided in absence ofreference_reads_R1andreference_reads_R1.reference_assembly_gbk: (OPTIONAL) path to Genbank file (gbff) of previously annotated genome. Must matchreference_assemblyif used.alignment_method: must be one ofska(aligning to closely-related reference) orbwa(aligning to distantly-related reference)shovill_params: additional parameters for shovill e.g.--trimksize: integer, k-mer size used for SKA2snpEFF_config: path to SnpEff config file, usuallyscripts/snpEff.configbakta_db: path to bakta database. This needs to be downloaded and updated prior to running (see here)bakta_extra_settings: enables addition of extra bakta parameters e.g. adding your own trusted protein annotations using--protein file.fasta. This can be generated from a gbff using the scriptextract_fasta_from_gbff.py
Run snakemake:
snakemake --cores <cores>
All workflows output the following directories/files:
shovill: assembly files of reference strain.bakta: bakta annotations of reference strain.reads_list: input files for SKA2.vcfs: VCF files generated by SKA2/bwa, non-reference read files mapped to reference assembly.snpeff: annotated VCFs generated by SnpEff.annotated_variants.csv: presence/absence matrix of variants in non-reference files, annotated by SnpEfflogs: log files.
If running ska, you'll get an additional directory:
ska_build: SKA2 sketches of non-reference read files.
If running bwa, you'll get an alternative additional directory
bwa_mem: alignment files generated by bwa mem.