Under active development — interfaces may evolve
CLI toolkit to (1) scan PWMs genome-wide, (2) compute accessibility-derived site scores s_l, (3) annotate with TSS/conservation/TPM, and (4) fit and evaluate a Gaussian EM with a logistic prior.
Full user and API documentation: https://yangli04.github.io/RAD-seq_EM_model/
- Tested on Ubuntu 24.04 LTS (kernel 6.17, x86_64).
- Expected to work on any recent Linux distribution and macOS (12+). Not tested on Windows; use WSL2.
- Python ≥ 3.12.
polars1.41,numpy2.4,scipy1.17,pandas3.0,pyarrow24pyranges11.3.8metagene0.0.16,scikit-learn1.8,numba0.65,pybigwig0.3.25matplotlib3.10,seaborn0.13typer0.26,tqdm4.67
A minimum-version manifest is in pyproject.toml.
None. CPU-only; no GPU required. ~4 GB RAM is sufficient for the demo; full genome-wide runs scale with the number of motif hits and benefit from ≥16 GB RAM.
Create a fresh Python 3.12 environment (micromamba shown; conda/venv work the same):
micromamba create -y -n accmix python=3.12 -c conda-forge
micromamba activate accmixInstall from source (development install):
git clone https://github.com/yangli04/RAD-seq_EM_model.git
cd RAD-seq_EM_model
pip install -e .or from PyPI:
pip install accmixVerify the CLI:
accmix --helpOn a normal desktop (Ubuntu 24.04, 8-core CPU, warm pip cache, broadband):
- Environment creation: ~10 s
pip install -e .(orpip install accmix): ~30 s- Total: ~40 s (first-ever install with a cold pip cache: ~2–3 min)
All files referenced under data/... are bundled in one archive. Download from Box:
https://uchicago.box.com/s/tsmy_wbgi_mvl4_7w9n_ynq6_6hur_3c8d_nirvNote for humans: Box link is mildly obfuscated to deter automated scrapers. Before using, remove all underscores
_from the path segment after/s/. The resulting URL should contain only letters and digits in the share token.
Unpack at the repo root so the structure matches:
RAD-seq_EM_model/
├── data/
│ ├── fasta/test.fa
│ ├── pwms/M00124_example.txt
│ ├── clipseq/ELAVL1_HeLa.bed
│ ├── AS/ANC1C.hisat3n_table.bed6
│ ├── AS/ANC1xC.hisat3n_table.bed6
│ └── evaluation/...
└── ...
Four-step pipeline using the bundled example PWM M00124:
# (1) Scan PWM and GC across the test genome
accmix scan \
-f data/fasta/test.fa \
-p data/pwms/M00124_example.txt \
-o results/M00124_example
# (2) Compute accessibility-derived s_l from AS native/fixed tables
accmix annotate-acc \
-n data/AS/ANC1C.hisat3n_table.bed6 \
-f data/AS/ANC1xC.hisat3n_table.bed6 \
-t results/M00124_example_topA.tsv.gz \
-o results/M00124_example_sl.parquet
# (3) Add TSS / conservation / TPM annotations
accmix annotate-tss \
-i results/M00124_example_sl.parquet \
-o results/M00124_example_annotated.parquet \
-r data/evaluation/RNAseq_HeLa_TPM.parquet \
-c data/evaluation/phastCons100way.bed.gz \
-p data/evaluation/phastCons100way.parquet \
-y data/evaluation/phyloP100way.parquet \
-R data/fasta/test.fa
# (4) Fit the EM model
accmix model \
-i results/M00124_example_annotated.parquet \
-o results/RBP_Motif \
-r ExampleRBP \
-m M00124After step (1):
results/M00124_example_topA.tsv.gz— top-strand PWM hits with inner/outer scoresresults/M00124_example_botB.tsv.gz— bottom-strand PWM hits
After step (2):
results/M00124_example_sl.parquet— sites with thes_laccessibility score column
After step (3):
results/M00124_example_annotated.parquet— sites with addedtss_dist,phastCons,phyloP,TPMcolumns
After step (4):
results/RBP_Motif.<id>.model.parquet— original sites annotated withprior_pandposterior_rresults/RBP_Motif.<id>.model.json— fitted Gaussian + logistic-prior parameters
To run on your own data, replace the demo inputs with files in the same formats:
| Stage | Input | Format |
|---|---|---|
scan |
Genome FASTA, PWM | .fa (indexed or plain); PWM in TRANSFAC-style matrix |
annotate-acc |
AS native / fixed tables, PWM hits | .bed6 from hisat-3n table conversion; PWM hits from step 1 |
annotate-tss |
s_l parquet + annotation tracks | Parquet; phastCons/phyloP as parquet or .bed.gz; TPM as parquet |
model |
Annotated parquet | Output of annotate-tss |
evaluate |
Model parquet + CLIP/PIP-seq peaks | .bed peaks |
Each subcommand documents its full option list:
accmix scan --help
accmix annotate-acc --help
accmix annotate-tss --help
accmix model --help
accmix evaluate --helpKey tunables:
annotate-acc -M / -N: inner / outer flank size (defaults: 50 / 500). Outer flank length isN − M.model -r / -m: RBP and motif identifiers used to name outputs.
See the full CLI reference at https://yangli04.github.io/RAD-seq_EM_model/cli/ for every flag and default.
To reproduce the figures and tables from the manuscript:
- Install per Section 2.
- Download the full Box archive (Section 3, Demo) — this includes all example PWMs, AS tables, conservation tracks, and CLIP/PIP-seq evaluation data.
- Run the Snakemake pipeline driver:
snakemake -s Snakefile --cores 8
- Generated tables and figures will appear under
results/(or follow the symlinked path if you redirectresults/to a larger disk).
A step-by-step reproducibility walkthrough — including expected intermediate file sizes and the exact PWM/cell-line manifest — is maintained at:
https://yangli04.github.io/RAD-seq_EM_model/reproducibility/
The full user guide, API reference, and tutorials live at:
https://yangli04.github.io/RAD-seq_EM_model/
If you use this package, please cite:
Li Y., accmix: Accessibility Mixture Model, GitHub repository, https://github.com/yangli04/RAD-seq_EM_model