MalViT turns the multi-head attention of a byte-level transformer into a set
of grayscale images (one per attention head) and classifies an Android APK from
those images. A file is described by L × H complementary views instead of a
single static texture.
raw APK ─► byte tokenize ─► chunk (1024) ─► [frozen transformer] ─► per-head
attention ─► vocab-level 256×256 aggregation ─► RMS over chunks ─►
L·H grayscale images ─► [CNN] ─► benign / malware
| Component | File | Role |
|---|---|---|
| Central config | src/config.py |
Single source of truth; derived values (vocab_size, num_channels), path resolution. |
| Tokenizer | src/data/tokenizer.py |
Bytes → tokens (.npy). |
| Split | src/data/split.py |
File-level train/val/test manifest shared by both stages (prevents leakage). |
| Dataset | src/data/dataset.py |
Expands the manifest into chunk arrays per split. |
| Transformer | src/model/transformer.py |
Encoder-only BERT; trained once, then frozen and used as an attention extractor. |
| Image generator | src/model/image_generator.py |
Vocabulary aggregation (PAD excluded), RMS fusion, contrast, grayscale. |
| CNN classifier | src/model/classifier.py |
Multi-channel (256,256,L·H) CNN. |
| Metrics | src/evaluation/metrics.py |
Accuracy, Precision, Recall, F1, FPR, AUC-ROC, plots. |
pip install -r requirements.txtRun modules with -m from the repository root. Put APKs under
data/raw/benign/ and data/raw/malware/ first.
python -m src.data.tokenizer # 1. bytes -> data/processed/*.npy
python -m src.data.split # 2. build the file-level split manifest
python -m src.training.train_transformer # 3. train + freeze the transformer
python -m src.training.generate_images # 4. write L*H images per file, per split
python -m src.training.train_classifier # 5. train CNN, report metrics on the test split
python -m src.inference.predict path/to/app.apk # single-file inferenceThe split is computed once, at the file level, and persisted to
data/split_manifest.json. Both the transformer (build_transformer_dataset)
and the classifier (physically separated data/images/{train,val,test}/) read
that same manifest, so no APK — and none of its chunks — can appear in more than
one partition or leak between the two training stages. Set
split.strategy: temporal in configs/config.yaml (with a stem,year
data/metadata.csv) to train on 2019–2021 and evaluate on 2022–2023 instead.
The TensorFlow-free core (tokenizing, chunking, PAD-excluded aggregation, RMS, file-level split) is unit-tested and runs without a GPU:
python -m pytest tests/ -v