Skip to content
gaussssssPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

MalViT — Multi-Head Attention Visualization for Malware Detection

MalViT turns the multi-head attention of a byte-level transformer into a set of grayscale images (one per attention head) and classifies an Android APK from those images. A file is described by L × H complementary views instead of a single static texture.

raw APK ─► byte tokenize ─► chunk (1024) ─► [frozen transformer] ─► per-head
        attention ─► vocab-level 256×256 aggregation ─► RMS over chunks ─►
        L·H grayscale images ─► [CNN] ─► benign / malware

Architecture

Component File Role
Central config src/config.py Single source of truth; derived values (vocab_size, num_channels), path resolution.
Tokenizer src/data/tokenizer.py Bytes → tokens (.npy).
Split src/data/split.py File-level train/val/test manifest shared by both stages (prevents leakage).
Dataset src/data/dataset.py Expands the manifest into chunk arrays per split.
Transformer src/model/transformer.py Encoder-only BERT; trained once, then frozen and used as an attention extractor.
Image generator src/model/image_generator.py Vocabulary aggregation (PAD excluded), RMS fusion, contrast, grayscale.
CNN classifier src/model/classifier.py Multi-channel (256,256,L·H) CNN.
Metrics src/evaluation/metrics.py Accuracy, Precision, Recall, F1, FPR, AUC-ROC, plots.

Install

pip install -r requirements.txt

Run the pipeline

Run modules with -m from the repository root. Put APKs under data/raw/benign/ and data/raw/malware/ first.

python -m src.data.tokenizer            # 1. bytes -> data/processed/*.npy
python -m src.data.split                # 2. build the file-level split manifest
python -m src.training.train_transformer  # 3. train + freeze the transformer
python -m src.training.generate_images    # 4. write L*H images per file, per split
python -m src.training.train_classifier   # 5. train CNN, report metrics on the test split
python -m src.inference.predict path/to/app.apk   # single-file inference

Leakage control

The split is computed once, at the file level, and persisted to data/split_manifest.json. Both the transformer (build_transformer_dataset) and the classifier (physically separated data/images/{train,val,test}/) read that same manifest, so no APK — and none of its chunks — can appear in more than one partition or leak between the two training stages. Set split.strategy: temporal in configs/config.yaml (with a stem,year data/metadata.csv) to train on 2019–2021 and evaluate on 2022–2023 instead.

Tests

The TensorFlow-free core (tokenizing, chunking, PAD-excluded aggregation, RMS, file-level split) is unit-tested and runs without a GPU:

python -m pytest tests/ -v

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages