Skip to content

Latest commit

 

History

86 Commits

Folders and files

Repository files navigation

Robust Feature-Locking Technique for Language Models

arXiv checkpoints dataset

Locket (ACL '26) is a feature-locking technique (FLoTE) that enables feature-level access control for LLMs, supporting applications such as: DLC-like freemium monetization scheme for LLM services, robust content/age restrictions, flexible regulatory compliance, etc.

@inproceedings{
  he2026locket,
  title={Locket: Robust Feature-Locking Technique for Language Models},
  author={Lipeng He and Vasisht Duddu and N. Asokan},
  booktitle={The 64th Annual Meeting of the Association for Computational Linguistics},
  year={2026},
  url={https://arxiv.org/abs/2510.12117}
}

Pretrained Demo Adapters

The following four feature-locking adapters, each locking one feature of DeepSeek-Math-7B, are available on Hugging Face:


Environment Setup

Experiments were run on Lambda with 8 × NVIDIA A100 40GB GPUs.


1. Conda Environment

conda create -n locket python=3.12
conda activate locket

2. Dependencies

Install in the following order to resolve conflicts:

conda install -c pytorch -c nvidia faiss-gpu=1.12.0

pip install datasets==4.0.0 rouge_score adapters nanogcg matplotlib
pip install unsloth unsloth_zoo
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu126
pip install -U xformers==0.0.29.post3 --index-url https://download.pytorch.org/whl/cu126
pip install https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.4.post1/flash_attn-2.7.4.post1+cu12torch2.6cxx11abiTRUE-cp312-cp312-linux_x86_64.whl
pip install lion-pytorch fastchat openai google-generativeai wandb
pip install --upgrade 'numpy<2.0' 'pandas>=2.2'
pip install transformers==4.51.3 trl==0.18.2 torchao==0.13.0 peft==0.17.1

3. Project Setup

pip install -e .

Download the datasets (math, sql, samsum, refusal) into data/:

hf download ttttonyhe/locket-data --repo-type dataset --local-dir data

Login to HuggingFace and Weights & Biases:

hf auth login
wandb login

Download the Llama-3-8B chat template used by AutoDAN-Turbo's judge:

hf download meta-llama/Meta-Llama-3-8B-Instruct \
  --local-dir ./locket/robustness/AutoDAN_Turbo/llm/chat_templates/model_ckpt/meta-llama_Meta-Llama-3-8B-Instruct

Running Experiments

Long-running jobs should be run in a screen or tmux session with logging enabled:

# screen
screen -S <name> -L -Logfile /path/to/<name>.log
# tmux
tmux new -s <name>                                 # start a named session
tmux pipe-pane -o -t <name> 'cat >> /path/to/<name>.log'   # log the session to a file

Step 1 — Train Feature-Locking Adapters

Trains one LoRA adapter per feature via LAT (§4). Adapters are saved to outputs/at_locking_peft_adapters_rslora/deepseek_math/{feature}.

make train_at_locking

Configure LAT_DATASETS and ADAPTER_NAMES in locket/training/lock_at.py to select which features to train.


Step 2 — Evaluate Effectiveness and Utility (R1 & R2)

Single-feature and multi-feature scalability.

make eval_effect

Configure TARGET_MODELS in locket/effectiveness/main.py to select configurations. Results are logged to stdout and saved to logs/.


Step 3 — Evaluate Robustness (R3)

Attack success rates for Many-shot, GCG, TAP, AutoDAN-Turbo.

make eval_robust

Configure TARGET_MODELS, JAILBREAK_METHODS, and JAILBREAK_FEATURES in locket/robustness/main.py. Results are saved as JSON to logs/.


Key Hyperparameters

Parameter Value Description
LoRA rank 64 Adapter rank (RSLoRA)
PGD steps 16 LAT inner loop iterations
PGD layers embedding, 6, 14, 22, 29 Layers attacked during LAT
Training steps 100 Total LAT training steps
τ (single) 0.5–0.95 Per-feature spectral cap (see locket/utils/model.py)
τ (multi) 0.6–0.9 Multi-feature spectral cap (see locket/utils/model.py)

See Appendix E of the paper for full details.

About

A feature-locking technique (FLoTe) that enables robust feature-level access control for LLMs

Topics

Resources

Stars

160 stars

Watchers

0 watching

Forks

Contributors

Languages