Skip to content

Repository files navigation

nnsight

Interpret and manipulate the internals of deep learning models

Documentation | GitHub | Discord | Forum | Twitter | Paper

Ask DeepWiki


About

nnsight lets you get inside a model's forward pass. Open a with model.trace(...) block and write ordinary Python against any internal value — a layer's output, an attention pattern, a gradient — as if you already had it: read it, edit it, save it. You don't register hooks or refactor the model; you write the intervention in the order it happens, and nnsight runs it interleaved with the real forward pass.

The same trace runs on a model on your laptop or, with remote=True, on a model far too large for it via the NDIF infrastructure. nnsight works with any PyTorch model and ships wrappers for HuggingFace, diffusers, and vLLM.

📖 For how it works under the hood — tracing, interleaving, the envoy tree — read NNsight.md. Task recipes live under docs/.

Installation

pip install nnsight

Quick start

from nnsight import TransformersModel

model = TransformersModel("openai-community/gpt2", dispatch=True)

with model.trace("The Eiffel Tower is in the city of"):
    # edit a layer's output in place — the model computes on the edited value
    model.transformer.h[0].output[:] = 0

    # read a hidden state further down (a [batch, seq, hidden] tensor)
    hidden = model.transformer.h[6].output.save()

    # keep the final logits
    logits = model.output.logits.save()

print(hidden.shape)            # torch.Size([1, 10, 768])
print(logits.argmax(-1))       # next-token predictions

Inside the block you're not running the model — you're describing what to do when it runs. Reading .output gives you the real tensor once the model reaches that module; assigning to it splices your value in. Mark anything you want after the block with .save() (or nnsight.save(x)).

A GPT-2 block's .output is a plain tensor; some modules (like attention) return a tuple, so you'd index .output[0]. print(model) or print(module.source) shows the shape.

What you can do

Generate — generate returns token ids on tracer.result:

with model.generate("The Eiffel Tower is in", max_new_tokens=3) as tracer:
    ids = tracer.result.save()
print(model.tokenizer.decode(ids[0]))     # "The Eiffel Tower is in the middle of"

Reach into each generation step — save a container, append raw values (use a bounded range so code after the loop still runs):

import nnsight

with model.generate("Hello", max_new_tokens=5) as tracer:
    tokens = nnsight.save([])
    for step in tracer.iter[:5]:
        tokens.append(model.output.logits[0, -1].argmax(-1))

Batch several prompts in one pass — each invoke block sees only its own rows:

with model.trace() as tracer:
    with tracer.invoke("The Eiffel Tower is in"):
        eiffel = model.transformer.h[-1].output[:, -1].save()
    with tracer.invoke("The Great Wall is in"):
        wall = model.transformer.h[-1].output[:, -1].save()

Take gradients with respect to an internal value:

with model.trace("The Eiffel Tower is in the city of"):
    hidden = model.transformer.h[-1].output
    with model.output.logits.sum().backward():
        grad = hidden.grad.save()

Apply modules out of order (a logit lens — decode a middle layer through the head):

with model.trace("The Eiffel Tower is in the city of"):
    hidden = model.transformer.h[-1].output
    token = model.lm_head(model.transformer.ln_f(hidden))[0, -1].argmax(-1).save()
print(model.tokenizer.decode(token))      # " Paris"

Run it remotely on NDIF — the same trace, on a model you can't host:

from nnsight import CONFIG
CONFIG.set_default_api_key("YOUR_NDIF_KEY")

model = TransformersModel("meta-llama/Llama-3.1-8B")
with model.trace("The Eiffel Tower is in", remote=True):
    hidden = model.model.layers[-1].output.save()

There's more — source tracing into a module's forward (.source), persistent edit(), skip(), scan() for shapes, cache(), session(), and the vLLM and diffusion runtimes. See docs/ and NNsight.md.

Your own model

Any torch.nn.Module works — wrap it in NNsight and the whole tree becomes traceable:

import torch
from nnsight import NNsight

net = torch.nn.Sequential(torch.nn.Linear(5, 10), torch.nn.Linear(10, 2))
model = NNsight(net)

with model.trace(torch.rand(1, 5)):
    model[0].output[:] = 0                 # zero the first layer's output
    out = model.output.save()

print(out)                                 # [1, 2], computed with layer 0 zeroed

Using nnsight from an LLM agent

Give an agent up-to-date nnsight knowledge one of these ways:

  • Skills — in Claude Code: /plugin marketplace add https://github.com/ndif-team/skills.git then /plugin install nnsight@skills. In OpenAI Codex: skill-installer install https://github.com/ndif-team/skills.git.
  • Context7 MCP — add use context7 to prompts, or point your MCP client at https://mcp.context7.com/mcp (see Context7).
  • Docs in context — hand the agent CLAUDE.md (routes to the task docs) and NNsight.md (the internals).

Learn more

Citation

If you use nnsight in your research, please cite:

@article{fiottokaufman2024nnsightndifdemocratizingaccess,
      title={NNsight and NDIF: Democratizing Access to Foundation Model Internals},
      author={Jaden Fiotto-Kaufman and Alexander R Loftus and Eric Todd and Jannik Brinkmann and Caden Juang and Koyena Pal and Can Rager and Aaron Mueller and Samuel Marks and Arnab Sen Sharma and Francesca Lucchetti and Michael Ripa and Adam Belfki and Nikhil Prakash and Sumeet Multani and Carla Brodley and Arjun Guha and Jonathan Bell and Byron Wallace and David Bau},
      year={2024},
      eprint={2407.14561},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2407.14561},
}

About

The nnsight package enables interpreting and manipulating the internals of deep learned models.

Topics

Resources

Code of conduct

Contributing

Stars

1.1k stars

Watchers

3 watching

Forks

Releases

Used by

Contributors

Languages