A complete, educational implementation of the Transformer architecture from the paper "Attention is All You Need" (Vaswani et al., 2017) using TensorFlow.
This project implements a full Transformer model with:
- Multi-Head Attention mechanism
- Positional Encoding (sinusoidal)
- Encoder-Decoder architecture
- Feed-Forward Networks
- BPE Tokenization for efficient text processing
- French-to-English translation task as demonstration
src/
├── models/ # Core Transformer components
│ ├── attention.py
│ ├── multi_head_attention.py
│ ├── positional_encoding.py
│ ├── transformer_blocks.py
│ └── transformer.py
├── tokenization/ # Text tokenization utilities
│ └── tokenizer.py # BPE tokenizer
├── tasks/ # Application tasks
│ ├── translation_task.py # Word-level
│ └── translation_task_bpe.py # BPE (recommended)
└── utils/ # Utilities
├── training.py
└── inference.py
notebooks/
├── exploration.ipynb # Interactive components
└── translation.ipynb # Translation pipeline
docs/ # Documentation
weights/ # Model weights
data/ # Training data
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install -e .from src.tasks.translation_task_bpe import train_translation_model, translate
model = train_translation_model(epochs=50, save_weights=True)
output = translate(model, "bonjour comment allez vous")
print(output) # → "hello how are you"jupyter notebook notebooks/exploration.ipynb
jupyter notebook notebooks/translation.ipynbInput → Embedding + Positional Encoding
↓
[Encoder Block] × N (2 layers)
↓
[Decoder Block] × M (2 layers)
↓
Linear + Softmax → Output
| Component | Formula | Purpose |
|---|---|---|
| Scaled Dot-Product Attention | Compute context from all positions | |
| Multi-Head Attention | Multiple representation subspaces | |
| Positional Encoding | Preserve position information | |
| Feed-Forward | Non-linear transformation |
d_model = 64 # Model dimension
num_heads = 4 # Attention heads
dff = 256 # Feed-forward dimension
num_encoder_layers = 2 # Encoder blocks
num_decoder_layers = 2 # Decoder blocks
dropout_rate = 0.1 # RegularizationDataset: 29 French↔English sentence pairs
Tokenizer: BPE (77 FR tokens, 67 EN tokens)
Epochs: 50
| Metric | Value |
|---|---|
| Initial Loss | 4.90 |
| Final Loss | 0.98 |
| Final Accuracy | 76% |
Why BPE?
| Method | Memory | Order Preserved | Best For |
|---|---|---|---|
| One-Hot | 4MB ❌ | ✓ | Small ML |
| Bag of Words | 40KB |
✗ | Classification |
| BPE | 400B ✓✓✓ | ✓ | Modern NLP (BERT, GPT) |
BPE (Byte-Pair Encoding):
- ✓ Efficient memory usage
- ✓ Preserves word order (critical for translation)
- ✓ Handles unknown words via subword units
- ✓ Industry standard across all modern LLMs
from src.utils.inference import GreedyDecoder
decoder = GreedyDecoder(model, tokenizer, max_length=20)
result = decoder.decode("bonjour")from src.utils.inference import BeamSearchDecoder
decoder = BeamSearchDecoder(model, tokenizer, beam_width=5, max_length=20)
result = decoder.decode("bonjour")-
Attention is All You Need (Vaswani et al., 2017)
-
BERT: Pre-training of Deep Bidirectional Transformers (Devlin et al., 2018)
-
Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau et al., 2014)
MIT License - See LICENSE file
AI Research Student
pip install -r requirements.txt- 01_attention.py : Comprendre les maths de l'attention et le "Scaled Dot-Product"
- 02_multi_head_attention.py : Voir comment combiner plusieurs "têtes"
- 03_positional_encoding.py : Ajouter l'information de position
- 04_transformer_blocks.py : Assembler attention + feedforward
- 05_transformer.py : Le Transformer complet (Encoder + Decoder)
- 06_training.py : Entraîner sur une tâche simple (copie de séquence)
- 07_inference.py : Tester le modèle
- notebook_exploration.ipynb : Visualisations et expérimentations
On commencera par une tâche très simple : apprendre à copier une séquence d'entiers.
Exemple :
- Input:
[3, 1, 4, 1, 5] - Output:
[3, 1, 4, 1, 5]
Cette tâche simple permet de vérifier que chaque composant fonctionne correctement avant de passer à des tâches plus complexes.
- Scaled Dot-Product Attention
- Multi-Head Attention
- Positional Encoding (sinusöıdal)
- Encoder Stack
- Decoder Stack (avec Masked Attention)
- Training & Inference
# Entraîner le modèle
python 06_training.py
# Tester le modèle
python 07_inference.pyBon apprentissage ! 🚀