Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Stars

2 stars

Watchers

2 watching

Forks

Repository files navigation

Transformer from Scratch

A complete, educational implementation of the Transformer architecture from the paper "Attention is All You Need" (Vaswani et al., 2017) using TensorFlow.

Overview

This project implements a full Transformer model with:

  • Multi-Head Attention mechanism
  • Positional Encoding (sinusoidal)
  • Encoder-Decoder architecture
  • Feed-Forward Networks
  • BPE Tokenization for efficient text processing
  • French-to-English translation task as demonstration

Project Structure

src/
├── models/                  # Core Transformer components
│   ├── attention.py
│   ├── multi_head_attention.py
│   ├── positional_encoding.py
│   ├── transformer_blocks.py
│   └── transformer.py
├── tokenization/            # Text tokenization utilities
│   └── tokenizer.py         # BPE tokenizer
├── tasks/                   # Application tasks
│   ├── translation_task.py       # Word-level
│   └── translation_task_bpe.py   # BPE (recommended)
└── utils/                   # Utilities
    ├── training.py
    └── inference.py

notebooks/
├── exploration.ipynb        # Interactive components
└── translation.ipynb        # Translation pipeline

docs/                        # Documentation
weights/                     # Model weights
data/                        # Training data

Quick Start

Installation

python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install -e .

Run Translation

from src.tasks.translation_task_bpe import train_translation_model, translate

model = train_translation_model(epochs=50, save_weights=True)
output = translate(model, "bonjour comment allez vous")
print(output)  # → "hello how are you"

Explore in Jupyter

jupyter notebook notebooks/exploration.ipynb
jupyter notebook notebooks/translation.ipynb

Architecture

Transformer Pipeline

Input → Embedding + Positional Encoding
         ↓
    [Encoder Block] × N (2 layers)
         ↓
    [Decoder Block] × M (2 layers)
         ↓
    Linear + Softmax → Output

Key Components

Component Formula Purpose
Scaled Dot-Product Attention $\text{Attention}(Q,K,V) = \text{softmax}(\frac{QK^T}{\sqrt{d_k}})V$ Compute context from all positions
Multi-Head Attention $\text{MultiHead} = \text{Concat}(\text{head}_1,...,\text{head}_h)W^O$ Multiple representation subspaces
Positional Encoding $PE_{pos,2i} = \sin(\frac{pos}{10000^{2i/d}})$ Preserve position information
Feed-Forward $\text{FFN}(x) = \text{ReLU}(xW_1 + b_1)W_2 + b_2$ Non-linear transformation

Model Parameters

d_model = 64                    # Model dimension
num_heads = 4                   # Attention heads
dff = 256                       # Feed-forward dimension
num_encoder_layers = 2          # Encoder blocks
num_decoder_layers = 2          # Decoder blocks
dropout_rate = 0.1              # Regularization

Training Results

Dataset: 29 French↔English sentence pairs
Tokenizer: BPE (77 FR tokens, 67 EN tokens)
Epochs: 50

Metric Value
Initial Loss 4.90
Final Loss 0.98
Final Accuracy 76%

Tokenization Methods

Why BPE?

Method Memory Order Preserved Best For
One-Hot 4MB ❌ ✓ Small ML
Bag of Words 40KB ⚠️ ✗ Classification
BPE 400B ✓✓✓ ✓ Modern NLP (BERT, GPT)

BPE (Byte-Pair Encoding):

  • ✓ Efficient memory usage
  • ✓ Preserves word order (critical for translation)
  • ✓ Handles unknown words via subword units
  • ✓ Industry standard across all modern LLMs

Inference

Greedy Decoding

from src.utils.inference import GreedyDecoder
decoder = GreedyDecoder(model, tokenizer, max_length=20)
result = decoder.decode("bonjour")

Beam Search

from src.utils.inference import BeamSearchDecoder
decoder = BeamSearchDecoder(model, tokenizer, beam_width=5, max_length=20)
result = decoder.decode("bonjour")

References

  1. Attention is All You Need (Vaswani et al., 2017)

  2. BERT: Pre-training of Deep Bidirectional Transformers (Devlin et al., 2018)

  3. Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau et al., 2014)

License

MIT License - See LICENSE file

Author

AI Research Student

Installation

pip install -r requirements.txt

Comment suivre le projet

  1. 01_attention.py : Comprendre les maths de l'attention et le "Scaled Dot-Product"
  2. 02_multi_head_attention.py : Voir comment combiner plusieurs "têtes"
  3. 03_positional_encoding.py : Ajouter l'information de position
  4. 04_transformer_blocks.py : Assembler attention + feedforward
  5. 05_transformer.py : Le Transformer complet (Encoder + Decoder)
  6. 06_training.py : Entraîner sur une tâche simple (copie de séquence)
  7. 07_inference.py : Tester le modèle
  8. notebook_exploration.ipynb : Visualisations et expérimentations

Tâche pédagogique

On commencera par une tâche très simple : apprendre à copier une séquence d'entiers.

Exemple :

  • Input: [3, 1, 4, 1, 5]
  • Output: [3, 1, 4, 1, 5]

Cette tâche simple permet de vérifier que chaque composant fonctionne correctement avant de passer à des tâches plus complexes.

Concepts clés explorés

  • Scaled Dot-Product Attention
  • Multi-Head Attention
  • Positional Encoding (sinusöıdal)
  • Encoder Stack
  • Decoder Stack (avec Masked Attention)
  • Training & Inference

Exécution

# Entraîner le modèle
python 06_training.py

# Tester le modèle
python 07_inference.py

Bon apprentissage ! 🚀

About

No description, website, or topics provided.

Resources

Contributing

Stars

2 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages