This project is an attempt to implement the Transformer architecture completely from scratch in C++ — without using deep-learning frameworks such as PyTorch or TensorFlow.
The implementation follows the Attention is All You Need paper and includes:
- Token and vector embedding layer
- Sinusoidal positional encoding
- Multi-Head Self-Attention (query–key–value formulation)
- Scaled dot-product attention
- Add & Norm / Layer Normalization
- Feed-Forward (MLP) block
- Stacked encoder architecture
- Cross-entropy loss and gradient descent
All tensor operations (matrix multiplication, transpose, softmax, etc.) are written manually in C++ to avoid dependence on external ML libraries.
To speed up training and attention operations: CUDA kernels are being written for matrix multiplication and softmax.
The forward pass of the encoder stack is stable and tested with dummy input sequences.
The main bottleneck is implementing a complete and numerically stable automatic differentiation system for the backward pass. I am exploring two options:
-
Writing my own computational graph + autodiff system
-
Implementing a more manual gradient tracking mechanism for each layer
This project is ongoing. Due to the complexity of GPU acceleration and writing a gradient system from scratch, development is slower but highly educational. Contributions, suggestions, and benchmarking ideas are welcome.