2021

MiniTorch: inside a learning framework

Implementing the machinery beneath a neural network, from automatic differentiation to tensor kernels and convolutions.

In brief

My contribution: core operators within the provided MiniTorch teaching framework: gradients, tensor layout, accelerated CPU kernels, and neural-network primitives.

Evidence: implementation walkthrough, archived training logs, and an activation visualization. No new benchmark or speedup is claimed.

I implemented components of MiniTorch, an educational framework for understanding how deep-learning libraries work. Starting from the provided framework and assignments, my work filled in core operations across its modules.

Four feature maps from an MNIST hidden layer, showing different responses to a handwritten digit three.
MNIST hidden-layer visualization recorded in my Module 4 project README.

What I implemented

  • Automatic differentiation: reverse traversal of a computation graph and gradient accumulation through scalar and tensor operations.
  • Tensor operations: indexing, shapes, strides, broadcasting, maps, reductions, and matrix multiplication.
  • Efficient execution: Numba-parallelized CPU operations and CUDA kernel exercises.
  • Neural-network building blocks: one- and two-dimensional convolutions, pooling, softmax, log-softmax, and dropout.

From operators to models

The final module connected these components to handwritten-digit classification and sentiment classification. The saved work includes training logs and a visualization of hidden-layer activations from the MNIST model.

One engineering decision: layout before speed

A tensor's shape is not its storage layout. My indexing implementation maps coordinates to storage through strides and handles broadcasting across trailing dimensions. The CPU map kernel uses a direct-storage fast path only when input and output shapes and strides match. Otherwise it computes the appropriate indices. That distinction matters: a fast loop that assumes contiguous storage can silently give incorrect results for a different view of the same data.

How correctness was checked

The archived test suite contains explicit convolution examples and property-based checks for batched and multichannel inputs. Its gradient checker compares backpropagation with a central finite difference at sampled input positions (epsilon 10-6, absolute and relative tolerance 10-2). These are the checks in the archive, not a claim that the historical environment has been rerun for this page.

Reverse-mode differentiation traverses the computation graph backward, accumulating derivative contributions by node before reaching the leaves. This accumulation is essential when a value is reused: a single downstream path does not account for its full effect on the output.

What the training logs do and do not show

The MNIST example connected two convolution layers, pooling, a hidden linear layer, dropout, and log-softmax. Although the code prepared a 500-image validation slice, its logging loop checked only the first 16 images. A saved count of 16 correct predictions is therefore a small development check, not 100% accuracy on the MNIST test set. A stronger evaluation would run every held-out batch with training-only behavior disabled and report the full denominator.

The interesting part of this project is the implementation: connecting tensor layout, gradient propagation, and model training in one small system. The archived logs are development runs, not a benchmark claim about model accuracy or GPU performance.

Project context

My contribution was an implementation of MiniTorch's core operations, not the creation of the original framework. The MiniTorch documentation describes the teaching project and its modules. My archived assignment repository is not currently publicly accessible.

Python / NumPy / Numba / automatic differentiation / tensor programming