LEARN WITH DIRECTION

Small steps. Deeper understanding.

A sequence beats a random walk. Choose a path and work through it at your own pace.

01 / Your priority path

OpenAI ML coding prep

Fifty focused implementations: matrices and gradients, vectorized neighbors, entropy, attention, training bugs, and data evaluation. Independent preparation, not confirmed interview questions.

Start this plan
0 of 50 completed0%
01Matrix-vector multiplication and shape contractsUnrated02Multiclass cross-entropy and manual linear-layer gradientsUnrated03A custom autograd operationUnrated04Central-difference gradient checkingUnrated05MLP BackwardHard06Find the missing gradientEasy07Repair a small classifier training stepUnrated08Pairwise squared distancesUnrated09Nearest neighbors with deterministic tiesUnrated10cumsum · each output chooses a prefixMedium11Entropy and exact discrete KL divergenceUnrated12Online mean and varianceUnrated13Scaled attention with fully masked rowsUnrated14Multi-head causal self-attention from scratchUnrated15Causal attention with padding and empty rowsUnrated16KV-cached causal attention and offset masksUnrated17A two-layer NumPy MLP and its complete backward passUnrated18A classifier training step that actually changes parametersUnrated19Overfit one batchEasy20Audit labels, duplicates and split leakageUnrated21Audit duplicates, group leakage and classifier metricsUnrated22Binary confusion metrics with explicit edge policiesUnrated23Fit preprocessing on training data onlyUnrated24Choose a threshold using validation dataUnrated25Stable log-sum-expUnrated26Stable softmax over the correct axisUnrated27Softmax backward without a quadratic JacobianUnrated28Binary cross-entropy from logitsUnrated29Linear regression loss and gradientsUnrated30One k-means assignment/update stepUnrated31Masked mean with an empty-row policyUnrated32Scatter-add with repeated indicesUnrated33Embedding lookup and repeated-token gradientsUnrated34Pad a variable-length batch without guessing validity from token IDUnrated35Masked token loss with correct weightingUnrated36Inverted dropout and train/eval behaviorUnrated37Batch normalization with running statisticsUnrated38Layer normalization from primitive operationsUnrated39SGD with momentumUnrated40Adam and decoupled AdamWUnrated41Correct accumulation for unequal microbatchesUnrated42Global gradient norm clippingUnrated43A complete tiny decoder-only TransformerUnrated44Temperature, top-k and nucleus sampling probabilitiesUnrated45Reproducible minibatchesUnrated46ROC-AUC with tied scoresUnrated47Expected calibration errorUnrated48Covariance with samples as rowsUnrated49Categorical sampling by inverse CDFUnrated50Gaussian maximum-likelihood estimatesUnrated
02 / Start here

Tensor fluency

Twenty-one small puzzles for a big shift in how you think about arrays. Work with shapes, broadcasting, indexing, and vectorization.

Start this plan
0 of 21 completed0%
03 / Build the fundamentals

PyTorch foundations

Build the fundamentals of a working training loop: tensors, gradients, modules, optimizers, and evaluation.

Start this plan
0 of 20 completed0%
04 / Go deeper

Inside the transformer

Move from shape manipulation to attention and the building blocks of modern language models. Implement each piece before putting it together.

Start this plan
0 of 20 completed0%
05 / The complete archived collection

TensorGym, recovered

All 30 original descriptions, with the 20 recovered starters and solutions. Original test data is preserved; incomplete or inconsistent cases are labeled.

Start this plan
0 of 30 completed0%
06 / Strengthen your intuition

ML from first principles

Rebuild your mathematical intuition through classical algorithms, probability, and evaluation. Focus on why each implementation works.

Start this plan
0 of 24 completed0%