Neural Network
A neural network as a function: every neuron visible, ReLU kinks as building blocks of the output, weights editable per neuron, parameter count; curve, plane and digits modes.
Neural networksTry itCollection
From neurons and gradients to transformers, RLHF and agents: machine learning you can touch.
Apps
A neural network as a function: every neuron visible, ReLU kinks as building blocks of the output, weights editable per neuron, parameter count; curve, plane and digits modes.
Neural networksTry itLogits become a distribution via softmax(z/T): temperature, greedy, sampling, top-k and top-p with a visible cut-off line, entropy, a decision tree of the generation and loops at T = 0.
Neural networksTry itText splits into tokens: byte-level BPE is trained live, the merge slider shows frequent words fusing into one token, plus token details, a language comparison and the share of the context window.
Linguistics & phoneticsTry itText from counts of the last n − 1 tokens: count table P(w|h), the path from gibberish to plausible text to verbatim copying, the explosion of possible contexts, smoothing and backoff; German and English corpora.
Linguistics & phoneticsTry itA model is a form with parameters: polynomials of degree 0 to 12 by least squares, residuals as squares (area = loss), held-out error over the degree and parameters that explode when overfitting.
Probability & statisticsTry itOne artificial neuron: weighted sum, activation, a decision line in the input plane with the w arrow and bias offset; perceptron rule and gradient step solve AND and OR, XOR stays at 75 %.
Neural networksTry itA drawn digit flows through a real network trained on MNIST (784 → 64 → 32 → 10): weight images per neuron, forward wave, softmax bars and test digits it gets confidently wrong.
Neural networksTry itTraining lab for neural networks in 3D: neurons as tiles showing their output, edges by weight, spiral, XOR, moons, layers, activations and regularisation live; plus a digits mode.
Neural networksTry itLoss surface of a real line fit in 3D: θ ← θ − η∇L with stability limit 2/λmax, two valleys and a saddle, GD, momentum and Adam racing, SGD noise and the fitted line in the model panel.
Neural networksTry itBackpropagation on a 2-2-2 network with real numbers in six steps: forward pass, loss, δ at the output and in the hidden layer, gradient per weight, only the update lowers the loss — checked numerically.
Neural networksTry itReading training dynamics from the loss curves: learning rate, batch size, epochs, training and validation loss, early stopping at the validation minimum, L2, dropout, augmentation and more data against overfitting.
Neural networksTry itReal word vectors — English GloVe, German fastText (cc.de.300): relations as parallel arrows, king − man + woman ≈ queen with hits, near misses and learned bias, cosine against distance and sentences as a path with next-word candidates from GPT-2 or GerPT2.
Linguistics & phoneticsTry itWords as vectors in a space of meaning: similarity as cosine in the full space, analogy arrows, neighbours and three 3D shadows (map, PCA, axes) with how much neighbourhood each one keeps.
Neural networksTry itStatic vs. contextual word vector in 3D: an ambiguous word starts between two islands of meaning and is pulled layer by layer by the context words, in proportion to their attention weights.
Linguistics & phoneticsTry itSelf-attention made visible: query, keys, softmax(QKᵀ/√d)·V, five specialised heads, a matrix heat map with causal mask, a second layer and a by-hand calculation with three tokens.
Neural networksTry itThe GPT-2 architecture to walk through: residual stream through 12 blocks, masked attention as bridges, MLP d → 4d → d, unembedding and softmax, one token at a time; parameters per component and GPT-3 for comparison.
Neural networksTry itYou compare answers: a reward model learns from them (σ(r_A − r_B)), the policy shifts under a KL leash; with a loose leash the flattering answer wins (reward hacking); plus the DPO route.
Neural networksTry itRetrieval-augmented generation step by step: chunks, embeddings, search by cosine, BM25 or hybrid, top-k cut-off, prompt with token budget and an answer with sources; missed chunks lead to hallucination.
Neural networksTry itElementary cellular automata with all 256 rules: rule 90 jumps to any generation by formula, rule 30 must compute every step (irreducibility); plus a prediction game and an LLM's compute budget per token.
Computer scienceTry itLeast-squares line on lecture data: errors as squares, SST = SSR + SSE as a right triangle, SSE as a bowl over slope and position; Anscombe, a parabola and a confounder show what R² does not reveal.
Probability & statisticsTry itMany training sets from the same world as a family of models: average model against the truth (bias), spread of the models (variance) and the exact decomposition test error = bias² + variance + noise over polynomial degree or ridge λ.
StatisticsTry itA query point, its k neighbours with distance shape and majority vote: decision regions from k = 1 (jagged) to k = 40 (smooth), train and test accuracy over k, metric and feature scaling change the neighbourhood.
Machine LearningTry itTwo features, a linear score, a sigmoid surface in 3D: the threshold plane cuts it in a straight line, log loss per point, gradient descent step by step, XOR as the model's limit.
StatisticsTry itPrior times per-feature likelihoods: posterior field with marginal densities, axis-aligned ellipses and a quadratic boundary; correlated data show what the independence assumption misses; plus a spam mode with log-odds.
Machine LearningTry itMaximum margin and support vectors, solved exactly with SMO: points without a ring do not move the boundary, soft margin C, slack ξ, linear, lift (paraboloid in 3D) and RBF kernels.
Machine LearningTry itAxis-parallel splits by information gain: ΔI for every candidate threshold, tree diagram linked to the plane, depth, minimum leaf size and pruning, train against test accuracy over the depth.
Machine LearningTry itEight classifiers learn on the same data: kNN, logistic regression, linear and RBF SVM, naive Bayes, decision tree, random forest and a neural network — their decision boundaries reveal the model family, train and test accuracy side by side expose overfitting.
Machine LearningTry itSplit, validation curve, k-fold cross-validation and a one-time test on a polynomial regression — plus data leakage and tuning on the test set as counterexamples.
Machine LearningTry itSlide a classifier's threshold and watch cases move between TP, FP, FN and TN — accuracy, precision, recall, specificity and F1 with their formulas, the precision–recall curve and the accuracy paradox with rare positives.
Machine LearningTry itTwo score distributions, the threshold sweeps: the ROC curve is traced point by point, AUC as the share of correctly ranked pairs, ROC stays the same at any prevalence, two models with crossing curves compared.
Machine LearningTry itMany random decision trees vote: the majority smooths the jagged boundaries of single trees and lowers the test error — bootstrap samples, feature randomness, out-of-bag error and test error over the number of trees.
Machine LearningTry itDecision stumps trained one after another on weighted data: mistakes grow, the next stump targets them, vote weight α per round, the ensemble's exact staircase boundary, train and test error over up to 60 rounds.
Machine LearningTry itLloyd's algorithm as a chain of states without labels: assign and move centres, Voronoi cells, inertia per half-step, random vs. k-means++, local minima, elbow with silhouette and typical failure cases.
Machine LearningTry itPrincipal axes of a 3D point cloud as eigenvectors of the covariance: projection onto k components with perpendiculars, reconstruction error = sum of the discarded λ, scree plot, standardising and the difference from the regression plane.
Machine LearningTry itThe same prompt through pretraining, SFT, RLHF/DPO and LoRA: the knowledge comes from pretraining, the later stages mainly shift which answer is likely.
Neural networksTry itAn LLM agent is a loop: the model writes tool calls as text, a program executes them and every observation grows into the context — with and without tools compared, including a tool failure and a prompt injection.
Neural networksTry itA model trained on past hiring decisions sorts 200 applications. Delete the group from the data and the ratio barely improves, because a proxy feature carries the bias on.
AI literacyTry itA random number generator is a fixed rule: same seed, same sequence, and eventually the loop closes. Middle-square, a linear congruential generator with freely adjustable parameters, RANDU, a shift register and xorshift side by side — good ones and deliberately bad ones. Four views reveal what a list of numbers hides: the rule with values plugged in, the cycle as a rho shape, the scatter of consecutive pairs plus a rotatable cloud of triples (RANDU collapses into 15 planes), and a histogram.
CryptographyTry itAn AI answer sounds certain and cites sources. Break it into claims and check each against the original document: supported, refuted or made up.
AI literacyTry itCollections
Mathematical questions open for decades, for which an AI model proposed solutions in 2026, as interactive apps: understand the problem, experience the result.
24 appsOpen collectionDNA, replication, transcription, translation and cell division as explorable 3D models.
11 appsOpen collectionPerspective, ornament, structure and space: how art and architecture work with mathematics.
9 appsOpen collectionHyperloop, autonomous vehicles, warehouses and conveyors: logistics as interactive simulation.
17 appsOpen collectionHeart, eye, voice, skeleton and nerves: the human body in 3D, layer by layer.
8 appsOpen collectionMonty Hall, Simpson, Berkson and base rates: where intuition fails us with numbers.
14 appsOpen collectionPlanets, stars, black holes and space probes: the cosmos to explore.
6 appsOpen collectionRiemann, P vs NP, Poincaré and other great questions of mathematics, made tangible.
16 appsOpen collectionCounting, arithmetic and shapes: maths for beginners, playful and clear.
6 appsOpen collectionWant to use these apps in your courses? Let’s find out what works for you.
One form, three matters. We answer personally.
We show heyprof on your own material. Tell us what you teach and we will prepare the conversation around it.