Forward, Backward, Update
Key concept: Training cycles through forward values, backward gradients, and a parameter update θ ← θ − η∇θ.
Train a 2→3→3→2 ReLU network on XOR. Each click advances Forward, Backward, or Update; the loss curve records each update.
What to look for: hidden-layer values gradually separate the two XOR classes. The network is learning a feature map.
Open the demo in a new tab for more room.