30 — Sequence Models

GRU — Gated Recurrent Unit#

Sequential✓ Mathematical
◆ The PatternLSTM's streamlined sibling — two gates, one state vector

GRU simplifies LSTM by merging cell+hidden state and using only two gates: update and reset.

zₜ = σ(Wz·[hₜ₋₁,xₜ])    rₜ = σ(Wr·[hₜ₋₁,xₜ])
z = update gate  |  r = reset gate
h̃ₜ = tanh(W·[rₜ⊙hₜ₋₁, xₜ])
Candidate hidden state — gated by reset
hₜ = (1−zₜ)⊙hₜ₋₁ + zₜ⊙h̃ₜ
Interpolate between old and new state via update gate
// RNN vs GRU vs LSTM — parameter count comparison
When to use which: RNN — quick baseline. GRU — best default for seq tasks. LSTM — complex long-range deps. Transformer — lots of data + GPU.
Pattern bridge: GRU merges forget and input into a single update gate — elegant reduction. In markets, RSI compresses momentum into one number.
← Previous
LSTM
Open in the full reader, with the topic sidebar →