MicroGPT::AttentionHead
Constructors
Instance methods
attn_weights
Sourceattn_weights=(attn_weights : Mat | Nil)
SourceBackward: returns {dq, dk, dv} for the fused projection backward
forward(q : Mat, k : Mat, v : Mat, mask : Mat | Nil = nil) : Mat
Forward: compute attention from pre-projected Q, K, V If mask is provided, it's added to scores (pre-computed -inf pattern)
head_dim
Sourcek=(k : Mat | Nil)
Sourceq=(q : Mat | Nil)
Sourcev=(v : Mat | Nil)
Source