算法面试常考【手撕MHA】
·
🎯 什么是 MHA(Multi-Head Attention)?
MHA 是 Transformer 中最关键的一层,它的思想是:
“将输入映射到多个不同的子空间(多个 Head),分别做 Attention,再拼接在一起。”
它的核心组件是 Scaled Dot-Product Attention + 多个 Head 并行处理。
✅ 整体公式结构(以矩阵形式)
给定:
-
输入矩阵 X∈Rn×dmodelX \in \mathbb{R}^{n \times d_{\text{model}}}X∈Rn×dmodel
-
对每个 head,我们使用不同的线性变换权重:
- Q=XWQQ = XW^QQ=XWQ, K=XWKK = XW^KK=XWK, V=XWVV = XW^VV=XWV
-
然后计算每个 head 的 attention:
Attention(Q,K,V)=softmax(QKTdk)V \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V Attention(Q,K,V)=softmax(dkQKT)V
-
多个 head 的输出拼接后再乘一个输出矩阵:
MHA(X)=Concat(head1,…,headh)WO \text{MHA}(X) = \text{Concat}(head_1, \dots, head_h) W^O MHA(X)=Concat(head1,…,headh)WO
✅ 用 Python 手撕简易版 MHA(不依赖框架)
我们用最基本的 Python+NumPy(因为矩阵运算太繁琐了,没 numpy 会非常啰嗦)
import numpy as np
def softmax(x):
# x shape: (n, n)
x = x - np.max(x, axis=-1, keepdims=True)
exp_x = np.exp(x)
return exp_x / np.sum(exp_x, axis=-1, keepdims=True)
def scaled_dot_product_attention(Q, K, V):
d_k = Q.shape[-1]
scores = Q @ K.T / np.sqrt(d_k) # (n, n)
weights = softmax(scores) # (n, n)
return weights @ V # (n, d_v)
def mha(X, num_heads=2):
n, d_model = X.shape
assert d_model % num_heads == 0
d_k = d_v = d_model // num_heads
# 初始化参数:W^Q, W^K, W^V for each head
W_Q = np.random.randn(num_heads, d_model, d_k)
W_K = np.random.randn(num_heads, d_model, d_k)
W_V = np.random.randn(num_heads, d_model, d_v)
W_O = np.random.randn(num_heads * d_v, d_model)
heads = []
for i in range(num_heads):
Q = X @ W_Q[i] # (n, d_k)
K = X @ W_K[i] # (n, d_k)
V = X @ W_V[i] # (n, d_v)
head = scaled_dot_product_attention(Q, K, V) # (n, d_v)
heads.append(head)
concat = np.concatenate(heads, axis=-1) # (n, num_heads * d_v)
output = concat @ W_O # (n, d_model)
return output
🔍 示例用法
X = np.random.randn(4, 8) # 假设有 4 个 token,每个是 8 维
out = mha(X, num_heads=2)
print(out.shape) # (4, 8)
✅ 面试可能问的 follow-up:
| 面试点 | 说明 |
|---|---|
| 为什么要用多头? | 多个子空间学习不同的注意力模式(方向),提升表达能力 |
| 为什么要缩放(√d)? | 防止 dot-product 结果过大导致 softmax 梯度消失 |
| Attention 是线性的吗? | Attention 本质上是非线性的,因为 softmax 不是线性变换 |
| 能不能共享 Q/K/V 参数? | 可以,但会损失表达能力 |
| MHA 为什么比单头好? | 单头可能关注局部模式,多头能并行关注不同区域/语义 |
更多推荐



所有评论(0)