← 返回 bytedance 的题目列表PyTorch Self-Attention
类型:online_judge
Manually implement the Self-Attention mechanism in PyTorch. Specify the dimensions of several important matrices and compute the time complexity. When applying multiple layers of Self-Attention, how do you stabilize the training of Transformers and solve the issues of gradient explosion and GPU memory explosion?
Example
Input
(2, 10, 64)
[(batch_size, seq_len, d_k)]