neural_network.sliding_window_attention¶
– - - - - - - - - - - - - - - - - - - - - - -
Name - - sliding_window_attention.py Goal - - Implement a neural network architecture using sliding
window attention for sequence modeling tasks.
- Detail: Total 5 layers neural network
Input layer
Sliding Window Attention Layer
Feedforward Layer
Output Layer
Author: Stephen Lee Github: 245885195@qq.com Date: 2024.10.20 References:
Choromanska, A., et al. (2020). “On the Importance of Initialization and Momentum in Deep Learning.” Proceedings of the 37th International Conference on Machine Learning.
Dai, Z., et al. (2020). “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.” arXiv preprint arXiv:2006.16236.
[Attention Mechanisms in Neural Networks](https://en.wikipedia.org/wiki/Attention_(machine_learning))
– - - - - - - - - - - - - - - - - - - - - - -
Attributes¶
Classes¶
Sliding Window Attention Module. |
Module Contents¶
- class neural_network.sliding_window_attention.SlidingWindowAttention(embed_dim: int, window_size: int)¶
Sliding Window Attention Module.
This class implements a sliding window attention mechanism where the model attends to a fixed-size window of context around each token.
- Attributes:
window_size (int): The size of the attention window. embed_dim (int): The dimensionality of the input embeddings.
- forward(input_tensor: numpy.ndarray) numpy.ndarray¶
Forward pass for the sliding window attention.
- Args:
- input_tensor (np.ndarray): Input tensor of shape (batch_size,
seq_length, embed_dim).
- Returns:
np.ndarray: Output tensor of shape (batch_size, seq_length, embed_dim).
>>> x = np.random.randn(2, 10, 4) # Batch size 2, sequence >>> attention = SlidingWindowAttention(embed_dim=4, window_size=3) >>> output = attention.forward(x) >>> output.shape (2, 10, 4) >>> (output.sum() != 0).item() # Check if output is non-zero True
- attention_weights¶
- embed_dim¶
- window_size¶
- neural_network.sliding_window_attention.rng¶