Forward intermediates of the attention computation, required by the backward pass.
Cache
01 Syntax
Microsoft.VisualBasic.MachineLearning.Transformer.MultiHeadAttention.Cache
02 Fields
| Name | Overloads | Summary |
|---|---|---|
| Qf | 2 | The projected query tensors, one per head. |
| Kf | 2 | The projected key tensors, one per head. |
| Vf | 2 | The projected value tensors, one per head. |
| Probs | 2 | The attention probabilities (softmax output) of every head. |
| Concat | 2 | The heads after concatenation along the last dimension. |
| QueriesInput | 2 | The query side input (identical to the K/V input for self attention). |
| KInput | 2 | The key side input. |
| VInput | 2 | The value side input. |
| CrossAttention | 2 | Whether this is cross attention, i.e. |
03 Members
Qf
The projected query tensors, one per head.
Kf
The projected key tensors, one per head.
Vf
The projected value tensors, one per head.
Probs
The attention probabilities (softmax output) of every head.
Concat
The heads after concatenation along the last dimension.
QueriesInput
The query side input (identical to the K/V input for self attention).
KInput
The key side input.
VInput
The value side input.
CrossAttention
Whether this is cross attention, i.e. queries and keys/values come from different inputs.
Qf
Kf
Vf
Probs
Concat
QueriesInput
KInput
VInput
CrossAttention