Accessing Attention on MLX #3246
Replies: 1 comment
|
MLX doesn't have PyTorch-style forward hooks, but since models are plain Python objects you can reach intermediates directly. Hidden states / activations: any Attention weights specifically are trickier: mlx-lm's attention uses the fused scores = mx.softmax((q @ k.swapaxes(-1, -2)) * scale + mask, axis=-1)
out = scores @ v # stash `scores` before this lineGrab the module ( |
Uh oh!
There was an error while loading. Please reload this page.
“How can intermediate attention activations (e.g., attention heads or attention weights) be accessed or inspected when running transformer models with MLX?”
All reactions