Finding 1
The Transformer never attended across words
batch_first=True but fed (batch, sequence) tensors. So self-attention mixed the 16 claims in a mini-batch, never the words within one claim. Run on a single claim, the whole network reduces exactly to a lookup table of per-word scores. That makes every prediction on this site fully explainable.