Self-attention looks at all positions at once, which means it has no idea which came first. Position has to be added to the input explicitly. Without it, the most sophisticated architecture in the field falls back to the least sophisticated representation there is.
no position, no order