A model with perplexity 100 is as unsure as if it were picking uniformly from 100 words at every step. That is the whole intuition. It also means perplexity is only comparable between models that share a tokenizer, which is why leaderboard numbers so often are not.
same tokenizer, or no comparison