Can somebody explain to me the sinus wave positional encoding thing? The naïve approach would be to just add number indices to the tokens, wouldn’t it?
- It should output a unique encoding for each time-step (word’s position in a sentence)
- Distance between any two time-steps should be consistent across sentences with different lengths.
- Our model should generalize to longer sentences without any efforts. Its values should be bounded.
- It must be deterministic.
Your example contradicts the 'values should be bounded' criterion as it generalizes to longer sentences.They use indices for models like vision transformers with a fixed number of patches but for variable length context I think it's more beneficial to use encodings that can also capture the relative distance.