Gemma 3n Architectural Innovations – Speculation and poking around in the model
old.reddit.com
old.reddit.com
Having more than one embedding is something I've tried myself, but not separate ones for each layer.
I'm guessing it's something like h_{l+1} = MultiHeadSelfAttentionWithPositionEncodingBakedIn(MLP(h_l) + embed_l(token_ids)). So it's probably really easy to implement on toy problems to see if it works.
Then you'd figure out a set of toy tasks that you like and think are important.
In this particular case you take something like NanoGPT, go to model.py, go to class GPT, go to __init__, modify the self.transformer ModuleDict by changing nn.Embedding to a ModuleList of nn.Embedding, then you change the for loop at line 180 to loop over a range, modify forward by adding x = x + self.transformer.wte[i], something like that I think.
I haven't tried yet though (I've got a terrible cold, so I am on social media instead of doing anything sensible).
"4x gated residual streams" look quite weird. Is there any paper or technique report for this?