You can read this for a comprehensive deep dive. https://arxiv.org/pdf/2502.0163...

yorwba · 2025-05-20T20:37:44 1747773464

The paper you link to is about a different way to create embeddings at the input layer. In no way does it match your claimed description.

ankit219 · 2025-05-20T21:02:57 1747774977

I simplified what i wrote. There is an off accelerator memory where the embeddings are stored and queried at inference time, i did not want to get into details. That is how you reduce the in memory RAM. There are definitely more things going on in the paper as it builds upon the concept I described. The central idea remains the same: you have input embedding layers which map text to continuous vectors. Instead of loading all these layers at runtime, you can break it per layer at training time, and then fetch the required ones from a separate store during inference. Would not be in RAM. Per layer is not mentioned in the paper. But surely it's not a great leap from the paper itself?

yorwba · 2025-05-21T05:45:07 1747806307

The name "per-layer embeddings" is all we have to go on, and there are currently no published papers (that I'm aware of) using any similar mechanism, so, yes, it's a huge leap from a paper that doesn't mention per-layer anything.

It's fine to speculate based on the name, but don't pretend that it's a known technique when it clearly isn't.

krackers · 2025-05-21T06:46:29 1747809989

Someone [1] inspected dimensions of the embedding component of model and it seems GP was on the right track. Assuming I understood correctly in [2], it does seem to be the embedding of the input tokens which is passed directly into each layer.

I have not looked at the model but since the embedding dimension of 256 seems quite small (for reference according to [3] the old Gemma 1B had 1152 dimension input embedding), I'm guessing that this is not done _in lieu_ of the main input embedding to first layer, but in addition to it.

[1] https://twitter.com/cccntu/status/1925043973170856393

[2] https://news.ycombinator.com/edit?id=44048662

[3] https://developers.googleblog.com/en/gemma-explained-whats-n...