Due to nuances of how data is split across the various model files, the implementation involved some non-trivial steps in order for everything to be allocated correctly. The changes are here [1].
[1] https://github.com/ggerganov/llama.cpp/commit/5b8023d9354010...