10,660 karma · joined March 2, 2010
albzey@gmail.com
http://www.az2000.de
https://github.com/albertz
https://twitter.com/albertzeyer
Research Scientist, PhD; Deep Learning, Speech Recognition, NLP: http://www-i6.informatik.rwth-aachen.de/~zeyer/ https://scholar.google.com/citations?user=qrh5CBEAAAAJ&hl=en
[ my public key: https://keybase.io/albertzeyer; my proof: https://keybase.io/albertzeyer/sigs/1NQp461oCwDiD4vB8HKfiZAkztYALFalSYwHFsnX8IU ]
https://github.com/albertz/PyCParser/blob/master/demos/disse...
The code seems based on llama.cpp and GGML.
I don't fully understand why it is a standalone project. The readme discusses this: DwarfStar 4 is a small native inference engine specific for DeepSeek V4 Flash. It is intentionally narrow: ...
I think the only bigger difference in DeepSeek V4 vs other models is maybe the type of self-attention. And that leads to: KV cache is actually a first-class disk citizen.
But I still feel like those changes could have been implemented as part of some of the other local engines.
I also assume more models will come out, not just from DeepSeek but also from others, and they might share similar self-attention approaches, that would benefit from a similar KV cache implementation.
> The temperature in Finnish saunas is 80 to 110 °C (176 to 230 °F), usually 80–90 °C (176–194 °F)
And with that temperature, I think 10–15 minutes are pretty standard.
I know many computer science colleagues who were not exposed to programming during that age and only later came to it.
I feel kind of lucky that somewhat randomly I stumbled into computer programming (because XtreeGold could show the content of files, and I was learning to understand BAT-files by looking into them) during that age, and that's what I do now.
There are probably a lot of things you were not exposed during that age, that could have been the perfect match.
There are also lots of kids who just play games, or video games, do sports, watch films or so during that age, without really being exposed to any "potential useful" activities. Some parents would maybe even say that this is how it should be.
As a parent, I guess a good advice would be to try to expose your child to as much things as possible, without forcing it to do anything of course.
v0.5 123M Parameters
v1: 700M Parameters
v2mini-eval1: 300M Parameters
I would not call this LLM. This is not large. It's just a normal-sized LM. Or even small.
(It's also not a small LLM.)
This is released under GPL.
I wonder, who is K1n9_Duk3? Does he have the rights to actually release this, and put it under GPL?
What does "reconstructed" mean? Is this disassembled? And if so, is it really ok to put this under GPL then?
In that discussion, most of the same points as in this article were already discussed, specifically some async DNS alternatives.
See also here the discussion: https://github.com/crystal-lang/crystal/issues/13619
This reminds me of variational noise (https://www.cs.toronto.edu/~graves/nips_2011.pdf).
If it is random noise on the input, it would be like many of the SSL methods, e.g. DINO (https://arxiv.org/abs/2104.14294), right?
How does this fit together with a startup? Would investors happily invest into this knowing not to expect anything in return for at least the next 5-10 years?
However, I regret this decision. Git-Annex is not usable anymore on my data because the amount of files has grown so much (millions) and Git-Annex is just too slow (it takes minutes up to even hours for some Git operation, and the FS is decently fast). I assume I would not have had those problems with Perkeep.
https://www.vulture.com/2019/07/motion-smoothing-is-ruining-... https://www.filmindependent.org/blog/hacking-film-24-frames-...
Once you consider that you anyway need to write very different kind of code for RPython, then maybe just using Nim or some other language is a better idea?
I also had a very similar bug a while ago, broken gradients due to non-contiguous data for masked_select: https://github.com/pytorch/pytorch/issues/99638
In my case, it was easier to identify: I had another implementation of my loss function before that did not use masked_select. But then I thought I can be clever and use masked_select to take out the non-masked frames and calculate the loss only on those. But it wasn't working. Also, it only happened for some models, not for all. It turns out, it was always happening when the data coming out of the model was non-contiguous.
I think the bugs with non-contiguous data are not so uncommon. I wonder how much of that we still have.