808 karma · joined November 11, 2020
edit: I know that it mentions its not modern, but these kinds of details have major implications in terms of the representations a model can learn, which is in many ways the most important part!
[1] https://youtu.be/H7_d_sgui6o?t=4436 (timestamped url)
Also related to what I was saying see these:
https://arxiv.org/pdf/2609.16372v1
https://arxiv.org/pdf/2609.01449v1 (this one is quite mindblowing cause they noise they the actual input space every time but it still works)
Not implausible but what exactly because sometimes when people throw that phrase its just aggrieved cope
[1] https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...
The video is timestamped to open at the comparison frame. I don't think an LLM can tell the quality difference without direct reference for comparison
Yes but for generation LingBot seems uniquely compelling https://technology.robbyant.com/lingbot-vision because it has a very strong spatial prior