Yikes...
30 karma · joined June 28, 2023
Yikes...
Hopefully, future models can be trained to be even more aware of external knowledge, accessible through web search / RAG / whatever it will be then, and might not need to internalize much knowledge at all.
But in the end, it's all up to the units/quantities we choose to measure, no? If we, say, decided to measure "Squenergy" in Sqoules, with 1Sq² = 1J, then suddenly, squenergy does increase linearly with speed! The formula for kinetic Squenergy becomes sqrt(m/2)v.
Of course this complicates other stuff, like potential Squenergy becoming sqrt(MgH), it not being additive, etc.
How do you folks deal with these massively increased threats to self-hosted open source apps?
So this means the sequence of μₙ will perform a kind of random walk that can stray arbitrarily far from 0 and is almost sure to eventually do so.
ELMo: https://arxiv.org/abs/1802.05365 BERT: https://arxiv.org/abs/1810.04805 ERNIE: https://arxiv.org/abs/1904.09223v1 Big Bird: https://arxiv.org/abs/2007.14062
[1] https://support.google.com/business/answer/7690269?hl=en
The current workarounds to make this happen in python are quite ugly imho, e.g. Pytorch spawns multiple python processes and then pushes data between the processes through shared memory, which incurs quite some overhead. Tensorflow on the other hand requires you to stick to their Tensor-dsl so that it can run within their graph engine. If native concurrency were a thing, data loading would be much more straightforward to implement without such hacks.