460 karma · joined September 4, 2015
Exxa website: https://withexxa.com
Contact: etienne -at- withexxa.com
--
meet.hn/city/fr-Grenoble
Interests: AI/ML, Research, Robotics, Startups
---
- "Benevolent Dictators" of companies or projects have to obey the law - They can't forbid competition or alternatives - Every participant can leave at any time - If they burn the organization to the ground, the worst case scenario is the organization get replaced and people move on
I think it shows that we're using the word "dictator" way too casually in that case.
However, it doesn't seem trivial to do deduplication in that case without removing relevant/necessary context.
While high torque motors got way cheaper, especially with MIT Cheetah "clones" getting easily available, they're still at least 200-500 a pop (depending on the torque needed for each articulation) from what I could find.
I might not know where to search for the real gems though. Where do you search for cheap powerful servomotors?
From Nvidia and AMD, I read sparse fp8 at 7 PFLOPs for B100 [0] vs 5.22 PFLOPs for mi325x [1]
Nvidia doesn't give the dense fp8 so that's the easiest comparison I could get.
[0] https://resources.nvidia.com/en-us-blackwell-architecture [1] https://www.amd.com/en/products/accelerators/instinct/mi300/...
We don't have to imagine far, it's slowly happening. Pytorch for ROCm is getting better and better!
Then they will have to fix the split between data-center and consumer GPU for sure. From what I understand, this is on the roadmap with the convergence of both GPU lines on the UDNA architecture.
There is a good reason why the ML community took Python as the favorite language overall.
There is a recent paper from Meta that propose a way to train a model to backtrack its generation to improve generation alignment [0].
RNN are constantly updating and overwriting their memory. It means they need to be able to predict what is going to be useful in order to store it for later.
This is a massive advantage for Transformers in interactive use cases like in ChatGPT. You give it context and ask questions in multiple turns. Which part of the context was important for a given question only becomes known later in the token sequence.
To be more precise, I should say it's an advantage of Attention-based models, because there are also hybrid models successfully mixing both approaches, like Jamba.
That would be interesting to know if his solution could match the 4k$ in term of usability or if there is some issue like refreshing rate that make the piezo based system necessary for a good user experience.
We have "garde à vue" and "détention provisoire":
- "Garde a vue" is similar to being in police custody and it's limited to 24-48 hours normally, but there is longer duration for specific crime, the maximum being 6 days for terrorism investigations [0].
- Then, a judge can decide that the defendant should stay in "détention provisoire" before a trial. It doesn't have a duration limit but should be motivated and can be re-examined multiple time [1].
0: https://fr.wikipedia.org/wiki/Garde_%C3%A0_vue_en_droit_fran... 1: https://fr.wikipedia.org/wiki/D%C3%A9tention_provisoire_en_F...
Without any quantization our current price is 30cts ingest and 50cts output per million tokens. [1]
Hard coded dialogs often feel very unnatural and limiting. I can see why people want to explore LLM to try to make new experiences possible.
I can see it becoming a new dimension of game design, open vs closed dialogs, like there is currently open vs closed world. And as in the open vs closed world, they will probably coexist instead of one type replacing the other.
I don't think it should be ignored, especially when the power consumption is similar.
They should probably show separately the throughput per completion as the tensor parallelism is often used for that purpose in addition to the doubling the VRAM.
I did an old experiment on a scrollable whiteboard with replay that I built after watching a khan academy style video and wanting to scroll to back to a formula without pausing the audio. This makes me want to dig it back ^^
But I'm very hopeful. They're slowly chipping away at Nvidia massive lead on the software front. And now everything is starting to align through the efforts of AMD but also the work from the Pytorch team to make it easier to build new backends.
We all benefit in the end if they manage to get toe to toe with Nvidia.
But it's still early and I guess that they kept just enough mystery to kickstart the conversation around this new model!
Unsupervised is a confusing term as there is always an underlying loss being optimized and working as a supervision signal, even for good old kmeans. But generative models are generally considered to be part of unsupervised methods.
They are sending messages of unity, not fealty to Sam Altman.
Working together brought them to the top of the world both in term of research and product. This is a dream come true.
What would surprise me is that they wouldn't do everything they can to avoid breaking their team.
The board tried to get a merger with Anthropic, probably the second best team in the world for this type of work.
They don't seem to think they can "just re-hire".