But I too wouldn't use this. X is playing fast and loose with ... everything, so having a business rely on their product seems risky.
With LLMs (and AGI) it's really that simple: the company with the best model wins regardless of all else.
It's going to be a combination of price, performance, quality, reliability, availability etc.
And since the prompts need to be optimised for each model there is a degree of vendor lock-in.
I'm talking about the next huge step forward that only 1 company will achieve, because it simply has the most GPUs (in limited supply) + energy source first and keeps that advantage.
At some point this becomes a run-away self amplifying differentiator and it will make that company win regardless of all else.
My money is on xAI in 2025.
PS: the only reason prompts need to be optimized for each model is a symptom of models simply not being that good yet. This need will vanish in the near future as you get way better models. A recent hint of what I mean: mid-journey needed very elaborate prompt (and even loras) to get what you want. In flux that prompt can be much shorter (without loras) and it still gets closer to what you want. Same will happen with LLMs. Another example: with ChatGPT 4 you need to literally beg a model to only return what you ask for (for example JSON) or put it in a certain mode (JSON mode), in Claude Sonnet 3.5 it will simply just listen to what you ask for. So again: that's not "because every model needs model-specific fine-tuning" that's because previous models where simply not as good.
Sometimes having a fast enough model at a low enough price makes you the obvious choice e.g. I know Claude is better than gpt-4o-mini but I use the latter for a lot more data processing because it's significantly cheaper and faster and the gains I'd get out of Claude seem somewhat marginal for my use case
Best at product / market fit. And that space is very very wide. Does the GenAI serve as a feature in a larger product (like realtime “reasoning” on X or in Apple’s case in iOS)? Is it a standalone product that general public or enterprises use? Does it play in a niche area? Etc.
https://www.theverge.com/2024/8/12/24219121/donald-trump-elo...
If that's true it's not exactly the sort of behaviour you want from an API you're depending on.
Evidence for DDOS:
- Elon said so
- the event in question very clearly had huge technical issues
Evidence Against DDOS:
- Elon said so
- People who worked at Twitter said it was bullshit
- every other spaces event that was run at the same time was unaffected.
- no other part of the website was impacted in any way whatsoever.
Aren’t these last two an argument FOR a ddos attack? It seems reasonable to assume we’re there a ddos attack at that time it would be against the Elon/Trump stream explicitly.
So no I think it was just a straight up technical failure on their end.
No, we have no idea from The Verge article whether the sources are even qualified to make such statements or if the statements are even true. In fact on the basis of the 99 percent speculative quote we can disregard the source quotes altogether. I'll say this, I work on far less significant software than X and we get DDOSed all the time.
> every other spaces event that was run at the same time was unaffected.
That's not true, I wasn't even able to load my feed during the initial part of the stream.
You baselessly accuse journalists of straight up making things up and then go on to give some anecdotal evidence that conveniently nobody can disprove.
The Register has found no evidence of a denial of service attack directed at X. Check Point Software's live cyber threat map does not record unusual levels of activity at the time of writing. NetScout's real-time DDoS map recorded only small attacks on the US.
If a DDoS was indeed the reason for the delayed start of the event, it appears not to have impacted the rest of X's operations – there were plenty of posts commenting on the problems with the Space occupied by the interview. And Musk was tweeting from the very network said to be under attack.
Elon Musk claims live Trump interview on X derailed by DDoS https://www.theregister.com/2024/08/13/trump_musk_livestream...They also threw shade on the numbers:
The interview commenced some 40 minutes after its advertised time. Live audience statistics reported 1.1 to 1.3 million attendees during the portions of the event The Register observed – although during the stream Trump claimed that the event had an audience of 60 million or more, exceeding targets of 25 million.And unfortunately both of these men are known for bullshitting more than anything else and have been now for a long time.
Was Trump a fool to count the people that took one look and changed channels, or a knowledgable and deliberate deceiver?
The Verge has no political bias, has a good reputation and thus deserve the benefit of the doubt.
If they were actually neutral, they'd phrase it more like: with technical difficulties.
We know nothing about the sources, and writers are not above making stuff up. I could just as easily spin it on them: there's a 99 percent chance they made up the sources.
From a black box perspective, LLMs are pretty simple, you put text or images in, (possibly structured) text comes out, maybe with some tool invocations.
If you use a good library for this, like Python's litellm for example, all it takes is changing one string in your code or config, as the library exposes most APIs of most providers under a simple, uniform interface.
You might need to modify your prompt and run some evals on whatever task your app is solving, but even large companies regularly deprecate old models and introduce vastly better ones, so you should have a pipeline for that anyway.
These models have very little "stickiness" or lock-in. If your app is a Twitter client and is built around the Twitter API, turning it into a Mastodon client built around the Mastodon API would take a lot of work. If your app uses Grok and is designed properly, switching over to a different model is so simple that it might be worth doing for half an hour during an outage.
Trying it isn't exactly locking you into anything