- This OpenAI outage[1] we notified 4 minutes before they acknowledged.
- The last AWS outage[2], we notified 28 minutes before they acknowledged
- There is def an Azure outage[3] now yet they have still not updated their status page. We notified 35 minutes ago.
1. https://statusgator.com/services/openai
2. https://statusgator.com/blog/amazon-cognito-outage-december-...
Does Azure have any other state? :-)
PS. That, and "check engine light management", as a concept ...>Impact Statement: Starting at 18:44 UTC on 26 Dec 2024, you have been identified as a customer who was impacted by a power incident in South Central US and may experience a degraded experience.
>Current Status: There was a power incident in the South Central US AZ03 which affected multiple services. We have applied mitigation and are actively validating recovery to the impacted services. Further updates will be provided in 60 minutes, or sooner as events warrant.
The times are the same for OpenAI - first notice from 11:00 PST (19:00 UTC)
They are lucky they were able to walk away, but the facility was dark till someone could get in there and give the power equipment the green light.
derecho storms hammered the area and killed power. external power lines in failed, and the ATS hung or died when switching to the N+1 diesel generators.
since it never got switched to diesel, the UPS systems kept things going for the standard interval (e.g. ~3-5 minutes) and then ran out of power, and then everything went down. AWS died and IIRC it took a lot of stuff with it, most notably reddit, etc.
Tl;Dr fire in one data center hall was put out with water, water leaked into other hall's power generator and battery area. Turns out loads of water and power generation equipment don't mix well, and servers don't like sitting in puddles of water.
Another one was where the whole building burnt down rather than "just" the power equipment. https://www.datacenterdynamics.com/en/news/fire-destroys-ovh...
That was an epic mess. Especially considering they positioned themselves as the most technically sophisticated colo in the region. So, yeah, it happens.
What an eloquent way to say "our service isn't working".
It’s always the same abstract invisible hand that just keeps affecting everyone! Scott Alexander’s Moloch perhaps :)
I almost wish I hadn't read that blog post because now I see Moloch and and his invisible handprints everywhere.
Anyway here it is: https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
1. In normal and honest language you state things you have done, and their consequences.
2. Passive voice attempts to deflect blame by only stating the consequences as if they just magically occurred through divine intervention.
3. In the next stage they don't even acknowledge the consequences, and instead place the entire issue inside your experience of the facts.
We don't know what exactly caused the power issue and they might not have had a root cause at the time either. Let's assume that their power redundancy equipment failed, say, due to insufficient maintenance. This is not an active action, it's a passive one (they didn't do their maintenance duties properly and now it blew). So there is nothing to say for point #1 and #2.
There's also the part where they say that the customers they identified as impacted may be experiencing a service degradation. This may sound pedantic, but I think it is not an entirely unreasonable phrasing. Maybe my business isn't actively relying on the resources I have deployed in that datacenter. How would they know (#3)? Should I clean those resources up? Possibly. Depends on my access patterns and other considerations.
It reads like face (and ass) saving legal esque language. But there's a reason face and ass saving legalese sounds like it does.
1. Larger Context Window (128K)
With Pro-Mode, I can paste an entire firmware file—thousands of lines of code—along with detailed hardware references, and the model can actually process it. This isn’t possible with the smaller context window on the $20 tier. On the Pro plan, I’ve pasted like 30+ pages of MCU datasheet information plus multiple header files in a single go. The model is then reasonably capable to provide accurate, bit-twiddled code, many times on the first try. Is it always working on the first go? Sure sometimes, but often there's still debugging, and I don't expect people that haven't actually tried to do it before without AI could do it effectively. However, I can do a code diff using tools like beyond compare (necessary for this workflow) to find bugs and/or explain what happened to pro-mode perhaps with some top level nudge for a strategy to fix it, and generally 2-3 tries later we've made progress.
2. Deeper understanding, real solutions
When I describe a complex hardware/software setup—like the power optimization for the product which is a LiPo-rechargeable fan/flashlight, the Pro-Mode model can understand the entire system better and synthesize troubleshooting approaches into a near-finished solution, with 95–100% usable results. By contrast, the non-pro plan can give good suggestions in smaller chunks, but it can’t grasp the entire system context due to its limited memory.
3. Practical Engineering Impact
I’m working on essentially the fourth generation of a LiPo-battery hardware product. Since upgrading, the Pro-Mode model helped us pinpoint power issues and cut standby battery drain from 20 days to over a year. Like, this week it guided me to discover a stealth 800 µA draw from the fan itself when the device was supposed to be in deep sleep. We were consuming ~1000 µA of power when it should be about ~200 µA. Finally, discovered the fan issue and achieved 190 µA without it in the system, so now we have a move forward to add a load switch so the MCU can isolate it from the system before it sleeps. Bingo we just went from a dead battery in ~70 days (we'd already cut it from 20 days to 70 days with firmware changes alone) to now it should take about 1 year for it to drain. This is the difference between end users having zero charge when the open the box to being able to use the product immediately.
4. Value vs. Traditional Consulting
I’ve hired $20K short-term consultants who didn’t deliver half the insights I’ve gotten in a single subscription month. It might sound like an overstatement, but Pro-Mode has been the best $200 I’ve spent—especially given how quickly it has helped resolve engineering hurdles.
In short: Probably the biggest advantage is the vastly higher context window, which allows the model to handle large, interrelated hardware/software details all at once. If you work on complex firmware or detailed electronics designs, Pro-Mode can feel like an invaluable engineering partner.
I've been thinking of using another service for bigger contexts. But this may not make sense then.
I only use GPT via the API anyway so it's pay as you go. But as far as I remember there's limits there too, only big spenders get access to the top shelf stuff. I only spend a couple dollars a month because I use my llama server most of the time. It's not as good as ChatGPT obviously but it's mine and doesn't leak my conversations.
It wasn’t an accusation (I don’t think it actually matters in the end), so much as to understand why do it — in a post about ChatGPT usage, it helps understand context: if OP values using it for stuff I wouldn’t value using it for, for example, then it will change the variables.
(also sucks for non-native speakers or even speakers of other dialects, like delve - apparently it is a common word for Nigerian English)
- With a statically typed language and a compiler, it's quite easy to automatically assemble a meaningful context with 1-2 nested calls of recursive 'Go To Definition' and including the source from that. You can use various heuristics (either from compile time or runtime). It's quite easy to implement, we've done this for older, non-AI stuff a while ago, for trying to figure out the impact of code changes. If you have a compiler running, I'm pretty sure you could do this in a couple days. This makes the long context not super necessary.
- In my experience, long context models can't really use their contexts that well. They were trained to do well on 'needle-in-the-haystack' benchmarks, that is, to retrieve information that might be scattered anywhere in the context, which might be good enough here, but asking complex questions that require the understanding the entire context trips the models up. I tried some fiction writing with long context models, and I often found that they forgot things and messed up cause and effect. Not sure if this applies to current state of the art models, but I bet it does, since sequencing and theory-of-mind (it's established in the story that Alice is the killer, but Bob doesn't know that at that point, models often mess this up and assume he does) are still active research topics, and current models kinda suck at it.
For writing fiction, I found that the sliding window of short-context models was much better, with long-context ones often bringing up irrelevant details, and ignoring newer, more relevant ones.
Again, not sure how this affects the business of writing firmware code, but limitations do exist.
Was this context across a single datasheet or was Pro-Mode able to deduce from how multiple parts were connected/programmed? Did it identify the problem, or just suggest where to look?
No worries that you run put of prompts for o1. which allows for more experimentation and creativity.
Not sure if it’d work for your workflow, but it’s really nice if it does.
That’s setting aside when they hosted the status red/yellow/green indicator images on s3, so the surest sign that s3 was having issues was that the status indicators didn’t load at all.
“Boss, can I change this indicator from green to red?”
“How much will that cost us?”
“About a million dollars an hour.”
“No.”
Not your lost profit!
Everyone assumes the latter -- that they'll be compensated -- but in reality they'll be refunded $3.27 for the storage account that's got a few gigabytes of ultra-critical build scripts and static web content, without which their multi-million dollar business stops dead.
"We can refund you with loose change, or a gift card for a coffee."
That would be something entirely different -- buying a form of insurance, basically, that would be expensive.
SLA's aren't meant to make your whole as a business, generally speaking. They're meant to incentivize the provider to take uptime really seriously, so that downtime eats some of their profit. Which means you can take their expected uptime estimates as a decent ballpark.
Like are there contracts bound to the uptime and at the same time bound to them self reporting it? That would seem strange.
"Sir, are you absolutely sure, it does mean changing the bulb."
edit: downdetector confirms:
Stackoverflow will have duplicates, approximates and what not and sometimes that works. But at other times, you hunt for a half hour before you figure it out.
You can throw the problem at ChatGPT, it may go wrong but your course correct it with simple instructions and slowly but steadily you move towards your goal with minimal noise of the irrelevant discussions.
What stands between the solution and you then is your ability to figure out when it is hallucinating and guide it to the right direction. But as a solution developer you should have that insight anyways
I do wonder if we're the last generation that will be able to effectively do such "course correct" operations -- feels like a good chunk of the next generation of programmers will be bootstrapped using such LLMs, so their ability to "have that insight" will be lacking, or be very challenging to bootstrap. As analogy, do you find yourself having to "course correct" the compiler very often?
I asked it a simple non programming question. My last paycheck was December 20, 2024. I get paid biweekly. In which year will I get paid 27 times. It got it wrong ... very articulately.
I run into this every single day.
To do this reliably, prepend your request to invoke a tool like OpenAI's Code Interpreter (e.g. "Code the answer to this: My last paycheck was December 20, 2024. I get paid biweekly. In which year will I get paid 27 times.") to get the correct response of 2027.
> I get paycheck every 2 weeks. Last paycheck was December 20, 2024. Which year will I have 27 paychecks?
I sent it again and it bombed again. It seems your prompt and my prompt are quite similar, but I realize the suggestion (or direction) to it to code.
OpenAI's Code Interpreter was the first thing I saw which helped me understand that we really won't understand the impact of LLMs until they're released from their sandbox. This is why I find Apple's efforts to create standard interfaces to iOS/macOS apps and their data via App Intents so interesting. Even if Apple's on-device models can't beat competitors' cloud models, I think there's magic in that union of models and tools.
And by no means I am giving up on stackoverflow, it is just another tool, but its primacy may be in doubt. Just like for the last couple of years I would search for some information by pointing google to reddit, I will now have a mental map of when to go to chatter, when to go to SO, and when to go to reddit.
These services are accumulative and the chance that any one of of those services being down at any given time is increasing every time we add a new online dependency.
Can't wait for this postmortem to be released.
For logical reasoning of course chain-of-thought models like the O1 family are better.
He runs highly competent firms.
Option 1.
Fix cars not to crash
Option 2.
Buy president and have reporting agencies closed.
Five times in two weeks I've asked OpenAI some basic factual information and it didn't get even close on any of them.
Ask it for reasoning, u habe to bring the facts.
Also he's a horrible human for many reasons and I'd prefer not to support him if I can avoid it. (You know it's bad when Google is ethical in comparison.)
Nor building AI tools for the pentagon to bomb Yemeni weddings more efficiently: https://en.wikipedia.org/wiki/Project_Maven
Didn't you read the part where he wrote "(You know it's bad when Google is ethical in comparison.)"?
Do you know of anyone at Google who overpaid billions of dollars for a popular widely used communication platform just to use it to publicly humiliate, deadname, misgender, and bully their own child in front of millions of people?
And do you actually think he has the self control not to inject his own prejudices into the LLM he made for that very purpose? Of course it's ingested the sewage of content from Twitter, which is FULL of his own jabs against people he doesn't like, including his own child. He gives his own tweets extra weight, so don't you think he does the same with training Grok?
https://www.nbcnews.com/tech/tech-news/elon-musk-transgender...
>Elon Musk's transgender daughter, in first interview, says he berated her for being queer as a child. In an exclusive interview, Vivian Jenna Wilson said her father’s recent statements, including that she is “not a girl,” inspired her to speak out: “I’m not just gonna let that slide.”
It got him the adviser role of the president, which in turn might save and make him billions.
But his main motivation might have been indeed to fight "the woke terror".
Maybe he'll give Trump some advice on putting Don Jr. and Eric in their place.
Disclosure: I work at Google.
I ask cos I use it very liberally and haven't had any issues that have made me consider adding a key, except when I made it read my whole codebase on every request
Plus mistral-nemo in particular has a large context window, so you can cook up some shell scripts to throw a bunch of context into the buffer before your question. One I use a lot takes the name of a manpage and a question about it, then the LLM has the whole manpage to reference.