Ms. Murati wrote a private memo to Mr. Altman raising questions about his management and also shared her concerns with the board. That move helped to propel the board’s decision to force him out.
https://www.nytimes.com/2024/03/07/technology/openai-executi...
It should be no surprise if Sam Altman wants executives who opposed his leadership, like Mira and Ilya, out of the company. When you're firing a high-level executive in a polite way, it's common to let them announce their own departure and frame it the way they want.
And John Schulman, and Peter Deng are out already. Yet the company is still shipping, like no other. Recent multimodal integrations and benchmarks of o1 are outstanding.
It's a very relevant fact that Greg Brockman recently left on his own volition.
Greg was aligned with Sam during the coup. So, the fact that Greg left lends more credence to the idea that Murati is leaving on her own volition.
Except that isn’t true. He has not resigned from OpenAI. He’s on extended leave until the end of the year.
That could become an official resignation later, and I agree that that seems more likely than not. But stating that he’s left for good as of right now is misleading.
>> Yet the company is still shipping, like no other.
this is factually wrong. Just today Meta (which I despise) shipped more than openAI in a long time.
If executives / high level architects / researchers are working on this quarter's features something is very wrong. The higher you get the more ahead you need to be working, C-level departures should only have an impact about a year down the line, at a company of this size.
Meta, Anthropic, Google, and others all are shipping state of the art models.
I'm not trying to be dismissive of OpenAI's work, but they are absolutely not the only company shipping very large foundation models.
Big fan of Greg, and I think the motivation behind AGI is sound here. Even what we have now is a fantastic tool, if people decide to use it.
Really? Anthropic seems to be popping off right now.
Kagi isn’t exactly in the AI space, but they ship features pretty frequently.
OpenAI is shipping incremental improvements to its chatgpt product.
- Prompt Cache, 90% savings on large system prompts for 5 mins of calls. This is amazing
- Contexual RAG, while not ground breaking idea, is important thinking and method for better vector retrieval
I don't see it for OpenAI, I do see it for the competition. They have shipped incremental improvements, however, they are watering down their current models (my guess is they are trying to save on compute?). Copilot has turned into garbage and for coding related stuff, Claude is now better than gpt-4.
Honestly, their outlook is bleak.
I understand why they are doing it, but honestly if they cancel GPT-4, many people will just cancel their subscription.
You also give them some distance in time from the drama so the two appear unconnected under cursory inspection.
Now that she's seen exactly what prompted the previous board to fire Altman, she fires herself because she understands their decision now.
Just to clear one thing up, the designated function of a board of directors is to appoint or replace the executive of an organisation, and openAI in particular is structured such that the non-profit part of the organisation controls the LLC.
The coup was the executive, together with the investors, effectively turning that on its head by force.
Also highly cynical.
Some folks are professional and mature. In the best organisations, the management team sets the highest possible standard, in terms of tone and culture. If done well, this tends to trickle down to all areas of the organization.
Another speculation would be that she's resigning for complicated reasons which are personal. I've had to do the same in my past. The real pro's give the benefit of the doubt.
Please no speculative pieces, rumor nor hearsay.
CEOs get fired all the time and company puts out a statement.
I've never seen "we won't tell you why we fired our CEO" anywhere.
now he is back making totally ridiculous statments like 'AI is going to solve all of physics' or that 'AI is going to clone my brain by 2027'
This is a strange company.
Because the old guard wanted it to remain a cliquey non-profit filled to the brim with EA, AI Alignment, and OpenPhilanthropy types, but the current OpenAI is now an enterprise company.
This is just Sam Altman cleaning house after the attempted corporate coup a year ago.
The board’s only reason to exist is effectively to fire the CEO.
There are some juicy rumors about what actually happened too. much more belivable lol .
Neither did she though... To my knowledge.
Can you provide any evidence that she tried to do that? I would ask that it be non-speculative in nature please.
This piece is built on conjecture from a source whose identify is withheld. The sources version of events is openly refuted by the parties in question. Offering it as evidence that Mirati intentionally made political moves in order to get Altman ousted is an indefensible position.
'Mr. Sutskever’s lawyer, Alex Weingarten, said claims that he had approached the board were “categorically false.”'
'Marc H. Axelbaum, a lawyer for Ms. Murati, said in a statement: “The claims that she approached the board in an effort to get Mr. Altman fired last year or supported the board’s actions are flat wrong. She was perplexed at the board’s decision then, but is not surprised that some former board members are now attempting to shift the blame to her.” In a message to OpenAI employees after publication of this article, Ms. Murati said she and Mr. Altman “have a strong and productive partnership and I have not been shy about sharing feedback with him directly.”
She added that she did not reach out to the board but “when individual board members reached out directly to me for feedback about Sam, I provided it — all feedback Sam already knew,” and that did not mean she was “responsible for or supported the old board’s actions.”'
This part of NYT piece is supported by evidence:
'Ms. Murati wrote a private memo to Mr. Altman raising questions about his management and also shared her concerns with the board. That move helped to propel the board’s decision to force him out.'
INTENT matters. Mirati says the board asked for her concerns about Altmans. She provided it and had already brought it to Altmans attention... in writing. Her actions demonstrate transparency and professionalism.
Organizational performance metrics.
Frequency of scientific breakthroughs.
Frequency and quality of product updates.
History of consistently setting the state of the art in artificial intelligence.
Demonstrated ability to attract world class talent.
Released the fastest growing software product in the history of humanity.
I could write paragraphs...
Why the rain clouds?
Instead, you get to see grey area after grey area.
The fact that it has a structure that subordinates it to the board of a non-profit would be only tangential to the interests involved even if that was meaningful and not just rhe lingering vestige of the (arguably, deceptive) founding that the combined organization was working on getting rid of.
Not to dunk on Mira Murati, because this note is pretty cookie cutter, but it exemplifies this perfectly. It says nothing about her motivations for resigning. It bends over backwards to kiss the asses of the people she's leaving behind. It could ultimately be condensed into two words: "I've resigned."
Never spook the horses. Never show the team, or the public, what's going on behind the curtain.. or even that there is anything going on. At all time present the appearance of a swan gliding serenely across a lake.
Because if you show humanity, those other humans might cotton on to the fact that you're not much different to them, and have done little to earn or justify your position of authority.
And that wouldn't do at all.
Probably why AI sludge is so well suited to this particular cultural moment.
See flirting as a more basic example.
Edit: Shoot, look at the general level of distrust that the populous puts in politicians.
Mira's latest one liner tweet 'OpenAI is nothing without it's people" speaks volumes.
Saying I was wrong should not be this complicated, or saying we failed.
I do however agree that there is nothing to be gained and everything to be risked. So why do it.
You are asking the question, why are politicians not honest?
Altmans quote was that "it's possible that we will have superintelligence in a few thousand days", which sounds a lot more optimistic on the surface than it actually is. A few thousand days could be interpreted as 10 years or more, and by adding the "possibly" qualifier he didn't even really commit to that prediction.
It's hype with no substance, but vaguely gesturing that something earth-shattering is coming does serve to convince investors to keep dumping endless $billions into his unprofitable company, without risking the reputational damage of missing a deadline since he never actually gave one. Just keep signing those 9 digit checks and we'll totally build AGI... eventually. Honest.
I think he was referring to ASI, not AGI.
Old-school AI was already specialised. Nobody can agree what "sentient" is, and if sentience includes a capacity to feel emotions/qualia etc. then we'd only willingly choose that over non-sentient for brain uploading not "mere" assistants.
By all the standards I had growing up, ChatGPT is already AGI. It's almost certainly not as economically transformative as it needs to be to meet OpenAI's stated definition.
OTOH that may be due to limited availability rather than limited quality: if all the 20 USD/month for Plus gets spent on electricity to run the servers, at $0.10/kWh, that's about 274 W average consumption. Scaled up to the world population, that's approximately the entire global electricity supply. Which is kinda why there's also all the stories about AI data centres getting dedicated power plants.
We made a thing that exhibits the emergent property of intelligence. A level of intelligence that trades blows with humans. The fact that our brains do lots of other things to make us into self-contained autonomous beings is cool and maybe answers some questions about what being sentient means but memory and self-learning aren't the same thing as intelligence.
I think it's cool that we got there before simulating an already existing brain and that intelligence can exist separate from consciousness.
Has been this way since calculation machines were invented hundreds of years ago.
A range I'd agree with; for me, "pessimism" is the shortest part of that range, but even then you have to be very confident the specific metaphorical horse you're betting on is going to be both victorious in its own right and not, because there's no suitable existing metaphor, secretly an ICBM wearing a patomime costume.
2 (or even 3) you use "a couple"
A few is almost always > 3 and one could argue that upper limit 15
So, 10 years to 50 years
But the mere fact you say 15 is arguable does indeed broaden the range, just as me saying 1 broadens it in the opposite extent.
15 is too high to be a "few" except in contexts of a few out of tens of thousands of items.
Realistically I interpret this as 3-7 thousands of days (8 to 19 years), which is largely consensus prediction range anyway.
That said, I think people are possibly overanalysing this very vague barely-even-a-claim just a little. Realistically, when a tech company makes a vague claim about what'll happen in 10 years, that should be given precisely zero weight; based on historical precedent you might as well ask a magic 8-ball.
But really. o1 has been very whelming, nothing like the step up from 3.5 to 4. Still prefer sonnet3.5 and opus.
There, that's my conspiracy theory quota for 2024 in one comment.
Sure, a few thousand days and a few trillion $ away. We'll also have full self driving next month. This is just like the fusion is the energy of the future joke: it's 30 years away and it will always be.
[1] https://slate.com/technology/2019/02/openai-gpt2-text-genera...
It makes us feel understood in the same ways John Edward used to in daytime tv, its all about how language makes us feel
true AGI...unfortunately we're not even close
Or even human intelligence
"Your argument is just a reductive rhetorical strategy."
"a probabilistic syllable generator is not intelligence, it does not understand us, it cannot reason" is a strong statement and I highly doubt it's backed by any sort of substance other than "feelz".
For example, even the dumbest dog has a memory, a strikingly advanced concept model of the world [1], a persistent state beyond the last conversation history, and an ability to reason (that doesn't require re-running the same conversation sixteen bajillion times in a row). Transformer models do not. It's really cool that they can input and barf out realistic-sounding text, but let's keep in mind the obvious truths about what they are doing.
[1] "I like food. Something that smells like food is in the square thing on the floor. Maybe if I tip it over food will come out, and I will find food. Oh no, the person looked at me strangely when I got close to the square thing! I am in trouble! I will have to do it when they're not looking."
Lets assume the dog visual systems run at 60 frames per second. If it takes 1 second to flip a bowl of food over then that's 60 datapoints of cause-effect data that the dog's brain learned from.
Assuming it's the same for humans, lets say I go on a trip to the grocery store for 1 hour. That's 216,000 data points from one trip. Not to mention auditory data, touch, smell, and even taste.
> ability to reason [...] Transformer models do not
Can you tell me what reasoning is? Why can't transformers reason? Note I said transformers not llm's. You could make a reasonable (hah) case that current LLMs cannot reason (or at least very well) but why are transformers as an architecture doomed?
What about chain of thought? Some have made the claim that chain of thought adds recurrence to transformer models. That's a pretty big shift, but you've already decided transformers are a dead end so no chance of that making a difference right?
There's a huge amount of circuitry between the input and the output of the model. How do you know what it does or doesn't do?
Humans brains "just" output the next couple milliseconds of muscle activation, given sensory input and internal state.
Edit: Interestingly, this is getting downvotes even though 1) my last sentence is a precise and accurate statement of the state of the art in neuroscience and 2) it is completely isomorphic to what the parent post presented as an argument against current models being AGI.
To clarify, I don't believe we're very close to AGI, but parent's argument is just confused.
Apologies for the wording but I think you got it and the point stands.
I'm not a native speaker and mostly use English in a professional science related setting, that's why I sound like that sometimes.
isomorphic - being of identical or similar form, shape, or structure (m-w). Here metaphorically applied to the structure of an argument.
Yeah - but it's just a stack of transformer layers. No looping, no memory, no self-modification (learning). Also, no magic.
Neuroscience hasn't found the magic dust in our brains yet, either. ;)
There is no real runtime learning - certainly no weight updates. The weights are all derived from pre-training, and so the runtime model just represents a frozen chunk of learning. Maybe you are thinking of "in-context learning", which doesn't update the weights, but is rather the ability of the model to use whatever is in the context, including having that "reinforced" by repetition. This is all a poor substitute for what an animal does - continuously learning from experience and exploration.
The "magic dust" in our brains, relative to LLMs, is just a more advanced and structure architecture, and operational dynamics. e.g. We've got the thalamo-cortical loop, massive amounts of top-down feedback for incremental learning from prediction failure, working memory, innate drives such as curiosity (prediction uncertainty) and boredom to drive exploration and learning, etc, etc. No magic, just architecture.
But, people in this thread are making philosophically very poor points about why that is supposedly so.
It's not "just" sequence prediction, because sequence prediction is the very essence of what the human brain does.
Your points on learning and memory are similarly weak word play. Memory means holding some quantity constant over time in the internal state of a model. Learning means being able to update those quantities. LLMs obviously do both.
You're probably going to be thinking of all sorts of obvious ways in which LLMs and humans are different.
But no one's claiming there's an artificial human. What does exist is increasingly powerful data processing software that progressively encroaches on domains previously thought to be that of humans only.
And there may be all sorts of limitations to that, but those (sequences, learning, memory) aren't them.
Agree wrt the brain.
Sure, LLMs are also sequence predictors, and this is a large part of why they appear intelligent (intelligence = learning + prediction). The other part is that they are trained to mimic their training data, which came from a system of greater intelligence than their own, so by mimicking a more intelligent system they appear to be punching above their weight.
I'm not sure that "JUST sequence predictors" is so inappropriate though - sure sequence prediction is a powerful and critical capability (the core of intelligence), but that is ALL that LLMs can do, so "just" is appropriate.
Of course additionally not all sequence predictors are of equal capability, so we can't even say, "well, at least as far as being sequence predictors goes, they are equal to humans", but that's a difficult comparison to make.
> Your points on learning and memory are similarly weak word play. Memory means holding some quantity constant over time in the internal state of a model. Learning means being able to update those quantities. LLMs obviously do both.
Well, no...
1) LLMs do NOT "hold some quantity constant over time in the internal state of the model". It is a pass-thru architecture with zero internal storage. When each token is generated it is appended to the input, and the updated input sequence is fed into the model and everything is calculated from scratch (other than the KV cache optimization). The model appears to be have internal memory due to the coherence of the sequence of tokens it is outputting, but in reality everything is recalculated from scratch, and the coherence is due to the fact that adding one token to the end of a sequence doesn't change the meaning of the sequence by much, and most of what is recalculated will therefore be the same as before.
2) If the model has learnt something, then it should have remembered it from one use to another, but LLMs don't do this. Once the context is gone and the user starts a new conversation/session, then all memory of the prior session is gone - the model has NOT updated itself to remember anything about what happened previously. If this was an employee (an AI coder, perhaps) then it would be perpetual groundhog day. Every day it came to work it'd be repeating the same mistakes it made the day before, and would have forgotten everything you might have taught it. This is not my definition of learning, and more to the point the lack of such incremental permanent learning is what'll make LLMs useless for very many jobs. It's not an easy fix, which is why we're stuck with massively expensive infrequent retrainings from scratch rather than incremental learning.
This is also true of those with advanced Alzheimer's disease. Are they not conscious as well? If we believe they are conscious then memory and learning must not be essential ingredients.
I thought we're talking about intelligence, not consciousness, and limitations of the LLM/transformer architecture that limit their intelligence compared to humans.
In fact LLMs are not only architecturally limited, but they also give the impression of being far more intelligent than they actually are due to mimicking training sources that are more intelligent than the LLM itself is.
If you want to bring consciousness into the discussion, then that is basically just the brain modelling itself and the subjective experience that gives rise to. I expect it arose due to evolutionary adaptive benefit - part of being a better predictor (i.e. more intelligent) is being better able to model your own behavior and experiences, but that's not a must-have for intelligence.
a pretrained transformer in the limit does not converge on any collective or consensus state in that sense and in fact, pre-training actually punishes this. It learns to predict the words of Feynman as readily as the dumbass across the street.
When i say that GPT does not mimic, i mean that the training objective literally optimizes for beyond that.
Consider <Hash, plaintext> pairs. You can't predict this without cracking the hash algorithm, but you could easily fool a GAN's discriminator(one that has learnt to compute hash functions) just by generating typical instances.
# Consider that some of the text on the Internet isn't humans casually chatting or extemporaneous speech. It's the results section of a science paper. It's news stories that say what happened on a particular day. It's text that people crafted over hours or days.
These models can engage in multistep logical reasoning, solve complex problems, and generate novel ideas - going far beyond simply predicting the next syllable. They can follow intricate chains of thought and arrive at non-obvious conclusions. And OpenAI has now showed us that fine-tuning a model specifically to plan step by step dramatically improves its ability to solve problems that were previously the domain of human experts.
Although there is no definitive evidence that state-of-the-art language models have a comprehensive "world model" in the way humans do, several studies and observations suggest that large language models (LLMs) may possess some elements or precursors of a world model.
For example, Tegmark and Gurnee [1] found that LLMs learn linear representations of space and time across multiple scales. These representations appear to be robust to prompting variations and unified across different entity types. This suggests that modern LLMs may learn rich spatiotemporal representations of the real world, which could be considered basic ingredients of a world model.
And even if we look at much smaller models like Stable Diffusion XL, it's clear that they encode a rich understanding of optics [2] within just a few billion parameters (3.5 billion to be precise). Generative video models like OpenAI's Sora clearly have a world model as they are able to simulate gravity, collisions between objects, and other concepts necessary to render a coherent scene.
As for AGI, the consensus on Metaculus is that it will arrive in 2023. But consider that before GPT-4 arrived, the consensus was that full AGI was not coming until 2041 [3]. The consensus for the arrival date of "weakly general" AGI is 2027 [4] (i.e AGI that doesn't have a robotic physical world component). The best tool for achieving AGI is the transformer and its derivatives; its scaling keeps going with no end in sight.
Citations:
[1] https://paperswithcode.com/paper/language-models-represent-s...
[2] https://www.reddit.com/r/StableDiffusion/comments/15he3f4/el...
[3] https://www.metaculus.com/questions/5121/date-of-artificial-...
[4] https://www.metaculus.com/questions/3479/date-weakly-general...
I won't expand on the rest, but this is simply nonsensical.
The fact that Sora generates output that matches its training data doesn't show that it has a concept of gravity, collision between object, or anything else. It has a "world model" the same way a photocopier has a "document model".
The ability of video models to generate novel video consistent with physical reality shows that they have extracted important invariants - physical law - out of the data.
It's probably better not to muddle the discussion with ill defined terms such as "intelligence" or "understanding".
I have my own beef with the AGI is nigh crowd, but this criticism amounts to word play.
Are you saying that it is not possible to learn about dynamics in a higher dimensional space from a lower dimensional projection? This is clearly not true in general.
E.g., video models learn that even though they're only ever seeing and outputting 2d data, objects have different sides in a fashio that is consistent with our 3d reality.
The distinctions you (and others in this thread) are making is purely one of degree - how much generalization has been achieved, and how well - versus one of category.
Not only are we within eyesight of the end, we're more or less there. o1 isn't just scaling up parameter count 10x again and making GPT-5, because that's not really an effective approach at this point in the exponential curve of parameter count and model performance.
I agree with the broader point: I'm not sure it isn't consistent with current neuroscience that our brains aren't doing anything more than predicting next inputs in a broadly similar way, and any categorical distinction between AI and human intelligence seems quite challenging.
I disagree that we can draw a line from scaling current transformer models to AGI, however. A model that is great for communicating with people in natural language may not be the best for deep reasoning, abstraction, unified creative visions over long-form generations, motor control, planning, etc. The history of computer science is littered with simple extrapolations from existing technology that completely missed the need for a paradigm shift.
I definitely agree that AGI isn't just a matter of scaling transformers, and also as you say that they "may not be the best" for such tasks. (Vanilla transformers are extremely inefficient.) But the really important point is that transformers can do things such as abstract, reason, form world models and theories of minds, etc, to a significant degree (a much greater degree than virtually anyone would have predicted 5-10 years ago), all learnt automatically. It shows these problems are actually tractable for connectionist machine learning, without a paradigm shift as you and many others allege. That is the part I disagree with. But more breakthroughs needed.
(Translated from Chinese) > According to industry insiders, OpenAI originally actively negotiated with TSMC to build a dedicated wafer factory. However, after evaluating the development benefits, it shelved the plan to build a dedicated wafer factory. Strategically, OpenAI sought cooperation with American companies such as Broadcom and Marvell for its own ASIC chips. Development, among which OpenAI is expected to become Broadcom's top four customers.
[1] https://money.udn.com/money/story/5612/8200070 (Chinese)
Even if OpenAI doesn't build its own fab -- a wise move, if you ask me -- the investment required to develop an ASIC on the very latest node is eye watering. Most people - even people in tech - just don't have a good understanding of how "out there" semiconductor manufacturing has become. It's basically a dark art at this point.
For instance, TSMC themselves [2] don't even know at this point whether the A16 node chosen by OpenAI will require using the forthcoming High NA lithography machines from ASML. The High NA machines cost nearly twice as much as the already exceptional Extreme Ultraviolet (EUV) machines do. At close to $400M each, this is simply eye watering.
I'm sure some gurus here on HN have a more up to date idea of the picture around A16, but the fundamental news is this: If OpenAI doesn't think scaling will be needed to get to AGI, then why would they be considering spending many billions on the latest semiconductor tech?
Citations: [1] https://www.phonearena.com/news/apple-paid-twice-as-much-for... [2] https://www.asiabusinessoutlook.com/news/tsmc-to-mass-produc...
Based on capabilities alone, current LLMs demonstrate many of the capabilities practitioners ten years ago would have tossed into the AGI bucket.
What are some top capabilities (meaning inputs and outputs) you think are missing on the path between what we have now and AGI?
I think it's more productive to think about AI in terms of "effectiveness" or "capability". If you ask it, "what is the capital of France?", and it replies "Paris" - it doesn't matter whether it is intelligent or not, it is effective/capable at identifying the capital of France.
Same goes for producing an image, writing SQL code that works, automating some % of intellectual labor, giving medical advice, solving an equation, piloting a drone, building and managing a profitable company. It is capable of various things to various degrees. If these capabilities are enough to make money, create risks, change the world in some significant way - that is the part that matters.
Whether we call it "intelligence" or "probabilistically generaring syllables" is not important.
I understand the fear, but the knee jerk response “its just predicting the next token thus could never be intelligent” makes you look more like a stochastic parrot than these models are.
I think your use of the "goalposts" metaphor is telling. You see this as a team sport; you see yourself on the offensive, or the defensive, or whatever. Neither is conducive to a balanced, objective view of reality. Modern LLMs are shockingly "smart" in many ways, but if you think they're general intelligence in the same way humans have general intelligence (even disregarding agency, learning, etc.), that's a you problem.
^[1] I feel the implicit suggestion that there was some sort of broad consensus on this in the before-times is revisionism.
How is it a me problem? The idea of these models being intelligent is shared with a large number of researchers and engineers in the field. Such is clearly evident when you can ask o1 some random completely novel question about a hypothetical scenario and it gets the implication you're trying to make with it very well.
I feel that simultaneously praising their abilities while claiming that they still aren't intelligent "in the way humans are" is just obscure semantic judo meant to stake an unfalsifiable claim. There will always be somewhat of a difference between large neural networks and human brains, but the significance of the difference is a subjective opinion depending on what you're focusing on. I think it's much more important to focus on the realm of "useful, hard things that are unique to intelligent systems and their ability to understand the world" is more important than "Possesses the special kind of intelligence that only humans have".
This is a common strawman that appears in these conversations—you try to reframe my comments as if I'm claiming human intelligence runs on some kind of unfalsifiable magic that a machine could never replicate. Of course, I've suggested no such thing, nor have I suggested that AI systems aren't useful.
Attempts at autonomous AI agents are still failing spectacularly because the models don't actually have any thought or memory. Context is provided to them via prefixing the prompt with all previous prompts which obviously causes significant info loss after a few interaction loops. The level of intellectual complexity at play here is on par with nematodes in a lab (which btw still can't be digitally emulated after decades of research). This isn't a diss on all the smart people working in AI today, bc I'm not talking about the quality of any specific model available today.
LLM's do have memory and thought. I've invented a few somewhat unusual games, described it to Sonnet 3.5 and it reproduces it in code almost perfectly. Likewise its memory has been scaling. Just a couple years ago context windows were 8000 tokens maximum, now they're reaching the millions.
I feel like you're approaching all these capabilities with a myopic viewpoint, then playing semantic judo to obfuscate the nature of these increases as "not counting" since they can be vaguely mapped to something that has a negative connotation.
>A lot of people don't even consider the ability to solve problems to be a reliable indicator of intelligence
That's a very bold statement, as lots of smart people have said that the very definition of intelligence is the ability to solve problems. If fear of the effectiveness of LLM's in behaving genuinely intelligently leads you to making extreme sweeping claims on what intelligence doesn't count as, then you're forcing yourself into a smaller and smaller corner as AI SOTA capabilities predictably increase month after month.
I am not 100% sure that they are still clearly leading the technology part, but agree in all other accounts.
There is a secondary market for OpenAI stock.
It's not a public market so nobody knows how much you're making if you sell, but if you look at current valuations it must be a lot.
In that context, it would be quite hard not to leave and sell or stay and sell. What if oai loses the lead? What if open source wins? Keeping the stock seems like the actual hard thing to me and I expect to see many others leave (like early googlers or Facebook employees)
Sure it's worth more if you hang on to it, but many think "how many hundreds of M's do I actually need? Better to derisk and sell"
a) you had more money than you'll ever need in your lifetime
b) you think AI abundance is just around the corner, likely making everything cheaper
c) you realize you still only have a finite time left on this planet
d) you have non-AGI dreams of your own that you'd like to work on
e) you can get funding for anything you want, based on your name alone
Do you keep working at OpenAI?
https://en.wikipedia.org/wiki/Mira_Murati
Point me to a single credential where you feel confident of putting your money on her?
It's a pity that HN crowd doesn't go one-level deep and truly understand on first principles
This is a very strange company to say the least.
Most hiring in the foundational AI/model space is very nepotistic and biased towards people in that clique.
Also, Elon Musk used to be the primary patron for OpenAI before losing interest during the AI Winter in the late 2010s.
Doesn't she have a dual bachelors in Mathematics and Mechanical Engineering?
But my point is that she does have a technical background.
No idea because she scrubbed her linkedin profile. But afaik she didn't have "years of experience leading projects" to get a job as leadpm at tesla. That was her first job as PM.
probably not coincidence that she resigned at almost the same time the rumors about OpenAI completely removing the non-profit board are getting confirmed - https://www.reuters.com/technology/artificial-intelligence/o...
Maybe, just maybe, we reached diminishing returns with AI, for now at least.
Whether or not we're leveling out, only time will tell. That's definitely what it looks like, but it might just be a plateau.
the curse of dimensionality though...
I just tried Gemini and it was useless.
Their laughably overzealous nanny-state censorship, paired with a model so appallingly inept it would embarrass a chatbot from the 90s, makes it nothing short of highway robbery that this digital dumpster fire is permitted to masquerade as a product fit for public consumption.
The sheer gall of Google to foist this steaming pile of silicon refuse onto unsuspecting users borders on fraudulent.
Someone says "X is the model that really impressive. Y is good too."
Then someone responds "What?! I just used Z and it was terrible!"
I see this at least once in practically every AI thread
They just won’t be the hottest thing since smartphones.
I think actually the best use case for LLMs is "explainer".
When combined with RAG, it's fantastic at taking a complex corpus of information and distilling it down into more digestible summaries.
I think that RAG and RAG-based tooling around LLMs is gonna be the clear way forward for most companies with a properly constructed knowledge base but I wonder what you mean by "explainer"?.
Are you talking about asking an LLM something like "in which way did the teams working on project X deal with Y problem?" and then having it breaking it down for you? Or is there something more to it?
1. I got this medical provider that has a webapp that downloads graphql data(basically json) to the frontend and shows some of the data to the template as a result while hiding the rest. Furthermore, I see that they hide even more info after I pay the bill. I download all the data, combine it with other historical data that I have downloaded and dumped it into the LLM. It spits out interesting insights about my health history, ways in which I have been unusually charged by my insurance, and the speed at which the company operates based on all the historical data showing time between appointment and the bill adjusted for the time of year. It then formats everything into an open format that is easy for me to self host. (HTML + JS tables). Its a tiny way to wrestle back control from the company until they wise up.
2. Companies are increasingly allowing customers to receive a "backup" of all the data they have on them(Thanks EU and California). For example Burger King/Wendys allow this. What do they give you when you request data? A zip file filled with just a bunch of crud from their internal system. No worries: Dump it into the LLM and it tells you everything that the company knows about you in an easy to understand format (Bullet points in this case). You know when the company managed to track you, how much they "remember", how much money they got out of you, your behaviors, etc.
I don't understand enough about #2 to comment, but it's certainly interesting.
2. I think if the data was significantly large, the llm would alias a ton of potentially important info.
Some trials have their protocols published.
Here's an example trial: https://clinicaltrials.gov/study/NCT06613256
And here's the protocol: https://cdn.clinicaltrials.gov/large-docs/56/NCT06613256/Pro... It's actually relatively short at 33 pages. Some larger trials (especially oncology trials) can have protocols that are 200 pages long.
One of the big challenges with clinical trials is making this information more accessible to both patients (for informed consent) and the trial site staff (to avoid making mistakes, helping answer patient questions, even asking the right questions when negotiating the contract with a sponsor).
The gist of it here is exactly like you said: RAG to pull back the relevant chunks of a complex document like this and then LLM to explain and summarize the information in those chunks that makes it easier to digest. That response can be tuned to the level of the reader by adding simple phrases like "explain it to me at a high school level".
The last system, I led one team competing for the Transcelerate Shared Investigator Portal (we were one of the finalist vendors).
Little side project: https://zeeq.ai
It’s not like it’ll do it consistently.
Just a marketing stunt.
It was trained to generate text in the lean language (https://www.lean-lang.org/) which is specifically used for formal proofs.
It’s not a natural language model.
Source: https://deepmind.google/discover/blog/ai-solves-imo-problems...
Google seems to mainly be playing the game of more specialized models (AlphaGo, AlphaProof) with general training methods (AlphaZero)
I do think it’s kind of funny that they mention AGI in that article, but the model is specifically not general.
"oi i need lik a scrip or somfing 2 take pic of me screen evry sec for min, mac"
with an actual (and usually functional) script to be "glorified grammar corrector", then sure.
"write a TCL/tk script file that is a "frontend" to the ls command: It should provide checkboxes and dropdowns for the different options available in bash ls and a button "RUN" to run the configured ls command. The output of the ls command should be displayed in a Text box inside the interface. The script must be runnable using tclsh"
It didn't get it right the first time (for some reason wants to put a `mainloop` instruction) but after several corrections I got an ugly but pretty functional UI.
Imagine a Linux Distro that uses some kind of LLM generated interfaces to make its power more accessible. Maybe even "self healing".
LLMs don't stop amazing me personally.
Current LLMs being 80% to being 100% useful doesn't mean there's only 20% effort left.
It means we got the lowest-hanging 80% of utility.
Bridging that last 20% is going to take a ton of work. Indeed, maybe 4x the effort that getting this far required.
And people also overestimate the utility of a solution that's randomly wrong. It's exceedingly difficult to build reliable systems when you're stacking a 5% wrong solution on another 5% wrong solution on another 5% wrong solution...
>And people also overestimate the utility of a solution that's randomly wrong. It's exceedingly difficult to build reliable systems when you're stacking a 5% wrong solution on another 5% wrong solution on another 5% wrong solution...
I call this the merry go round of hell mixed with a cruel hall of mirrors. LLM spits out a solution with some errors, you tell it to fix the errors, it produces other errors or totally forgets important context from one prompt ago. You then fix those issues, it then introduces other issues or messes up the original fix. Rinse and repeat. God help you if you don't actually know what you are doing, you'll be trapped in that hall of mirrors for all of eternity slowly losing your sanity.
I wrote some data visualizations with Claude and aider.
For anything that someone would actually pay for (expecting the robustness of paid-for software) I don’t think we’re there.
The devil is in the details, after all. And detail is what you lose when running reality through a statistical model.
This is just one more in a series of massive red flags around this company, from the insanely convoluted governance scheme, over the board drama, to many executives and key people leaving afterwards. It feels like Sam is doing the cleanup and anyone who opposes him has no place at OpenAI.
This, coming around the time where there are rumors of possible change to the corporate structure to be more friendly to investors, is an interesting timing.
"Burn out" doesn't apply when the issue at hand is AGI (and, possibly, superintelligence).
That said, I don't doubt that this particular departure was more the result of company politics, whether a product of the earlier board upheaval, performance related or simply the decision to bring in a new CTO with a different skill set.
Success in not around any corner. It's pure insanity to even believe that AGI is possible, let alone close.
AI is incapable of any innovation. It accelerates human innovation, just like any other piece of software, but that's it. AI makes protein folding more efficient, but it can't ever come up with the concept of protein folding on its own. It's just software.
You simply cannot have general intelligence without self-driven innovation. Not improvement, innovation.
But if we look at much more simple concepts, 2029 is only 5 years (not even) away, so I'm pretty confident that anything that it cannot do right now it won't be able to do in 2029 either.
I don't know her personal life or her feelings, but it doesn't seem like a stretch to imagine that she was just done.
being in San Francisco for 6 years and success means getting hauled in front of Congress and European Parliament
cant think of a worse occupational nightmare after having an 8-figure nest egg already
1) She has a very good big picture view of the market. She has probably identified some very specific problems that need to be solved, or at least knows where the demand lies.
2) She has the senior exec OpenAI pedigree, which makes raising funds almost trivial.
3) She can probably make as much, if not more, by branching out on her own - while having more control, and working on more interesting stuff.
Squeezing out senior execs could be a way for him to maximize his claim on the stake. Notwithstanding, the execs may have disagreed with the shift in culture.
Now I am become an AI language model, destroyer of the internet.