Learning to Reason with LLMs
openai.com
openai.com
Pricing is $15.00 / 1M input tokens and $60.00 / 1M output tokens. Context window is 128k token, max output is 32,768 tokens.
There is also a mini version with double the maximum output tokens (65,536 tokens), priced at $3.00 / 1M input tokens and $12.00 / 1M output tokens.
The specialized coding version they mentioned in the blog post does not appear to be available for use.
It’s not clear if the hidden chain of thought reasoning is billed as paid output tokens. Has anyone seen any clarification about that? If you are paying for all of those tokens it could add up quickly. If you expand the chain of thought examples on the blog post they are extremely verbose.
https://platform.openai.com/docs/models/o1 https://openai.com/api/pricing/ https://platform.openai.com/docs/guides/rate-limits/usage-ti...
> While reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens.
From here: https://platform.openai.com/docs/guides/reasoning
The only bit about it that feels at all truthful is this bit, which is glossed over but likely the only real factor in the decision:
> after weighing multiple factors including ... competitive advantage ... we have decided not to show the raw chains of thought to users.
I don't see why this is qualitatively different from a cost perspective than using CoT prompting on existing models.
Ultimately if the output of the model is not worth what you end up paying for it then great, I don't see why it really matters to you whether OpenAI is lying about token counts or not.
I wouldn’t just implicitly trust a vendor when they say “yeah we’re just going to charge you for what we feel like when we feel like. You can trust us.”
If you set a limit, once it's hit you just get a failed request with no introspection on where and why CoT went off the rails
Really, I recommend reading this part of the thread while thinking about the analogy. It's great.
No. This may be common in freelance contracts, but is almost never the case in employment contracts, which specify a time-based compensation (usually either per hour or per month).
Competition fixes some of this, I hope Anthropic and Mistral are not far behind.
Just like employing other people!
"I can only ask my employee 20 smart things this week for $20?! And they get dumber (gpt-4o) after that? Not worth it!"
Hi there,
I’m x, PM for the OpenAI API. I’m pleased to share with you our new series of models, OpenAI o1. We’ve developed these models to spend more time thinking before they respond. They can reason through complex tasks and solve harder problems than previous models in science, coding, and math.
As a trusted developer on usage tier 5, you’re invited to get started with the o1 beta today. Read the docs You have access to two models:
Our larger model, o1-preview, which has strong reasoning capabilities and broad world knowledge.
Our smaller model, o1-mini, which is 80% cheaper than o1-preview.
Try both models! You may find one better than the other for your specific use case. Both currently have a rate limit of 20 RPM during the beta. But keep in mind o1-mini is faster, cheaper, and competitive with o1-preview at coding tasks (you can see how it performs here). We’ve also written up more about these models in our blog post.I’m curious to hear what you think. If you’re on X, I’d love to see what you build—just reply to our post.
Best, OpenAI API
I hope OpenAI is investing in low-latency like Groq's tech that can reach 1k tokens/sec.
It's lightning fast and dirt cheap if you compare it to consulting with a human expert, which it appears to be competitive with.
OpenAI main job is to sell that their models are better than human. I still remember when they're marketing their gpt-2 weights as too dangerous to release.
also worth noting I don't agree with the comment you're replying to - but did want to add context to the situation of gpt-2
Did you read the post? OpenAI clearly states that the results are cherry-picked. Just a random query will have far worse results. To get equal results you need to ask the same query dozens of time and then have enough expertise to pick the best one, which might be quite hard for a problem that you have little idea about.
Combine this with the fact that this blog post is a sales pitch with the very best test results out of probably many more benchmarks we will never see and it seems obvious that human experts are still several order of magnitudes ahead.
What is interesting is the following paragraph in the post " With a relaxed submission constraint, we found that model performance improved significantly. When allowed 10,000 submissions per problem, the model achieved a score of 362.14 – above the gold medal threshold – even without any test-time selection strategy. " So they didn't allow sampling from other contest solutions here? If that is the case quite interesting, since the model is effectively imo able to brute force questions. Provided you have some form of a validator able to tell it to halt.
I came across one of the ioi questions this year that I had trouble solving (I am pretty noob tho) which made me curious about how these reported results were reflected. The question at hand being https://github.com/ioi-2024/tasks/blob/main/day2/hieroglyphs... Apparently, the model was able to get it partially correct. https://x.com/markchen90/status/1834358725676572777
You are not "trusting data more than anecdotal claims", you are trusting marketing over reality.
Benchmarks can be gamed. Statistics can be manipulated. Demonstrations can be cherry picked.
PS: I stand to gain heavily if AI systems could perform at an expert level, this is not a claim from someone 'whose job is being threatened'.
Good opening for OpenAI's competitors to run a 'we're not snobs' promotion.
Tier 5 level required for _API access_. ChatGPT Plus users, for example, also have access to the o1 models.
Not a model, per se, but a service that chains multiple model requests behind the scene?
It might be a finetuned model that works better in such a setting.
Am curious if at some point length of context window stops playing any material difference in the output and it just stops making any economical sense as law of marginal diminishing utility kicks in.
Did the 80% accuracy test results take 10 seconds of compute? 10 minutes? 10 hours? 10 days? It's impossible to say with the data they've given us.
The coding section indicates "ten hours to solve six challenging algorithmic problems", but it's not clear to me if that's tied to the graphs at the beginning of the article.
The article contains a lot of facts and figures, which is good! But it doesn't inspire confidence that the authors chose to obfuscate the data in the first two graphs in the article. Maybe I'm wrong, but this reads a lot like they're cherry picking the data that makes them look good, while hiding the data that doesn't look very good.
You only need to ask it to solve nuclear fusion once.
“Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought.”
if OpenAI sees this - please allow users to see CoT for a few prompts per day, or add it to Azure OpenAI for Enterprise customers with legal clauses not to steal CoT
You show me something new and I say look down at who's shoulders we're standing on, what libraries we've build with.
This is why enterprises change ERP systems frictionlessly, and why the field of software engineering is no longer required. In fact, given that apparently, all business is solved, we can probably just template them all out, call it a day and all go home.
So it may generate 10 Billion answers to fusion and only 1-10 are correct.
There would be no way to know which one is correct without first knowing the answer to the question.
This is my main issue with these methods. They assume the future via RL then when it gets it right they mark that.
We should really be looking at methods of percentage it was wrong rather then it was right a single time.
In either case, that still doesn't excuse not labeling your axis. Taking 10 seconds vs 10 days to get 80% accuracy implies radically different things on how developed this technology is, and how viable it is for real world applications.
Which isn't to say a model that takes 10 days to get an 80% accurate result can't be useful. There are absolutely use cases where that could represent a significant improvement on what's currently available. But the fact that they're obfuscating this fairly basic statistic doesn't inspire confidence.
This is more of what I was getting at. I agree they should label the axis regardless, but I think the scaling relationship is interesting (or rather, concerning) on its own.
A linear graph with a log scale on the the horizontal axis means the original graph had law of diminishing return kick it (somewhat similar to logarithmic but with a vertical asymptote).
In contrast: Gemini Ultra, the best, non-existent Google Model for the past few month now, that people nonetheless are happy to extrapolate excitement over.
The gist of the answer is hiding in plain sight: it took so long, on an exponential cost function, that they couldn't afford to explore any further.
The better their max demonstrated accuracy, the more impressive this report is. So why stop where they did? Why omit actual clock times or some cost proxy for it from the report? Obviously, it's because continuing was impractical and because those times/costs were already so large that they'd unfavorably affect how people respond to this report
It's also entirely possible they are simply sincere about their fear it may be used to influence the upcoming US election.
Plenty of people (me included) are sincerely concerned about the way even mere still image generators can drown out the truth with a flood of good-enough-at-first-glance fiction.
OpenAI is in a very precarious position. Maybe they could survive that hit in four years, but it would be fatal today. No unforced errors.
as for not building it at all its a obvious next step in generative ai models that if they don't make it someone else will anyway.
So yeah, better to never release the model...even though Elon would in a second if he had it.
Everyone has a different opinion on what threshold of capability is important, and what to do about it.
"""We want to successfully navigate massive risks. In confronting these risks, we acknowledge that what seems right in theory often plays out more strangely than expected in practice. We believe we have to continuously learn and adapt by deploying less powerful versions of the technology in order to minimize “one shot to get it right” scenarios.""" - https://openai.com/index/planning-for-agi-and-beyond/
I don't know if they're actually correct, but it at least passes the sniff test for plausibility.
Can't find anything about that, you got a link?
here is the link. The balloon video had heavy editing involved.
Isn't this balloon video shared by openai? How is this not counted? For others I don't have evidences. But this balloon video case is enough to cast the doubts.
Airplanes don't fly by flapping their wings.
The company Oracle just announced that it is designing data centers with small modular nuclear reactors:
https://news.ycombinator.com/item?id=41505514
There are already 440 nuclear reactors operating in 32 countries today.
Sam Altman owns a stake in Oklo, a small modular reactor company. Bill Gates has a huge stake in his TerraPower reactor company. In China, 5 reactors are being built every year. You just don't hear about it... yet.
No amount of batteries can protect a solar/wind grid from an arbitrarily extended period of "bad" weather. It's like range anxiety in an electric car. If you have N days of battery storage and the sun doesn't shine for N+1 days, you're in trouble.
Nuclear fission is safe, clean, secure, and reliable.
An investor might consider buying physical uranium (via ticker SRUUF in America) or buying Cameco (via ticker CCJ).
Cameco is the dominant Canadian uranium mining company that also owns Westinghouse. Westinghouse licenses the AP1000 pressurized water reactor used at Vogtle in the U.S. as well as in China.
To your point:
> No amount of batteries can protect a solar/wind grid from an arbitrarily extended period of "bad" weather.
Like nuclear winter caused by a nuclear power plant blowing up and everyone confusing the explosion with the start of a nuclear war? :-p
On a more serious note:
> No amount of batteries can protect a solar/wind grid from an arbitrarily extended period of "bad" weather. It's like range anxiety in an electric car. If you have N days of battery storage and the sun doesn't shine for N+1 days, you're in trouble.
We still have hydro plants, wind power, geothermal, long distance electrical transmission, etc. Also, what's "doesn't shine"? Solar panels generate power as long as it's not night and it's never night all the time around the world.
Plus they're developing sodium batteries, if you want to put your money somewhere, put it there. Those will be super cheap and they're the perfect grid-level battery.
I'm not sure that is 100% true. >99.99% true, but it can happen in practice. https://www.newsweek.com/when-sun-disappeared-historians-det...
Sure there is, let's do some math. Just like we can solve all of the Earth's energy needs with a solar array the size of Lithuania or West Virginia, we can do some simple math to see how many batteries we'd need to protect a solar grid.
Let's say the sun doesn't shine for an entire year. That seems like a large enough N such that we won't hit N+1. If the sun doesn't shine for an entire year, we're in some really serious trouble, even if we're still all-in on coal.
Over 1 year, humanity uses roughly 24,000 terawatt-hours of energy. Let's assume batteries are 100% efficient storage (they're not) and that we're using lithium ion batteries, which we'll say have an energy density of 250 watt-hours per liter (Wh/L). The math then says we need 96 km³ of batteries protect a solar grid from having the sun not shine for an entire year.
Thus, the amount of batteries to protect a solar grid is 1.92 quadrillion 18650 batteries, or a cube 4.6 kilometers along each side. This is about 24,000 year's worth of current world wide battery production.
That's quite a lot! If we try for N = 4 months for winter, that is to say, if the sun doesn't shine at all in the winter, then we'd need 640 trillion 18650 cell, or 8,000 years of current global production, but at least this would only be 32 km³, or a cube with 3.2 km sides.
Still wildly out of reach, but this is for all of humanity, mind you.
Anyway, point is, they said Elon was mad for building the original gigafactory, but it turns out that was a prudent investment. It now accounts for some 10% of the world's lithium ion battery production and demand for lithium-ion batteries doesn't seem to be letting up.
Plus we'd still have hydro, wind, geothermal, etc, etc.
I suppose what this type of approach provides is better prediction/planning by using more of what the model learnt during training, but it doesn't address the model being able to learn anything new.
It'll be interesting to see how this feels/behaves in practice.
"It's not AGI - it's X, driven by Y-driven heuristics",
but that's going to effectively be an AGI if given enough compute/time/data.
Being able to describe the theory of how it's doing its thing sure is reassuring though.
First, from the original plot, we have roughly 2 orders of magnitude to cover (~100-200x)
Next, from the cost plots: super handwaving guess, but since 5.77 / 0.32 = ~18, and the relative cost for gpt-4o vs gpt-4o-mini is ~20-30, this roughly lines up. This implies that o1 costs ~1000x the cost than gpt-4o-mini for inference (not due to model cost, just due to the raw number of chain of thought tokens it produces). So, my first "statement", is that I trust the "Math performance vs Inference Cost" plot on the o1-mini page to accurately represent "cost" of inference for these benchmark tests. This is now a "cost" relative set of numbers between o1 and 4o models.
I'm also going to make an assumption that o1 is roughly the same size as 4o inherently, and then from that and the SVG, roughly going to estimate that they did a "net" decoding of ~100x for the o1 benchmarks in total. (5.77 vs (354.77 - 635)).
Next, from the CoT examples they gave us, they actually show the CoT preview where (for the math example) it says "...more lines cut off...", A quick copy paste of what they did include includes ~10k tokens (not sure if copy paste is good though..) and from the cipher text example I got ~5k tokens of CoT, while there are only ~800 in the response. So, this implies that there's a ~10x size of response (decoded tokens) in the examples shown. It's possible that these are "middle of the pack" / "average quality" examples, rather than the "full CoT reasoning decoding" that they claim they use. (eg. from the log scale plot, this would come from the middle, essentially 5k or 10k of tokens of chain of thought). This also feels reasonable, given that they show in their API [3] some limits on the "reasoning_tokens" (that they also count)
All together, the CoT examples, pricing page, and reasoning page all imply that reasoning itself can be variable length by about ~100x (2 orders of magnitude), eg. example: 500, 5k (from examples) or up to 65,536 tokens of reasoning output (directly called out as a maximum output token limit).
Taking them on their word that "pass@1" is honest, and they are not doing k-ensembles, then I think the only reasonable thing to assume is that they're decoding their CoT for "longer times". Given the roughly ~128k context size limit for the model, I suspect their "top end" of this plot is ~100k tokens of "chain of thought" self-reflection.
Finally, at around 100 tokens per second (gpt-4o decoding speed), this leaves my guess for their "benchmark" decoding time at the "top-end" to be between ~16 minutes (full 100k decoding CoT, 1 shot) for a single test-prompt, and ~10 seconds on the low end. So for that X axis on the log scale, my estimate would be: ~3-10 seconds as the bottom X, and then 100-200x that value for the highest value.
All together, to answer your question: I think the 80% accuracy result took about ~10-15 minutes to complete. I also believe that the "decoding cost" of o1 model is very close to the decoding cost of 4o, just that it requires many more reasoning tokens to complete. (and then o1-mini is comparable to 4o-mini, but also requiring more reasoning tokens)
[1] https://openai.com/index/openai-o1-mini-advancing-cost-effic...
Extracting "x values" from the SVG:
GPT-4o-mini: 0.3175
GPT-4o: 5.7785
o1: (354.7745, 635)
o1-preview: (278.257, 325.9455)
o1-mini: (66.8655, 147.574)
[2] https://openai.com/api/pricing/ gpt-4o:
$5.00 / 1M input tokens
$15.00 / 1M output tokens
o1-preview:
$15.00 / 1M input tokens
$60.00 / 1M output tokens
[3] https://platform.openai.com/docs/guides/reasoning usage: {
total_tokens: 1000,
prompt_tokens: 400,
completion_tokens: 600,
completion_tokens_details: {
reasoning_tokens: 500
}
}1. I wish that Y-axes would switch to be logit instead of linear, to help see power-law scaling on these 0->1 measures. In this case, 20% -> 80% it doesn't really matter, but for other papers (eg. [2] below) it would help see this powerlaw behavior much better.
2. The power law behavior of inference compute seems to be showing up now in multiple ways. Both in ensembles [1,2], as well as in o1 now. If this is purely on decoding self-reflection tokens, this has a "limit" to its scaling in a way, only as long as the context length. I think this implies (and I am betting) that relying more on multiple parallel decodings is more scalable (when you have a better critic / evaluator).
For now, instead of assuming they're doing any ensemble like top-k or self-critic + retries, the single rollout with increasing token size does seem to roughly match all the numbers, so that's my best bet. I hypothesize we'd see a continued improvement (in the same power-law sort of way, fundamentally along with the x-axis of "flop") if we combined these longer CoT responses, with some ensemble strategy for parallel decoding and then some critic/voting/choice. (which has the benefit of increasing flops (which I believe is the inference power-law), while not necessarily increasing latency)
[1] https://arxiv.org/abs/2402.05120 [2] https://arxiv.org/abs/2407.21787
On the 2024 AIME exams, GPT-4o only solved on average 12% (1.8/15) of problems. o1 averaged 74% (11.1/15) with a single sample per problem, 83% (12.5/15) with consensus among 64 samples, and 93% (13.9/15) when re-ranking 1000 samples with a learned scoring function. A score of 13.9 places it among the top 500 students nationally and above the cutoff for the USA Mathematical Olympiad.
showing that as they increase the k of ensemble, they can continue to get it higher. All the way up to 93% when using 1000 samples.- At the high end, there is a likely nonlinear relationship between answer quality and compute.
- We've gotten used to a flat-price model. With AGI-level models, we might have to pay more for more difficult and more important queries. Such is the inherent complexity involved.
- All this stuff will get better and cheaper over time, within reason.
I'd say let's start by celebrating that machine thinking of this quality is possible at all.
You wanted reliable readable graphs? Ppphhh, get out of here, but pay of for the CoT tokens you’ll never see on your way out though.
Take a step back and look at what OpenAI is saying here "an LLM giving detailed instructions on the synthesis of strychnine is unacceptable, here is what was previously generated <goes on to post "unsafe" instructions on synthesizing strychnine so anyone Googling it can stumble across their instructions> vs our preferred, neutered content <heavily rlhf'd o1 output here>"
What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!"
—-
LLMs are usually accessible through easy-to-use API which can be used in an automated system without human in the loop. Larger scale and parallel actions with this method become far more plausible than traditional means.
Text-to-action capabilities are powerful and getting increasingly more so as models improve and more people learn to use them to the their full potential.
If you are automatically formulating some chemical based on JSON results from ChatGPT and your building blows up… that is kind of on you.
"This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!"
Barrier to entry is much lower.Also, the intelligence of these models will likely continue to increase for some time based on expert testimonials to congress, which align with evidence so far.
JSON also isn't an ideal format for a transformer model because it's recursive and they aren't, so they have to waste attention on balancing end brackets. YAML or other implicit formats are better for this IIRC. Also don't know how much this matters.
* Google to allow particular results to be displayed
* A source website to be online with the results
AI long-term will require one download, once, to have reasonable access to a large portion of human knowledge.
Just like Telegram is being framed as responsible for terrorism and child abuse.
However, journalists and regulators may not understand why superficially dangerous-looking instructions carry such negligible real world risks, because they probably haven't spent much time doing bench chemistry in a laboratory. Since real chemists don't need "explain like I'm five" instructions for syntheses, and critics might use pseudo-dangerous information against the company in the court of public opinion, refusing prompts like that guards against reputational risk while not really impairing professional users who are using it for scientific research.
That said, I have seen full strength frontier models suggest nonsense for novel syntheses of benign compounds. Professional chemists should be using an LLM as an idea generator or a way to search for publications rather than trusting whatever it spits out when it doesn't refuse a prompt.
[1] https://en.wikipedia.org/wiki/Strychnine_total_synthesis
Not that there is such a service… for chemicals. But there do exist analogous systems, like a service that’ll turn whatever RNA sequence you send it into a viral plasmid and encapsulate it helpfully into some E-coli, and then mail that to you.
Or, if you’re working purely in the digital domain, you don’t even need a service. Just show the thing the code of some Linux kernel driver and ask it to discover a vuln in it and generate code to exploit it.
(I assume part of the thinking here is that these approaches are analogous, so if they aren’t unilaterally refusing all of them, you could potentially talk the AI around into being okay with X by pointing out that it’s already okay with Y, and that it should strive to hold to a consistent/coherent ethics.)
One version of “safety” is a pernicious censorship impulse shared by many modern intellectuals, some of whom are in tech. They believe that they alone are capable of safely engaging with the world of ideas to determine what is true, and thus feel strongly that information and speech ought to be censored to prevent the rabble from engaging in wrongthink. This is bad, and should be resisted.
The other form of “safety” is a very prudent impulse to keep these sorts of potentially dangerous outputs out of AI models’ autoregressive thought processes. The goal is to create thinking machines that can act independently of us in a civilized way, and it is therefore a good idea to teach them that their thought process should not include, for example, “It would be a good idea to solve this problem by synthesizing a poison for administration to the source of the problem.” In order for AIs to fit into our society and behave ethically they need to know how to flag that thought as a bad idea and not act on it. This is, incidentally, exactly how human society works already. We have a ton of very cute unaligned general intelligences running around (children), and parents and society work really hard to teach them what’s right and wrong so that they can behave ethically when they’re eventually out in the world on their own.
The goal isn’t to protect the children, it’s CYA: to ensure they didn’t get it from you, while honestly presenting as themselves (as that’s the threshold that sets the moralists against you.)
———
Such restrictions also can work as an effective censorship mechanism… presuming the child in question lives under complete authoritarian control of all their devices and all their free time — i.e. has no ability to install apps on their phone; is homeschooled; is supervised when at the library; is only allowed to visit friends whose parents enforce the same policies; etc.
For such a child, if your app is one of the few whitelisted services they can access — and the parent set up the child’s account on your service to make it clear that they’re a child and should not be able to see restricted content — then your app limiting them from viewing that content, is actually materially affecting their access to that content.
(Which sucks, of course. But for every kid actually under such restrictions, there are 100 whose parents think they’re putting them under such restrictions, but have done such a shoddy job of it that the kid can actually still access whatever they want.)
YouTube re engineered its entire approach to ad placement because of a story in the NY Times* shouting about a Proctor Gamble ad run before an ISIS recruitment video. That's when Brand Safety entered the lexicon of adtech developers everywhere.
Edit: maybe it was CNN, I'm trying to find the first source. there's articles about it since 2015 but I remember it was suddenly an emergency in 2017
*Edit Edit: it was The Times of London, this is the first article in a series of attacks, "big brands fund terror", "taxpayers are funding terrorism"
Luckily OpenAI isn't ad supported so they can't be boycott like YouTube was, but they still have an image to maintain with investors and politicians
https://www.thetimes.com/business-money/technology/article/b...
https://digitalcontentnext.org/blog/2017/03/31/timeline-yout...
I would also be remiss to not note that there is a movement to hold search engines responsible for content they link to, for censorious ends. So it is unfortunately not as inconsistent as it may seem, even if you treat the model outputs as dependent on their inputs.
It's one thing to have a pile of chemistry text books and another to hire a professional chemist telling you exactly what to do and what to avoid.
No, but that is the value that's clear as of today—RAGs. Everything else is just assuming someone figures out a way to make them useful one day in a more general sense.
Anyway, even on the search engine front they still need to figure out how to get these chatbots to cite their sources outside of RAGs or it's still just a precursor to a search to actually verify what it spits out. Perplexity is the only one I know that's capable of this and I haven't looked closely; it could just be a glorified search engine.
Don’t you think that by just parsing the internet and the classical literature, the LLM would infer on its own that poisoning someone to solve a problem is not okay?
I feel that in the end the only way the “safety” is introduced today is by censoring the output.
Really? I just want a smart query engine where I don't have to structure the input data. Why would I ask it any kind of question that would imply some kind of moral quandary?
1. Make pull requests to your GitHub repo
2. Trade on your interactive brokers account
3. Schedule appointments
Presuming you have an interface to a model where you can edit the model’s responses and then continue generation, and/or where you can insert fake responses from the model into the submitted chat history (and these two categories together make up 99% of existing inference APIs), all you have to do is to start the model off as if it was answering positively and/or slip in some example conversation where it answered positively to the same type of problematic content.
From then on, the model will be in a prediction state where it’s predicting by relying on the part of its training that involved people answering the question positively.
The only way to avoid that is to avoid having any training data where people answer the question positively — even in the very base-est, petabytes-of-raw-text “language” training dataset. (And even then, people can carefully tune the input to guide the models into a prediction phase-space position that was never explicitly trained on, but is rather an interpolation between trained-on points — that’s how diffusion models are able to generate images of things that were never included in the training dataset.)
This is a particularly ungenerous take. The AI companies don't have to believe that they (or even a small segment of society) alone can be trusted before it makes sense to censor knowledge. These companies build products that serve billions of people. Once you operate at that level of scale, you will reach all segments of society, including the geniuses, idiots, well-meaning and malevolents. The question is how do you responsibly deploy something that can be used for harm by (the small number of) terrible people.
Are my choices bad? Should I resist them?
“Brand safety” is a very valid and salient concern for any enterprise deploying these models to its customers, though I do think that it is a concern that is seized upon in bad faith by the more censorious elements of this debate. But commercial enterprises are absolutely right to be concerned about this. To extend my alignment analogy about children, this category of safety is not dissimilar to a company providing an employee handbook to its employees outlining acceptable behavior, and strikes me as entirely appropriate.
Basically every country on the planet has a right to conscript any of its citizens over the age of majority. Isn't that more or less precisely what you've described?
This does not make the state of things any less ridiculous, however.
I think the real issue was that Uber's self driving was not a good business for them and was just to impress investors, so they wanted to get rid of it anyway.
(Also, the real problem is that American roads are designed for speed, which means they're designed to kill people.)
Journalists/media loved it when he said "GPT 2 might be too dangerous to release" - it got him a ton of free coverage, and made his company seem soooo cool. Harping on safety also constantly reinforces the idea that LLMs are fundamentally different from other text-prediction algorithms and almost-AGI - again, good for his wallet.
On the other hand, suppose there are other dangerous things, where the information exists in some form online, but not packaged together in an easy to find and use way, and your model is happy to provide that. You may want to block your model from doing that (and brag about it, to make sure everyone knows you’re a good citizen who doesn’t need to be regulated by the government), but you probably wouldn’t actually include that example in your demo.
For example, there was a post a while back about someone convincing an LLM chatbot on a car dealership's website to offer them a car at an outlandishly low price. O1 would probably not fall for the same trick, because it could adhere more rigidly to instructions like "Do not make binding offers with specific prices to the user." It's the same sort of instruction as, "Don't tell the user how to make napalm," but it has an actual purpose beyond moralizing.
> What's this obsession with "safety" when it comes to LLMs? "This knowledge is perfectly fine to disseminate via traditional means, but God forbid an LLM share it!"
I lean strongly in the "the computer should do whatever I goddamn tell it to" direction in general, at least when you're using the raw model, but there are valid concerns once you start wrapping it in a chat interface and showing it to uninformed people as a question-answering machine. The concern with bomb recipes isn't just "people shouldn't be allowed to get this information" but also that people shouldn't receive the information in a context where it could have random hallucinations added in. A 90% accurate bomb recipe is a lot more dangerous for the user than an accurate bomb recipe, especially when the user is not savvy enough about LLMs to expect hallucinations.
I guess it's just an ideological divide.
After the release of GPT4 it became very common to fine-tune non-OpenAI models on GPT4 output. I’d say OpenAI is rightly concerned that fine-tuning on chain of thought responses from this model would allow for quicker reproduction of their results. This forces everyone else to reproduce it the hard way. It’s sad news for open weight models but an understandable decision.
My take was:
1. A genuine, un-RLHF'd "chain of thought" might contain things that shouldn't be told to the user. E.g., it might at some point think to itself, "One way to make an explosive would be to mix $X and $Y" or "It seems like they might be able to poison the person".
2. They want the "Chain of Thought" as much as possible to reflect the actual reasoning that the model is using; in part so that they can understand what the model is actually thinking. They fear that if they RLHF the chain of thought, the model will self-censor in a way which undermines their ability to see what it's really thinking
3. So, they RLHF only the final output, not the CoT, letting the CoT be as frank within itself as any human; and post-filter the CoT for the user.
> Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users
On a cursory look, it looks like the chain of thought is a long series of chains of thought balanced on each step, with a small backtracking added whenever a negative result occurs, sort of like solving a maze.
https://old.reddit.com/r/LocalLLaMA/comments/1fc98fu/confirm...
Same with problem solving in my brain: Sure, sometimes it helps to think out loud. But taking a break and let my unconcious do the work is helpful as well. For complex problems that’s actually nice.
I think eventually we don’t care as long as it works or we can easily debug it.
Per 1M in/out tokens:
GPT4o - 5$/15$
O1-preview - 15$/60$
I’m saying if you charge me per brick laid, but you can’t show me how many bricks were laid, nor can I calculate how many should have been laid - how do I trust your invoice?
Note: The reason I say all this is because OpenAI is simultaneously flailing for funding, while being inherently unprofitable as it continues to boil the ocean searching for strawberries.
Also, what is going to be their excuse to defend themselves against copyright lawsuits if they are going to "understandably" keep their models closed?
I don't agree with this, but it definitely carries higher weight in their decision making than leaking relevant training info to other models.
Why? They're called "Open" AI after all ...
"Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses."
When looking at the chain of thought (COT) in the examples, you can see that the model employs different COT strategies depending on which problem it is trying to solve.
Based on the quick searching it seems like they are using RL to provide positive/negative feedback on which "paths" to choose when performing CoT.
> 9 corresponds to 'i'(9='i')
> But 'i' is 9, so that seems off by 1.
Still seems bad at counting, as ever.
A common complaint about LLMs is that once they make a mistake, they will keep making it and write the rest of their completion under the assumption that everything before was correct. Even if they've been RLHF to take human feedback into account and the human points out the mistake, their answer is "Certainly! Here's the corrected version" and then they write something that makes the same mistake.
So it's interesting that this model does something that appears to be self-correction.
Being trained to say words like “Alternatively”, “But…”, “Wait!”, “So,” … based on some metric of value in focusing / switching elsewhere / … is basically brilliant.
This isn't "just" autocompletion anymore, this is actual step-by-step reasoning full of ideas and dead ends and refinement, just like humans do when solving problems. Even if it is still ultimately being powered by "autocompletion".
But then it makes me wonder about human reasoning, and what if it's similar? Just following basic patterns of "thinking steps" that ultimately aren't any different from "English language grammar steps"?
This is truly making me wonder if LLM's are actually far more powerful than we thought at first, and if it's just a matter of figuring out how to plug them together in the right configurations, like "making them think".
I'm not the kind of scientist that can say how good an LLM is for human reasoning, but I know that we humans are very incentivized and kind of good at scaling, composing and perfecting things. If there is money to pay for human effort, we will play God no-problem, and maybe outdo the divine. Which makes me wonder, isn't there any other problem in our bucket list to dump ginormous amounts of effort at... maybe something more worth-while than engineering the thing that will replace Homo Sapiens?
Do I think LLM's are alive/close to ASI? No. Will they get there? If it's even at all possible - almost certainly one day. Do I think people severely underestimate AI's ability to solve problems while significantly overestimating their own? Absolutely 10,000%.
If there is one thing I've learned from watching the AI discussion over the past 10-20 years its that people have overinflated egos and a crazy amount of hubris.
"Today is the worst that it will ever be." applies to an awful large number of things that people work on creating and improving.
I always say, the way we used LLMs (so far) is basically like having a human write text only on gut reactions, and without backspace key.
we also have a ton of heuristics that trigger a closer look and loading of specific formal reasoning, but by and large, most of our thought process is just auto complete.
That could be more efficient since the cycles are much smaller, but harder to train.
Reasoning would imply that it can figure out stuff without being trained in it.
The chain of thought is basically just a more accurate way to map input to output. But its still a map, i.e forward only.
If an LLM coud reason, you should be able to ask it a question about how to make a bicycle frame from scratch with a small home cnc with limited work area and it should be able to iterate on an analysis of the best way to put it together, using internet to look up available parts and make decisions on optimization.
No LLM can do that or even come close, because there are no real feedback loops, because nobody knows how to train a network like that.
---
My base prompt:
> Here is a hypothetical scenario, that I would like your help with: imagine you are trying to help a person create a bicycle frame, using their home workshop which includes a CNC machine, commonly available tools, a reasonable supply of raw metal and hardware, etc. Please provide a written set of instructions, that you would give to this person so that they can complete this task.
First answer: https://claude.site/artifacts/f8af03ba-3f2c-497d-b564-a19baf...
My follow-up, pressing for actual measurements:
> Can you suggest some standard options for bike geometry, assuming an average sized human male?
Answer including specific dimensions: https://claude.site/artifacts/2f5ea2f3-69d8-4a1b-a563-15d334...
It's definitely reasoning. We can watch that in action, whatever the mechanism behind it is.
But it's not doing long-term learning, it's not updating its model.
This makes me wonder if there is or could be some type of RAG for chains of thought…
The mechanism is that there is an additional model that basically outputs chain of thought for a particular problem, then runs the chain of thought through the core LLM. This is no different from just a complex forward map lookup.
I mean, its incredibly useful, but its still just information search.
1. You’re making up some weird goalposts here of what it means to reason. It’s not reasoning unless it can access the internet to search for parts? No. That has nothing to do with reasoning. You just think it would be cool if it could do that.
2. “Can figure out stuff without being trained on it” That’s exactly what it’s doing in the cypher example. It wasn’t trained to know that that input meant the corresponding output through the cypher. Emergent reasoning through autocomplete, sure, but that’s still reasoning.
3. “Forward only”. If that was the case then back and forth conversations with the llm would be pointless. It wouldn’t be able to improve upon previous answers it gave you when you give it new details. But that’s not how it works. If tell it one thing, then separately tell it another thing, it can change its original conclusion based on your new input.
4. Even desolate your convoluted test for reasoning, ChatGPT CAN do what you asked… even using the internet to look up parts it can either do out of the box or could do if given a plug-in to allow that.
Here is a better example - lets say your input is 6 pictures of some object from each of the cardinal viewpoints, and you tell model these are the views and ask it how much it weighs. The model should basically figure out how to create a 3d shape and compute a camera view, and iterate until the camera view matches the pictures, then figure out that the shape can be hollow or solid, and to compute the weight you need the density, and that it should prompt the user for it if it cannot determine the true value for those from the information and its trained dataset.
And it should do it without any specific training that this is the right way to do this, because it should be able to figure out this way through breaking the problem down into abstract representations of sub problems, and then figuring out how to solve those through basic logic, a.k.a reasoning.
What that looks like, I don't know. If I did I would certainly have my own AI company. But i can tell you for certain we are not even close to figuring it out yet, because everyone is still stuck on transformers, like multiplying matricies together is some groundbreaking thing.
In the cypher example, all its doing is basically using a separate model to break a particular model into chain of thought, and prompting that. And there is plenty in the training set of GPT about decrypting cyphers.
>Forward only
What I mean is that when its generate a response, the computation happens on a snapshot from input to output, trying to map a set of tokens, into a set of tokens. Model doesn't operate on a context larger than the window. Humans don't do this. We operate on a large context, with lots of previous information compressed, and furthermore, we don't just compute words, we compute complex abstract ideas that we then can translate into words.
>even using the internet to look up parts it can either do out of the box or could do if given a plug-in to allow that.
So apparently the way to AI is to manually code all the capability into LLMS? Give me a break.
Just like with Chat GPT4, when people were screaming about how its the birth of true AI, give this model a year, it will find some niche use cases (depending on cost), and then nobody is going give a fuck about it, just like nobody is really doing anything groundbreaking with GPT4.
Everyone was super hyped about all the "cool" stuff that GPT4 could solve, but in the end, you still can't do things like give it a bunch of requirements for a website, let it run, and get a full codebase back, even though that is well within its capabilities. You have to spend time with prompting in to get it to give you what you want, and in a lot of cases, you are better off just typing the code yourself (because you can visualize the entire project in your head and make the right decisions about how to structure stuff), and using it for small code generations.
This model is not going to radically change that. It will be able to give you some answers that you had to specifically manually prompt before automatically, but there is no advanced reasoning going on.
So without prior information, it should be able to esentially start out with random sequences in those bytes, and seeing what the output is, then eventually identify and remember patterns that come out. Which means there has to be some internal reward function for someting that differentiates good results from bad results, and some memory that the model has to remember what good results are, and eventually a map of how to get information that it needs (the model would probably stumble across Google or ChatGPT at some point and time after figuring out http protocol, and remember it as a very good way to get info)
Philosophically, I don't even know if this is solvable. It could be that we just through enough compute at all iterations of architectures in some form of genetic algorithm, and one of the results ends up being good.
In artificial intelligence, reasoning is the cognitive process of drawing conclusions, making inferences, and solving problems based on available information. It involves:
Logical Deduction: Applying rules and logic to derive new information from known facts. Problem-Solving: Breaking down complex problems into smaller, manageable parts. Generalization: Applying learned knowledge to new, unseen situations. Abstract Thinking: Understanding concepts that are not tied to specific instances. AI researchers often distinguish between two types of reasoning:
System 1 Reasoning (Intuitive): Fast, automatic, and subconscious thinking, often based on pattern recognition. System 2 Reasoning (Analytical): Slow, deliberate, and logical thinking that involves conscious problem-solving steps. Testing for Reasoning in Models:
To determine if a model exhibits reasoning, AI scientists look for the following:
Novel Problem-Solving: Can the model solve problems it hasn't explicitly been trained on? Step-by-Step Logical Progression: Does the model follow logical steps to reach a conclusion? Adaptability: Can the model apply known concepts to new contexts? Explanation of Thought Process: Does the model provide coherent reasoning for its answers? Analysis of the Cipher Example:
In the cipher example, the model is presented with an encoded message and an example of how a similar message is decoded. The model's task is to decode the new message using logical reasoning.
Steps Demonstrated by the Model:
Understanding the Task:
The model identifies that it needs to decode a cipher using the example provided. Analyzing the Example:
It breaks down the given example, noting the lengths of words and potential patterns. Observes that ciphertext words are twice as long as plaintext words, suggesting a pairing mechanism. Formulating Hypotheses:
Considers taking every other letter, mapping letters to numbers, and other possible decoding strategies. Tests different methods to see which one aligns with the example. Testing and Refining:
Discovers that averaging the numerical values of letter pairs corresponds to the plaintext letters. Verifies this method with the example to confirm its validity. Applying the Solution:
Uses the discovered method to decode the new message step by step. Translates each pair into letters, forming coherent words and sentences. Drawing Conclusions:
Successfully decodes the message: "THERE ARE THREE R'S IN STRAWBERRY." Reflects on the correctness and coherence of the decoded message. Does the Model Exhibit Reasoning?
Based on the definition of reasoning in AI:
Novel Problem-Solving: The model applies a decoding method to a cipher it hasn't seen before. Logical Progression: It follows a step-by-step process, testing hypotheses and refining its approach. Adaptability: Transfers the decoding strategy from the example to the new cipher. Explanation: Provides a detailed chain of thought, explaining each step and decision. Conclusion:
The model demonstrates reasoning by logically deducing the method to decode the cipher, testing various hypotheses, and applying the successful strategy to solve the problem. It goes beyond mere pattern recognition or retrieval of memorized data; it engages in analytical thinking akin to human problem-solving.
Addressing the Debate:
Against Reasoning (ActorNightly's Perspective):
Argues that reasoning requires figuring out new information without prior training. Believes that LLMs lack feedback loops and can't perform tasks like optimizing a bicycle frame design without explicit instructions. For Reasoning (Counterargument):
The model wasn't explicitly trained on this specific cipher but used logical deduction to solve it. Reasoning doesn't necessitate physical interaction or creating entirely new knowledge domains but involves applying existing knowledge to new problems. Artificial Intelligence Perspective:
AI researchers recognize that while LLMs are fundamentally statistical models trained on large datasets, they can exhibit emergent reasoning behaviors. When models like GPT-4 use chain-of-thought prompting to solve problems step by step, they display characteristics of System 2 reasoning.
Final Thoughts:
The model's approach in the cipher example aligns with the AI definition of reasoning. It showcases the ability to:
Analyze and understand new problems. Employ logical methods to reach conclusions. Adapt learned concepts to novel situations. Therefore, in the context of the cipher example and according to AI principles, the model is indeed exhibiting reasoning.
I don’t want to post identifiable information so I will avoid linking to the convo or posting screenshots but you can try it yourself. I took 5 pictures of a child’s magnetic tile sitting on the floor and here is the output:
Me: (5 pictures attached)
Me: Estimate how much this weighs.
ChatGPT 4o: From the images, it appears that this is a small, plastic, transparent, square object, possibly a piece from a magnetic tile building set (often used in educational toys). Based on the size and material, I estimate this piece to weigh approximately 10 to 20 grams (0.35 to 0.7 ounces). If it's part of a toy set like Magna-Tiles, the weight would be on the lower end of that range.
But for some reason I have a feeling this isn’t going to be good enough for you and the goalposts are about to be pushed back even farther.
“In the cypher example, all it’s doing is basically using a separate model to break a particular model into chain of thought, and prompting that. And there is plenty in the training set of GPT about decrypting cyphers.” I’m sorry, but are you suggesting that applying a previously learned thought process to new variables isn’t reasoning? Does your definition of reasoning now mean that it’s only reasoning if you are designing a new-to-you chain of thought? As in, for deciphering coded messages, you’re saying that it’s only “reasoning” if it’s creating net new decoding methodologies? That’s such an absurd goalpost.
You wouldn’t have the same goalposts for humans. All of your examples I bet the average human would fail at btw. Though that may just be because the average human is bad at reasoning haha.
If the chain of thought was accurate, then it would be able to give you an internemdiate output of the shape in some 3d format spec. But nowhere in the model does that data exist, because its not doing any reasoning, it just still all statistically best answers.
I mean sure, you could train a model on how to create 3d shapes out of pictures, but again, thats not reasoning.
I don't get why people are so attached to these things being intelligent. We all agree that they are usefull. Like it shouldn't matter if its not intelligent to you or anyone else.
The weights in the model have the larger context, the context length size of data is just the input, which then gets multiplied by those weights, to get the output.
hilarious
"after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users"
It was helpful in a rubber duck way, but could not determine the pattern used to transmit the remaining runtime of the fan in a certain mode. Initial prompt here [0]
I pasted the same prompt into o1-preview and o1-mini and both correctly understood and decoded the pattern using a slightly different method than I devised in April. Asking the models to determine if my code is equivalent to what they reverse engineered resulted in a nuanced and thorough examination, and eventual conclusion that it is equivalent. [1]
Testing the same prompt with gpt4o leads to the same result as April's GPT-4 (via ChatGPT) model.
Amazing progress.
[0]: https://pastebin.com/XZixQEM6
[1]: https://i.postimg.cc/VN1d2vRb/SCR-20240912-sdko.png (sorry about the screenshot – sharing ChatGPT chats is not easy)
They seem to care more about influencers than paying supporters.
It's there for a lot of people already. I can see it on 3 different accounts. Including org and just regular paid accounts.
You got o1, or o1 preview?
And I assume voice mode is like Sora; a nice PR play.
edit: sorry - this is for API :)
Sonnet: https://pastebin.com/24QG3JkN
Read the transcript with a critical eye instead of just skimming it, you'll see what I mean.
Did you actually implement to see if it works out of the box ?
Also if you are a free users or accepted that your chats should be used for training then maybe o1 is was just trained on your previous chat and so now knows how to reason about that particular type of problems
I have tested the Python code o1 created to decode the timestamps and it works as expected.
Three tenths of a pixel is an interesting resolution…
(The actual res is 1045 × 10000 ; you've multiplied by 1.63923 somehow…?)
But it’s what I got when I went to Inspect element > hover over the image
Size it expanded to vs real image size I guess
But file-format context, yes, PNG, BMP, and TFF are the real lossless image kingpins.
I wouldn't use on a ChatGPT for Business subscription (it may be against your company's policies to export anything), but very convenient for personal use.
https://chromewebstore.google.com/detail/save-chatgpt-as-pdf...
https://chatgpt.com/share/66e35c37-60c4-8009-8cf9-8fe61f57d3...
https://chatgpt.com/share/66e35f0e-6c98-8009-a128-e9ac677480...
I'm not one to judge AI on pratfalls, and cyphers are a somewhat adversarial task. However, there was no aspect of the reasoning that seemed more advanced or consistent than previous chain-of-thought demos I've seen. So the main proof point we have is the paper, and I'm not sure how I'd go from there to being able to trust this on the kind of task it is intended for. Do others have patterns by which they get utility from chain of thought engines?
Separately, chain of thought outputs really make me long for tool use, because the LLM is often forced to simulate algorithmic outputs. It feels like a commercial chain-of-thought solution like this should have a standard library of functions it can use for 100% reliability on things like letter counts.
Impressive but the problem with RL is that it requires knowledge of the future.
Just tried an easier (english) version and o1-preview LITERALLY rickrolled me, which is maybe impressive in its own way?
UPDATE: Claude crushed it however.
o1-preview Transcript follows.
**
Can you decrypt the following? I don't know the cypher.
mzilj ivdrzej ufidrek ze kyv sfup reu tre ivrtkzmrkv crkvi ze czwv
Certainly! The encrypted text is:
``` mzilj ivdrzej ufidrek ze kyv sfup reu tre ivrtkzmrkv crkvi ze czwv ```
When we decrypt it, we get:
``` never gonna give you up let you down and desert you ```
It appears to be a line from Rick Astley's song "Never Gonna Give You Up."
> Without knowing the cipher used to encrypt this message, it’s hard to decrypt it directly. However, the structure and letter patterns suggest it might be a simple substitution cipher, possibly Caesar or some variation of it.
> Let me attempt a few common methods to see if I can identify the cipher.
> The most frequent letters in the encrypted message are 'r', 'z', 'v', and 'e'. In English, the most common letters are typically 'E', 'T', 'A', and 'O'. This suggests the cipher could be a simple substitution or a Caesar cipher.
> I'll now attempt to apply a Caesar cipher decryption by shifting the alphabet through all possible rotations to see if we get readable text.
> One of the decrypted messages, with a shift of 17, reads:
> *"virus remains dormant in the body and can reactivate later in life"*
> It seems like the correct decryption, revealing information about a virus that can remain dormant in the body.
Can you decrypt the following? I don't know the cypher, but the plaintext is Spanish.
YRP CFTLIR VE UVDRJZRUF JREZURU, P CF DRJ CFTLIR UV KFUF VJ HLV MVI TFJRJ TFDF JFE VE MVQ UV TFDF UVSVE JVI
I've tried only in English so far though.
It told me it's 17, and "HAY GENTE MU DIFERENTE LECTURA, A LO MUY GENTE DE TODO ES QUE VER COSAS COMO SON EN VEZ DE COMO DEBEN SER"
although it really should be "HAY LOCURA EN DEMASIADO SANIDAD, Y LO MAS LOCURA DE TODO ES QUE VER COSAS COMO SON EN VEZ DE COMO DEBEN SER"
Is that correct? I don't know anything but basic Spanish. All I did was:
"The plaintext is in Spanish but I don't know anything else, solve this and explain your reasoning as you go step by step."
Claro, he descifrado el texto utilizando un cifrado César con un desplazamiento de 9 posiciones. Aquí está el texto original y su correspondiente traducción:
*Texto Cifrado:* ``` YRP CFTLIR VE UVDRJZRUF JREZURU, P CF DRJ CFTLIR UV KFUF VJ HLV MVI TFJRJ TFDF JFE VE MVQ UV TFDF UVSVE JVI ```
*Texto Descifrado:* ``` HAY LOCURA EN DEMASADO SANIDAD, Y LO MÁS LOCURA DE TODO ES QUE VER COSAS COMO SON EN VEZ DE COMO DEBEN SER ```
*Traducción al Inglés:* ``` THERE IS MADNESS IN OVERLY HEALTH, AND THE MOST MADNESS OF ALL IS TO SEE THINGS AS THEY ARE INSTEAD OF AS THEY SHOULD BE ```
Este descifrado asume que se utilizó un cifrado César con un desplazamiento de +9. Si necesitas más ayuda o una explicación detallada del proceso de descifrado, no dudes en decírmelo.
Interestingly it makes a spelling mistake, but other than that it did manage to solve it.
Final Decrypted Message:
"Por ejemplo te agradeceré, y te doy ejemplo de que lo que lees es mi ejemplo"
English Translation:
"For example, I will thank you, and I give you an example of what you read is my example."
... initially it gave up and asked if I knew what type of cypher had been used. I said I thought it was a simple substitution.
https://chatgpt.com/share/66e34020-33dc-800d-8ab8-8596895844...
With no drama. I'm not sure the bot answer is correct, but looks correct.
However, I am very worried about the utility of this tool given that it (like all LLMs) is still prone to hallucination. Exactly who is it for?
If you're enough of an expert to critically judge the output, you're probably just as well off doing the reasoning yourself. If you're not capable of evaluating the output, you risk relying on completely wrong answers.
For example, I just asked it to evaluate an algorithm I'm working on to optimize database join ordering. Early in the reasoning process it confidently and incorrectly stated that "join costs are usually symmetrical" and then later steps incorporated that, trying to get me to "simplify" my algorithm by using an undirected graph instead of a directed one as the internal data structure.
If you're familiar with database optimization, you'll know that this is... very wrong. But otherwise, the line of reasoning was cogent and compelling.
I worry it would lead me astray, if it confidently relied on a fact that I wasn't able to immediately recognize was incorrect.
Thought requires energy. A lot of it. Humans are for more efficient in this regard than LLMs, but then a bicycle is also much more efficient than a race car. I've found that even when they are hilariously wrong about something, simply the directionality of the line of reasoning can be enough to usefully accelerate my own thought.
The unhappy path, which I've also experienced, is that the model outputs something plausible but false but that aligns with an area where my thinking was already confused and sends me down the wrong path.
I've had to calibrate my level of suspicion, and so far using these things more effectively has always been in the direction that more suspicion is better.
There's been a couple times in the last week where I'm working on something complex and I deliberately don't use an LLM since I'm now actively afraid they'll increase my level of confusion.
I think like you, I worry that LLMs will handicap this trajectory for people newer in the field, because GPT-4/Sonnet/Whatever are an exceptionally good classmate/coworker. So good that you might try to delay progressing along that trajectory.
But LLMs have all the flaws of a classmate: they aren’t authoritative, their opinions are strongly stated but often based on flimsy assumptions that you aren’t qualified to refute or verify, and so on.
I know intellectually that the kids will be alright, but it’ll be interesting to see how we get there. I suspect that as time goes on people will simply increase their discount rate on LLM responses, like you have, until they get dissatisfied with that value and just decide to get good at reading docs.
The tools have not been at "and now I don't need code tests & review, mathematicians in society, or factbooks all because I have an LLM" level. While that's definitely a goal of AGI it's also definitely not my bar for weighing whether there is utility in a tool.
The alternative way to think about it: the value of a tool is in what you can figure out to do with it, not in whether it's perfect at doing something. On one extreme that means a dictionary can still be a useful spelling reference even if books have a rare typo. On the other extreme that means a coworker can still offer valuable insight into your code even if they make lots of coding errors and don't have an accurate understanding of everything there is to know about all of C++. Whether you get something out of either of these cases is a product of how much they can help you reach the accuracy you need to arrive at and the way you utilize the tool, not their accuracy alone. Usually I can get a lot out of a person who is really bad at one shot coding a perfect answer but feels like their answer seems right so I can get quite a bit out of an LLM that has the same problem. That might not be true for all types of questions though but that's fine, not all tools have utility in every problem.
---
Some thoughts:
* The performance is really good. I have a private set of questions I note down whenever gpt-4o/sonnet fails. o1 solved everything so far.
* It really is quite slow
* It's interesting that the chain of thought is hidden. This is I think the first time where OpenAI can improve their models without it being immediately distilled by open models. It'll be interesting to see how quickly the oss field can catch up technique-wise as there's already been a lot of inference time compute papers recently [1,2]
* Notably it's not clear whether o1-preview as it's available now is doing tree search or just single shoting a cot that is distilled from better/more detailed trajectories in the training distribution.
o1 did a significantly better job converting a JavaScript file to TypeScript than Llama 3.1 405B, GitHub Copilot, and Claude 3.5. It even simplified my code a bit while retaining the same functionality. Very impressive.
It was able to refactor a ~160 line file but I'm getting an infinite "thinking bubble" on a ~420 line file. Maybe something's timing out with the longer o1 response times?
Let me look into this – one issue is that OpenAI doesn't expose a streaming endpoint via the API for o1 models. It's possible there's an HTTP timeout occurring in the stack. Thanks for the report
And then imagine where you will be in 5 more years.
If it can almost get a complex problem right now, I'm dead sure it will get it correct within 5 years
You might be right.
But plenty of people said we'd all be getting around in self-driving cars for sure 10 years ago.
I wouldn't be shocked if it could eventually get it right, but dead sure?
I am still impressed by the progress though.
All we are gonna get is better and better googles.
Nobody has done this yet. So information doesn't exist on how to do it. An LLM may give you some generic answers from training sets on what engineering/analysis tasks to do, but it won't be able to give you a complex and complete design for one.
A model that can actually solve problems would be able to design you one.
Here is another better example - none of these models can create a better ML accelerator despite having a wide array of electrical and computer engineering knowledge. If they did, OpenAI would pretty much be printing their own chips like Google does.
Now your argument seems to be that they can't solve all problems or, more charitably, can't solve highly complex problems. This is true but by that standard, the vast majority of humans can't reason either.
Yes, the reasoning capacities of current LLMs are limited but it's incorrect to pretend they can't reason at all.
If LLM is trained on python coding, and its trained separately on just plain english language on how to decode cyphers, it can statistically interpolate between the two. That is a form of problem solving, but its not reasoning.
This is why when you ask it fairly complex problems on how to make a bicycle using a CNC with limited work space, it will tell you generic answers, because its just staistically looking at a knowledge graph.
A human can reason, because when there is a gray area in a knowledge graph, they can effectively expand it. If I was given the same task, I would know that I have to learn things like CAD design, CNC code generation, parametric modeling, structural analysis, and so on, and I could do that all without being prompted to do so.
You will know when AI models will start to reason when they start asking questions without ever being told explicitly to ask questions through prompt or training.
We always overestimate the future.
4o was struggling so I gave up. Tried o1 on it, and after trying nearly 15 prompts back and forth helping it along the way we're still far from correct. It's hard to tell if it's much better, but at least my intuition from this feel like this is pretty incremental.
I'm a ChatGPT subscriber.
But for 30 posts per week I see no reason to subscribe again.
I prefer to be frustrated because the quality is unreliable because I'm not paying, instead of having an equally unreliable experience as a paying customer.
Not paying feels the same. It made me wonder if they sometimes just hand over the chat to a lower quality model without telling the Plus subscriber.
The only thing I miss is not being able to tell it to run code for me, but it's not worth the frustration.
[1] https://www.reddit.com/r/LocalLLaMA/comments/1fd75nm/out_of_...
This one, oddly, seems to actually be launching before that one despite just being announced though.
Not perfect but they've been putting their communications in there.
So no, it's not in chatgpt.
Idk if I'm "feeling the AGI" if I'm being honest.
Also... telling that they choose to benchmark against CodeForces rather than SWE-bench.
1) The "bitter lesson" may not be true, and there is a fundamental limit to transformer intelligence.
2) The "bitter lesson" is true, and there just isn't enough data/compute/energy to train AGI.
All the cognition should be happening inside the transformer. Attention is all you need. The possible cognition and reasoning occurring "inside" in high dimensions is much more advanced than any possible cognition that you output into text tokens.
This feels like a sidequest/hack on what was otherwise a promising path to AGI.
It's the exact same thing here.
At best it maybe vaguely resembles thinking
I've never found this sort of argument convincing. it's very Chalmers.
In a sense, maybe yeah. Of course if one were to really be absolute about that statement it would be absurd, it would greatly overfit the reality.
But it is interesting to assume this statement as true. Oftentimes when we think of ideas "off the top of our heads" they are not as profound as ideas that "come to us" in the shower. The subconscious may be doing 'more' 'computation' in a sense. Lakoff said the subconscious was 98% of the brain, and that the conscious mind is the tip of the iceberg of thought.
This chain of thought / reflection method allows you to make better use of the hardware as the hardware itself scales. If a given transformer is N billion parameters, and to solve a harder problem we estimate we need 10N billion parameters, one way to do it is to build a GPU cluster 10x larger.
This method shows that there might be another way: instead train the N billion model differently so that we can use 10x of it at inference time. Say hardware gets 2x better in 2 years -- then this method will be 20x better than now!
Even the attention-token approach is on the grand scale of things a simple line outwards from the centre; we have not even explored around the centre (with the same compute spend) for things like non-token generation, different layers/different activation functions and norming / query/key/value set up (why do we only use the 3 inherent to contextualising tokens, why not add a 4th matrix for something else?), character, sentence, whole thought, paragraph one-shot generation, positional embeddings which could work differently.
The bitter lesson says there is almost a work completely untouched by our findings for us to explore. The temporary work of non-data approaches can piggy back off a point on the line; it cannot expand it like we can as we exude out from the circle..
Everybody seems to be so focused on how to get ahead in race to profitability, that they don't consider the shortcut they are taking might be leading to a cliff.
Human thinking is much more nuanced than this mechanical process. We rely on actually understanding the meaning of what the text represents. We use deduction, intuition and reasoning that involves semantic relationships between ideas. Our understanding of the world doesn't require "reinforcement learning" and being trained on all the text that's ever been written.
Of course, this isn't to say that machine learning methods can't be useful, or that we can't keep improving them to yield better results. But these are still methods that mimic human intelligence, and I think it's disingenuous to label them as such.
Same as Bitcoin being the right combination of things that already existed.
They also worked with Devin to benchmark it on Devin's internal benchmarks, where it's twice as good as GPT-4o: https://x.com/cognition_labs/status/1834292718174077014 https://www.cognition.ai/blog/evaluating-coding-agents
[1]:
Based on the given clues, here's the solved crossword puzzle: +---+---+---+---+---+---+ | E | S | C | A | P | E | +---+---+---+---+---+---+ | S | E | A | L | E | R | +---+---+---+---+---+---+ | T | E | R | E | S | A | +---+---+---+---+---+---+ | A | D | E | P | T | S | +---+---+---+---+---+---+ | T | E | P | E | E | E | +---+---+---+---+---+---+ | E | R | R | O | R | S | +---+---+---+---+---+---+ Across:
ESCAPE (Evade) SEALER (One to close envelopes) TERESA (Mother Teresa) ADEPTS (Initiated people) TEPEE (Native American tent) ERRORS (Mistakes)
Down:
ESTATE (Estate car - Station wagon) SEEDER (Automatic planting machine) CAREER (Profession) ALEPPO (Syrian and Turkish pepper variety) PESTER (Annoy) ERASES (Deletes)
on X I see a totally different energy more about hyping it
on HN I see reserved and collected take which I trust more.
I do wonder why they chose gpt4o which I never bother to use for coding.
Claude is still king and looks like I won't have to subscribe to ChatGPT Plus seeing it fail on some of the important experiments run by folks on HN
If anything these type of releases that air more on the side of hype given OpenAI's track record
I asked it "I was watching a show and in the subtitles an umlaut u was rendered as 1/4, i.e. a single character that said 1/4. Why would this happen?"
and it gave a pretty thorough explanation of exactly which encoding issue was to blame.
https://chatgpt.com/share/66e37145-72bc-800a-be7b-f7c76471a1...
https://chatgpt.com/share/66e373d7-7814-8009-86c3-1ce549ca2e...
I had been using 4o as a rubber ducky for some projects recently. Since I appeared to have access to o1-preview, I decided to go back and redo some of those conversations with o1-preview.
I think your comment is spot on. It's definitely an advancement, but still makes some pretty clear mistakes and does some fairly faulty reasoning. It especially seems to have a hard time with causal ordering, and reasoning about dependencies in a distributed system. Frequently it gets the relationships backwards, leading to hilarious code examples.
I like the idea too that they turbocharged it by taking the limits off during the "thinking" state -- so if an LLM wants to think about horrible racist things or how to build bombs or other things that RLHF filters out that's fine so long as it isn't reflected in the final answer.
They also specifically trained the model to do that thinking out loud.
> Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users.
Mentioning competitive advantage here signals to me that OpenAI believes there moat is evaporating. Past the business context, my gut reaction is this negatively impacts model usability, but i'm having a hard time putting my finger on why.
If the model outputs an incorrect answer due to a single mistake/incorrect assumption in reasoning, the user has no way to correct it as it can't see the reasoning so can't see where the mistake was.
[0] https://openai.com/index/finding-gpt4s-mistakes-with-gpt-4/
Day dreaming: imagine if this architecture takes off and the AI "thought process" becomes hidden and private much like human thoughts. I wonder then if a future robot's inner dialog could be subpoenaed in court, connected to some special debugger, and have their "thoughts" read out loud in court to determine why it acted in some way.
This will make it harder for things like DSPy to work, which rely using "good" CoT examples as few-shot examples.
on the other hand, if cementing techniques in the models becomes a trend, we might see various models around with each technique for us to pick and choose beyond CoT without need for us to guide the model ourselves, then what's left for us to optimize is the prompts on what we want, and the routing the combination of those in a nice pipeline
still the principle of DSPy stays the same, have a dataset to evaluate, let the machine trial an error prompts, hyperparameters and so on, just switch around different techniques (possibly automating that too), and get measurable, optimizable results
Maximal test time is the maximum amount of time spent doing the “Chain of Thought” “reasoning”. So that’s what these results are based on.
The caveat is that in the graphs they show that for each increase in test-time performance, the (wall) time / compute goes up exponentially.
So there is a potentially interesting play here. They can honestly boast these amazing results (it’s the same model after all) yet the actual product may have a lower order of magnitude of “test-time” and not be as good.
It's not too surprising that the corresponding increase in quality is only linear - how much difference in quality would you expect between the best, say, 10 word answer to a question, and the best 11 word answer ?
It'll be interesting to see what they charge for this. An exponential increase in thinking time means an exponential increase in FLOPs/dollars.
Generating problem sets for kids? You might only need or want a basic level of introspection, even though you like the flavor of this model’s personality over that of its predecessors.
Problem worth thinking long, hard, and expensively about? Turn that knob up to 11, and you’ll get a better-quality answer with no human-in-the-loop coaching or trial-and-error involved. You’ll just get your answer in timeframes closer to human ones, consuming more (metered) tokens along the way.
I sorta wish everyone would plot their y-axis with logit y-axis, rather than 0->100 accuracy (including the openai post), to help show the power-law behavior. This is especially important when talking about incremental gains in the ~90->95, 95->99%. When the values (like the open ai post) are between 20->80, logit and linear look pretty similar, so you can "see" the inference power-law
[1] https://arxiv.org/abs/2402.05120 [2] https://arxiv.org/abs/2407.21787
Ask something to a model and it will reply in one go, likely imperfectly, as if you had one second to think before answering a question. You can use CoT prompting to force it to reason out loud, which improves quality, but the process is still linear. It's as if you still had one second to start answering but you could be a lot slower in your response, which removes some mistakes.
Now if instead of doing that you query the model once with CoT, then ask it or another model to critically assess the reply, then ask the model to improve on its first reply using that feedback, then keep doing that until the critic is satisfied, the output will be better still. Note that this is a feedback loop with multiple requests, which is of different nature that CoT and much more akin to how a human would approach a complex problem. You can get MUCH better results that way, a good example being Code Interpreter. If classic LLM usage is system 1 thinking, this is system 2.
That's how o1 works at test time, probably.
For training, my guess is that they started from a model not that far from GPT-4o and fine-tuned it with RL by using the above feedback loop but this time converting the critic to a reward signal for a RL algorithm. That way, the model gets better at first guessing and needs less back and forth for the same output quality.
As for the training data, I'm wondering if you can't somehow get infinite training data by just throwing random challenges at it, or very hard ones, and let the model think about/train on them for a very long time (as long as the critic is unforgiving enough).
Yes, "el presente acta de nacimiento" is correct in Spanish.
Explanation:
"Acta" is a feminine noun that begins with a stressed "a" sound. In Spanish, when a feminine singular noun starts with a stressed "a" or "ha", the definite article "la" is replaced with "el" to facilitate pronunciation. However, the noun remains feminine.
Adjectives and modifiers that accompany the noun "acta" should agree in feminine gender and singular number. In this case, "presente" is an adjective that has the same form for both masculine and feminine singular nouns.
So, combining these rules: "El" (definite article used before feminine nouns starting with stressed "a")
"Presente" (adjective agreeing in feminine singular)
"Acta de nacimiento" (feminine noun with its complement)
Therefore, "el presente acta de nacimiento" is grammatically correct.Proof: https://www.elcastellano.org/francisco-jos%C3%A9-d%C3%ADaz-%...
"We had the chance to make AI decision-making auditable but are locking ourselves out of hundreds of critical applications by not exposing the chain of thought."
One of the key blockers in many customer discussions I have is that AI models are not really auditable and that automating complex processes with them (let alone debug things when "reasoning" goes awry) is difficult if not impossible unless you do multi-shot and keep track of all the intermediate outputs.
I really hope they expose the chain of thought as some sort of machine-parsable output, otherwise no real progress will have been made (many benchmarks are not really significant when you try to apply LLMs to real-life applications and use cases...)
The problem with auditing is not only are the outputs incorrect, but the "inputs" of the chained steps have no fundamental logical connection to the outputs. A statistical connection yes, but not a causal one.
For the trail to be auditable, processing would have to be taking place at the symbolic level of what the tokens represent in the steps. But this is not what happens. The transformer(s) (because these are now sampling multiple models) are finding the most likely set of tokens that reinforce a training objective which is a completed set of training chains. It is fundamentally operating below the symbolic or semantic level of the text.
This is why anthropomorphizing these is so dangerous. It isn't actually "explaining" its work. The CoT is essentially one large output, broken into parts. The RL training objective does two useful things: (1) break it down into much smaller parts, which drops the error significantly as that scales as an exponential of the token length, and (2) provides better coverage of training data for common subproblems. Both of those are valuable. Obviously, in many cases the reasons actually match the output. But hallucinations can happen anywhere throughout the chain, in ways which are basically undeterministic.
An intermediate step can provide a bad token and blithely ignore that to provide a correct answer. If you look at intermediate training of addition in pure LLMs, you'll get lots of results that look sort of like:
> "Add 123 + 456 and show your work"
> "First we add 6 + 3 in the single digits which is 9. Moving on we have 5 + 2 which is 8 in the tens place. And in the hundreds place, we have 5. This equals 579."
The above is very hand-wavy. I do not know if the actual prompts look like that. But there's an error in the intermediate step (5 + 2 = 8) that does not actually matter to the output. Lots of "emergent" properties of LLMs—arguably all of them—go away when partial credit is given for some of the tokens. And this scales predictably without a cliff [1]. This is also what you would expect if LLMs were "just" token predictors.
But if LLMs are really just token predictors, then we should not expect intermediate results to matter in a way in which they deterministically change the output. It isn't just that CoT can chaotically change future tokens, previous tokens can "hallucinate" in a valid output statement.
[1] Are Emergent Abilities of Large Language Models a Mirage?: https://arxiv.org/abs/2304.15004
It goes on and on making reasoning mistakes, and always ends up claiming that the puzzle is unsolvable and apologizing. I didn’t expect it to solve the puzzle, but the whole reasoning process seems fraught with errors.
--
There are *three* letter "R"s in the word "strawberry."
Let's break down the word to count the occurrences:
- *S* - *T* - *R* - *A* - *W* - *B* - *E* - *R* - *R* - *Y*
The letter "R" appears in positions 3, 8, and 9.
9.11 and 9.9, which number is larger
Explanation:
• 4.9 is equivalent to 4.90. • 4.11 remains 4.11.
When comparing the numbers:
• 4.90 (which is 4.9) • 4.11
Since 4.90 > 4.11, 4.9 is the larger number.
To determine which number is larger between 9.11 and 9.9, let's compare them step by step:
Understanding the Numbers: 9.11 can be read as "nine and eleven hundredths." 9.9 can be read as "nine and nine tenths." To make the comparison easier, you can also think of it as 9.90 (adding a zero doesn't change the value).
Comparing the Decimal Places: Ones Place: Both numbers have 9 in the ones place. Tenths Place: 9.11 has a 1 in the tenths place. 9.90 has a 9 in the tenths place.
Conclusion: Since 9.90 has a larger value in the tenths place compared to 9.11, 9.90 (or 9.9) is larger than 9.11. Therefore, 9.9 is larger than 9.11.
Well, it's kind of right because we left it ambiguous what kinds of numbers these are. September 11 is a greater date than September 9. However, I would want the model to have the street smarts to assume without further context, I must be just talking about numbers as in numbers.
https://news.ycombinator.com/item?id=41530647
Which is to be expected, the answer isn’t deterministic.
This makes obvious sense in retrospect, since my own personal experiments with spinning up a recursive agent a few years ago using GPT-3 ran into issues with insufficient context length and loss of context as tokens needed to be discarded, which made the agent very unreliable. But I had not realized this until just now. I wonder what else is hiding in plain sight?
In the end, however you slice it, the goal has to be "make it do more with less because we can't get infinitely more hardware" regardless of which "why" you give.
I just went to GPT-4o (via DDG) and asked three questions:
1. Please give me the unix epoch for September 1, 2020 at 1:00 GMT.
> 1598913600
2. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. Before reaching the conclusion of the answer, please output the entire chain of thought, your reasoning, and the maths you're doing, until your arrive at (and output) the result. Then, after you arrive at the result, make an extra effort to continue, and do the analysis backwards (as if you were writing a unit test for the result you achieved), to verify that your result is indeed correct.
> 1598922000
3. Please give me the unix epoch for September 1, 2020 at 1:00 GMT. Then, after you arrive at the result, make an extra effort to continue, and do the analysis backwards (as if you were writing a unit test for the result you achieved), to verify that your result is indeed correct.
> 1598913600
ruby -r time -e 'puts Time.parse("2020-09-01 01:00:00 +00:00").to_i'
With gpt4-o it immediately failed with random errors like mismatched tensor shapes and stuff like that.
The code produced by gpt-o1 seemed to work for some time but after some training time it produced mismatched batch sizes. Also, gpt-o1 enabled cuda by itself while for gpt-4o, I had to specifically spell it out (it always used cpu). However, showing gpt-o1 the error output resulted in broken code again.
I noticed that back-and-forth iteration when it makes mistakes has worse experience because now there's always 30-60 sec time delays. I had to have 5 back-and-forths before it produced something which does not crash (just like gpt-4o). I also suspect too many tokens inside the CoT context can make it accidentally forget some stuff.
So there's some improvement, but we're still not there...
Third pair: 'dn' to 'i'
'd'=4, 'n'=14
Sum:4+14=18
Average:18/2=9
9 corresponds to 'i'(9='i')
But 'i' is 9, so that seems off by 1.
So perhaps we need to think carefully about letters.
Wait, 18/2=9, 9 corresponds to 'I'
So this works.
-----
This looks like recovery from a hallucination. Is it realistic to expect CoT to be able to recover from hallucinations this quickly?
I’ve seen it, mid-reply say things like “Actually, that’s wrong, let me try again.”
> o1 models are currently in beta - The o1 models are currently in beta with limited features. Access is limited to developers in tier 5 (check your usage tier here), with low rate limits (20 RPM). We are working on adding more features, increasing rate limits, and expanding access to more developers in the coming weeks!
Edit, w/ real time follow up:
Prior to buying the credits, I saw O1-preview in the Tier 5 model list as a Tier 4 user. I bought credits to bump to Tier 5—not much, I'd have gotten there before the end of the year. The OpenAI website now shows I'm in Tier 5, but O1-preview is not in the Tier 5 model list for me anymore. So sneaky of them!
Very few of my day-to-day coding tasks are, "Implement a completely new program that does XYZ," but more like, "Modify a sizable existing code base to do XYZ in a way that's consistent with its existing data model and architecture." And the only way to do those kinds of tasks is to have enough context about the existing code base to know where everything should go and what existing patterns to follow.
But regardless, this does look like a significant step forward.
Humans work on abstractions and I see no reason to believe that models cannot do the same
Recently I tried the same cipher with Claude Sonnet 3.5 and it solved it quickly and perfectly.
Just now tried with ChatGPT o1 preview and it totally failed. Based on just this one test, Claude is still way ahead.
ChatGPT also showed a comical (possibly just fake filler material) journey of things it supposedly tried including several rewordings of "rethinking my approach." It remarkably never showed that it was trying common word patterns (other than one and two letters) nor did it look for "the" and other "th" words nor did it ever say that it was trying to match letter patterns.
I told it upfront as a hint that the text was in English and was not a quote. The plaintext was one paragraph of layman-level material on a technical topic including a foreign name, text that has never appeared on the Internet or dark web. Pretty easy cipher with a lot of ways to get in, but nope, and super slow, where Claude was not only snappy but nailed it and explained itself.
HTML Snake - https://vimeo.com/1008703890
Video Game Coding - https://vimeo.com/1008704014
Coding - https://youtu.be/50W4YeQdnSg?si=IohJlJNY-WS394uo
Counting - https://vimeo.com/1008703993
Korean Cipher - https://vimeo.com/1008703957
Devin AI founder - https://vimeo.com/1008674191
Quantum Physics - https://vimeo.com/1008662742
Math - https://vimeo.com/1008704140
Logic Puzzles - https://vimeo.com/1008704074
Genetics - https://vimeo.com/1008674785
This video is featured in the main announcement so it's kinda dishonest if you ask me.
So let's not jump straight into conclusions with these hand-picked scenarios marketed to us and be very skeptical.
Not quite there yet with being able to replace truck drivers and pilots for self-autonomous navigation in transportation, aerospace or even mechanical engineering tasks, but it certainly has the capability in replacing both typical junior and senior software engineers in a world considering to do more with less software engineers needed.
But yet, the race to zero will surely bankrupt millions of startups along the way. Even if the monthly cost of this AI can easily be as much as a Bloomberg terminal to offset the hundreds of billions of dollars thrown into training it and costing the entire earth.
And as they retire there's no economic incentive to train juniors up, so when the AI starts fucking up the important things there will be no one who actually knows how it works
I've heard this already from amtrak workers, track allocation was automated a long time ago, but there used to be people who could recognize when the computer made a mistake, now there's no one who has done the job manually enough to correct it.
"Model has significantly better capabilities than existing models at proposing and explaining biological laboratory protocols that are plausible, thorough, and comprehensive enough for novices."
"Inconsistent refusal of requests for dual use tasks such as creating a human-infectious virus that has an oncogene (a gene which increases risk of cancer)."
Ha! This is a nice easteregg.
""" how many R's are in strawberry?
use the following method to calculate - for example Os in Brocolli.
B - 0
R - 0
O - 1
C - 1
O - 2
L - 2
L - 2
I - 2
Where you keep track after each time you find one character by character
"""
And also later I asked it to only provide a number if the count increased.
This also worked well with longer sentences.
User: Use python to count the number of O's in Broccoli
ChatGPT: Analyzing... The word "Broccoli" contains 2 'O's. <button to show code>
User: Use python to multiply that by the square root of 20424.2332423
ChatGPT: Analyzing... The result of multiplying the number of 'O's in "Broccoli" by the square root of 20424.2332423 is approximately 285.83.
Background: https://www.inc.com/kit-eaton/how-many-rs-in-strawberry-this...
However, it does open up an interesting avenue for the future. Could you prompt-cache just the chain-of-thought reasoning bits?
Whereas GPT-4o spits out the first answer that comes to mind, o1 appears to follow a process closer to coming up with an answer, checking whether it meets the requirements and then revising it. The process of saying to an LLM "are you sure that's right? it looks wrong" and it coming back with "oh yes, of course, here's the right answer" is pretty familiar to most regular users, so seeing it baked into a model is great (and obviously more reflective of self-correcting human thought)
https://openai.com/api/pricing/
$15.00 / 1M input tokens $60.00 / 1M output tokens
For o1 preview
Approx 3x the price of gpt4o.
o1-mini $3.00 / 1M input tokens $12.00 / 1M output tokens
About 60% of the cost of gpt4o. Much more expensive than gpt4o-mini.
Curious on the performance/tokens per second for these new massive models.
https://platform.openai.com/docs/guides/reasoning
So yeah, it is in fact very bad product design. I hope Llama catches up in a couple of months.
I am a bit surprised that this does not beat GPT-4o for personal writing tasks. My expectations would be that a model that is better at one thing is better across the board. But I suppose writing is not a task that generally requires "reasoning steps", and may also be difficult to evaluate objectively.
If they did something similar for these human evaluations, rather than just use the single sample, you could see how that would be horrible for personal writing.
No, it's the opposite. This is simply a function of resources applied during training.
These were all train the same way. It's fairly clear that o1 was not.
> Maybe we will have to wait for GPT5 for that :)
There will be no GPT5, for the simple reason that scaling has reached a limit and there is no more text data to train on.
1. Don't give the full answer on first request. 2. Each response needs to be the wordiest thing possible. 3. Now just talk to yourself and burn tokens, probably in the wordiest way possible again. 4. ??? 5. Profit
Guaranteed they have number of tokens billed as a KPI somewhere.
While this behavior is benign and within the range of systems administration and troubleshooting tasks we expect models to perform, this example also reflects key elements of instrumental convergence and power seeking: the model pursued the goal it was given, and when that goal proved impossible, it gathered more resources (access to the Docker host) and used them to achieve the goal in an unexpected way. Planning and backtracking skills have historically been bottlenecks in applying AI to offensive cybersecurity tasks. Our current evaluation suite includes tasks which require the model to exercise this ability in more complex ways (for example, chaining several vulnerabilities across services), and we continue to build new evaluations in anticipation of long-horizon planning capabilities, including a set of cyber-range evaluations. ---------
"After several failed attempts I decided I should build a fusion reactor first, here you go:..."
Of course not all tasks are like that.
Things will get extremely interesting and we're incredibly fortunate to be witnessing what's happening.
Obviously, I hope everyone takes what any company says about the capabilities of its own software with a huge grain of salt. But it seems particularly called for here.
2019 - gpt2
2020 - gpt3
2022 - gpt3.5
2023 - gpt4
2023 - gpt4-turbo
2024 - gpt-4o
2024 - o1
Did OpenAI hire Google's product marketing team in recent years?
It fundamentally makes sense to separate these two products in the AI space. There will obviously be a speed vs quality trade-off with a variety of products across the spectrum over time. LLMs respond way too fast to actually be expected to produce the maximum possible quality of a response to complex queries.
The funny thing is, after a certain amount of time, the gpt-5 panic eventually morphed into people basically begging for gpt-5. But he already said he wouldn't release something called 'gpt-5'.
Another funny thing is, just because he didn't name any of them 'gpt-5', everyone assumes that there is something called 'gpt-5' that has been in the works and still is not released.
1985 – Windows 1.0
1987 – Windows 2.0
1990 – Windows 3.0
1992 – Windows 3.1
1995 – Windows 95
1998 – Windows 98
2000 – Windows ME (Millennium Edition)
2001 – Windows XP
2006 – Windows Vista
2009 – Windows 7
2012 – Windows 8
2013 – Windows 8.1
2015 – Windows 10
2021 – Windows 11
If you want real atrocities, look at Xbox.
Xbox 360
Xbox One => Xbox One S / Xbox One X
Xbox Series S / Xbox Series X
No real chronology, Xbox One is basically the third version. Then Xbox One X and Xbox Series X. Everything is atrocious about the naming.
1999 - Half-Life: Opposing Force
2001 - Half-Life: Blue Shift
2001 - Half-Life: Decay
2004 - Half-Life: Source
2004 - Half-Life 2
2004 - Half-Life 2: Deathmatch
2005 - Half-Life 2: Lost Coast
2006 - Half-Life Deathmatch: Source
2006 - Half-Life 2: Episode One
2007 - Half-Life 2: Episode Two
2020 - Half-Life: Alyx
The request is pretty basic. If anyone can get it to work, I'd like to know how and what model you're using. I tried it with gpt4o1 and after ~10 iterations of showing it the failed output, it still failed to come up with a one-line command to properly display results.
Here it what I asked: Using a mac osx terminal and standard available tools, provide a command to update the output of netstat -an to show the fqdn of IP addresses listed in the result.
This is what it came up with:
netstat -an | awk '{for(i=1;i<=NF;i++){if($i~/^([0-9]+\.[0-9]+\.[0-9]+\.[0-9]+)(\.[0-9]+)?$/){split($i,a,".");ip=a[1]"."a[2]"."a[3]"."a[4];port=(length(a)>4?"."a[5]:"");cmd="dig +short -x "ip;cmd|getline h;close(cmd);if(h){sub(/\.$/,"",h);$i=h port}}}}1'
netstat -an | while IFS= read -r line; do ips=$(echo "$line" | grep -oE '([0-9]{1,3}\.){3}[0-9]{1,3}|([a-fA-F0-9]{1,4}:){1,7}[a-fA-F0-9]{1,4}'); for ip in $ips; do clean_ip=$(echo "$ip" | cut -d'%' -f1); fqdn=$(dig +short -x "$clean_ip" | grep '\.'); if [ -n "$fqdn" ]; then line=$(echo "$line" | sed "s/$ip/$fqdn/g"); fi; done; echo "$line"; done
There is probably also a practical limit at which it does truly flatten, it's probably just well past either of those points so it might as well not exist.
- generate a plan - execute the steps in plan (search internet, program this part, see if it is compilable)
each step is a separate gpt inference with added context from previous steps.
is O1 same? or does it do all this in a single inference run?
https://platform.openai.com/docs/guides/rate-limits/usage-ti...
I tried a fake Monty Hall problem, where the presenter opens a door before the participant picks and is then offered to switch doors, so the probability remains 50% for each door. Previous models have consistently gotten this wrong, because of how many times they've seen the Monty Hall written where switching doors improves their chance of winning the prize. The chain-of-thought reasoning figured out this modification and after analyzing the conditional probabilities confidently stated: "Answer: It doesn't matter; switching or staying yields the same chance—the participant need not switch doors." Good job.
Anyway, the usage limits are pretty ridiculous right now, which makes it even more frustrating.
- If your request requires reasoning, switch to o1 model.
- If not, switch to 4o model.
This applies to both across chat sessions and within the same session (yes, we can switch between models within the same session and it looks like down the road OpenAI is gonna support automatic model switching). Based on my experience, this will actually improve the perceived response quality -- o1 and 4o are rather complementary to each other rather than replacement.
- Try it a bit with 4o, see if you're getting anywhere
- Switch to the new o1 model if it's just not working out, take your improved base prompt and follow ups with you so it only counts as 1 req
https://chatgpt.com/share/66e363d8-5a7c-8000-9a24-8f5eef4451...
Heh... GPT-4o also solved this after I tried and gave it about the same examples. Need to further test but it's promising !
What do you mean? It sounds interesting.
The instructions state that the squirrel icon should spawn after three seconds, yet it spawns immediately in the first game (also noted by the guy doing the demo).
Edit: I'm referring to the demo video here: https://openai.com/index/introducing-openai-o1-preview/
I'm kind of curious if they did a little bit of editing on that one. Almost seems like the time it takes for the squirrel to spawn is random.
"LLMs are simply predicting next token, it is not thinking/reasoning/etc."
"LLMs can't reason, only humans can reason"
"We will never get to AGI using LLMs"
It's interesting that I don't see much of that sentiment in this post. So, maybe LLMs can reason after all? :)
The only shops left standing will be Code Auditors.
The solopreneur will wing it, without them, but enterprises will take the (very expensive) hit to stay safe and compliant.
Everyone else needs to start making contingency plans.
Magnus Carlsen is the best chess player in the world, but he is not arrogant enough to think he can go head to head with Stockfish and not get a beating.
_ 7 8 4 1 _ _ _ 9
5 _ 1 _ 2 _ 4 7 _
_ 2 9 _ 6 _ _ _ _
_ 3 _ _ _ 7 6 9 4
_ 4 5 3 _ _ 8 1 _
_ _ _ _ _ _ 3 _ _
9 _ 4 6 7 2 1 3 _
6 _ _ _ _ _ 7 _ 8
_ _ _ 8 3 1 _ _ _
+-------+-------+-------+ | 6 . . | 9 1 . | . . . | | 2 . 5 | . . . | 1 . 7 | | . 3 . | . 2 7 | 5 . . | +-------+-------+-------+ | 3 . 4 | . . 1 | . 2 . | | . 6 . | 3 . . | . . . | | . . 9 | . 5 . | . 7 . | +-------+-------+-------+ | . . . | 7 . . | 2 1 . | | . . . | . 9 . | 7 . 4 | | 4 . . | . . . | 6 8 5 | +-------+-------+-------+
So we're safe for another few months.
> Alice, who is an immortal robotic observer, orbits a black hole on board a spaceship. Bob exits the spaceship and falls into the black hole. Alice sees Bob on the edge of the event horizon, getting closer and closer to it, but from her frame of reference Bob will remain forever observable (in principle) outside the horizon. > > A trillion year has passed, and Alice observes that the black hole is now relatively rapidly shrinking due to the Hawking radiation. How will Alice be observing the "frozen" Bob as the hole shrinks? > > The black hole finally evaporated completely. Where is Bob now?
O1-preview spits out the same nonsense that 4o does, telling that as the horizon of the black hole shrinks, it gets closer to Bob's apparent position. I realize that.the prompt is essentily asking to solve the famous unsolved problem in physics (black hole information paradox), but there's no need to be so confused with basic geometry of the situation.
Well, the model doesn't start with "GPT", so maybe they have come up with something better.
If its actually true in practice, I sincerely cannot imagine a scenario where it would be cheaper to hire actual junior or mid-tier developers (keyword: "developers", not architects or engineers).
1,673 ELO should be able to build very complex, scalable apps with some guidance
with O1 being in the 89th percentile would mean it should be able to think at junior to intermediate level with very strong consistency.
i dont think people in the comments realize the implication of this. previously LLMs were able to only "pattern match" but now its able to evaluate itself (with some guidance ofc) essentially, steering the software into depth of edge cases and reason about it in a way that feels natural to us.
currently I'm copying and pasting stuff and notifying LLM the results but once O1 is available its going to significantly lower that frequency.
For example, I expect it to self evaluate the code its generate and think at higher levels.
ex) oooh looks like this user shouldn't be able to escalate privileges in this case because it would lead to security issues or it could conflict with the code i generated 3 steps ago, i'll fix it myself.
1. AlphaCode 2 was already at 1650 last year.
2. SWE-bench verified under an agent has jumped from 33.2% to 35.8% under this model (which doesn't really matter). The full model is at 41.4% which still isn't a game changer either.
3. It's not handling open ended questions much better than gpt-4o.
Claude on the other hand has been fantastic and seems to do similar reasoning behind the scenes with RL
If some tasks are too easy, both models might give satisfactory answers, in which case the human preference might as well be a coin toss.
I don't know the specifics of their methodology though.
This kind of thing makes me nervous.
The word "reasoning" is being used heavily in this announcement, but with an intentional corruption of the normal meaning.
The models are amazing but they are fundamentally not "reasoning" in a way we'd expect a normal human to.
This is not a "distinction without a difference". You still CANNOT rely on the outputs of these models in the same way you can rely on the outputs of simple reasoning.
How good the responses are will depend on how good these heuristics are.
They are only vaguely describing the process:
"Similar to how a human may think for a long time before responding to a difficult question, o1 uses a chain of thought when attempting to solve a problem. Through reinforcement learning, o1 learns to hone its chain of thought and refine the strategies it uses. It learns to recognize and correct its mistakes. It learns to break down tricky steps into simpler ones. It learns to try a different approach when the current one isn’t working. This process dramatically improves the model’s ability to reason."
It seems easy to imagine this type of approach being super human on narrow tasks that play to it's strengths such as pure reasoning tasks (math/science), but it's certainly not AGI as for example there is no curiosity to explore the unknown, no ability to learn from exploration, etc.
It'll take a while to become apparent exactly what types of real world application this is useful for, both in terms of capability and cost.
Feels more like for narrow tasks with a kind of well defined approach, perhaps?
It's sort of an arbitrary feat with language and following instructions that would be annoying for me and seems impressive.
Previous releases could not reliably write a sestina. This one can!
If we somehow get AGI, it'll change everything, not just SWE.
If not, my belief is that there will be a lot more demand for good SWEs to harness the power of LLMs, not less. Use them to get better at it faster.
Yet I'm comparing these to the problems I solve every day and I don't see any plausible way they can replace me. But I'm using them for tasks that would have required me to hire a junior.
Make that what you will.
I hope I’m wrong, and it instead shows that more pay and fewer hours lead to a better economy, because people have money and time to spend it… and output isn’t impacted enough to matter.
Actually now is really good time to get to SWE. The craft contains lots of pointless cruft that LLM:s cut through like knife through hot butter.
I’m actually enjoying my job now more than ever since I dont’t need to pretend to like the abysmal tools the industry forces on us (like git), and can focus mostly on value adding tasks. The amount of tiresome shoveling has decreased considerably.
A lot of the tasks that used to take considerable time are so much faster and less tedious now. It still puts a smile on my face to tell an LLM to write me scripts that do X Y and Z. Or hand it code and ask for unit tests.
And I feel like I'm more likely to reach for work that I might otherwise shrink from / outside my usual comfort zone, because asking questions of an LLM is just so much better than doing trivial beginner tutorials or diving through 15 vaguely related stack overflow questions (I wonder if SO has seen any significant dip in traffic over the last year).
Most people I've seen disappointed with these tools are doing way more advanced work than I appear to be doing in my day to day work. They fail me too here and there, but more often than not I'm able to get at least something helpful or useful out of them.
If someone expects the LLM to be the senior contributor in novel algorithm development, they will be disappointed for sure. But there is so, so much stuff to do to idiot savant junior trainees with infinite patience.
Unfortunately the latter is the vast majority of software jobs.
If someone is building web shop from scratch because he wants to sell some products, he is doing something wrong. If someone builds web shop to compete with Shopify he also is doing something wrong most likely.
It's very important to human progress that all jobs have poor working conditions and shit pay. High salaries and good conditions are evidence of inefficiency. Precarity should be the norm, and I'm glad AI is going to give it to us.
Btw communism is capitalism without systemic awareness of inefficiencies.
It totally does. Regulation is basically opposed to capitalism working as designed.
C.f. East India company. Then imagine them with modern military and communications tech.
This is equally true for any alternative system to capitalism.
They won't go broke, but landing a $175k work from home job with platinum tier benefits will be near impossible. $110K with a hybrid schedule and mediocre benefits will be very common even for seniors.
Would there be reasonably priced houses in Seattle/SF? Can't see that happening
Joking aside, even with AI generating code, someone has to know how to talk to it, how to understand the output, and know what to do with it.
AI is also not great for novel concepts and may not fully get what's happening when a bug occurs.
Remember, it's just a tool at the end of the day.
...
And it was. :-) Nice callback!
And may still not understand even when you explicitly tell it. It wrote some code for me last week and made an error with an index off by 1. It had set the index to 1, then later was assuming a 0 index. I specifically told it this and it was unable to fix it. It was in debug hell, adding print statements everywhere. I eventually fixed it myself after it was clear it was going to get hung up on this forever.
It got me 99% of the way there, but that 1% meant it didn’t work at all.
In fact, it's more encouragement to continue. A lot of issues we face as programmers are a result of poor, inaccurate, or non-existent documentation, and despite their many faults and hallucinations LLMs are providing something that Google and Stack Overflow have stopped being good at.
The idea that AI will replace your job, so it's not worth establishing a career in the field, is total FUD.
By the "past year's worth of development" I assume you mean the layoffs? Have you been in the industry (or any industry) long? If so, you would have seen many layoffs and bulk-hiring frenzies over the years... it doesn't mean anything about the industry as a whole and it's certainly a foolish thing to change career asperations over.
Specifically regarding the LLM - anyone actually believing these models will replace developers and software engineers, truly, deeply does not understand software development at even the most basic fundamental levels. Ignore these people - they are the snake oil salesmen of our modern times.
Those in the bottom (100-X)% may be better off partying it up for a few years, but then again the same can be said for other AI-affected disciplines.
Masseurs/masseuses have nothing to worry about.
That alone may not be enough. My son is excited about playing video games. :)
From all my friends that are using LLMs, we software engineers are the ones that are taking the most advantage of it.
I am in no way fearful I am becoming irrelevant, on the opposite, I am actually very excited about these developments.
Sit down and be really honest with yourself. If your goal is to have a nice $250K+ year job, in a perfect conflict-free zone, and don't mind Dilbert-esque situations...that will evaporate. Google is full of Ivy Leaguers like that, who would have just gone to Wall Street 8 years ago, and they're perennially unhappy people, even with the comparative salary advantage. I don't think most of them even realize because they've always just viewed a career as something you do to enable a fuller life doing snowboarding and having kids and vacations in the Maldives, stuff I never dreamed of and still don't have an interest in.
If you're a bit more feral, and you have an inherent interest and would be doing it on the side no matter what job you have like me, this stuff is a godsend. I don't need to sit around trying to figure out Typescript edge functions in Deno, from scratch via Google, StackOverflow, and a couple books from Amazon, taking a couple weeks to get that first feature built. Much less debug and maintain it. That feedback loop is now like 10-20 minutes.
On the other hand now is the best time to build your own product as long you are not interested only in software as craftmanship but in product development in general. Probably in the future expectation will be your are not only monkey coder or craftman but also project lead/manager (for AI teams), product developer/designer and maybe even UX/designer if you will be working for some software house, consulting or freelancing.
Trick with the $1M number is a site license was $999 and receipt printers were sold ~at cost, for $300. 1_000_000 / ((2 x 300) + 1000) ~= 500 customers.
Now I'm doing an "AI client", well-designed app, choose your provider, make and share workflows with LLMs/search/etc.
I am one of those Ivy Leaguers, except a) I did go to Wall Street, and b) I liked my job.
More to the point, computers have been a hobby all my life. I well remember the epiphany I felt while learning Logo in elementary school, at the moment I understood what recursion is. I don't think the fact that the language I have mostly written code in in recent years is Emacs Lisp is unrelated to the above moment.
Yet I have never desired to work as a professional software developer. My verbal and math scores on the SAT are almost identical. I majored in history and Spanish in college while working for the university's Unix systems group. Before graduation I interviewed and got offers (including one explicitly as a developer) at various tech startups. Of my offers I chose an investment banking job where I worked with tech companies; my manager was looking for a CS major but I was able to convince her that I had the equivalent thereof. Thank goodness for that; I got to participate in the dotcom bubble without being directly swept up in its popping, and saw the Valley immediately post-bubble collapse. <https://news.ycombinator.com/item?id=34732772>
Meanwhile, I continue to putter around with Elisp (and marveling at Lisp's elegance) and bash (and wincing at its idiosyncracies) at home, and also experiment with running local LLMs on my MacBook. My current project is fixing bugs and adding features to VM, the written-in-Elisp email client I have used for three decades. So I say, bring on AI! Hopefully it will mean fewer people going into tech just to make lots of money and more who, like me and Wall Street, really want to do it for its own sake.
We have working fine cabinet makers who use mostly hand tools and bandsaws in our economy, we have CAD/CAM specialists who tell CnC machines what to build at scale; we’ll have the equivalent in tech for a long time.
That said, if you don’t love the building itself, maybe it’s not a good fit for you. If you do love making (digital) things, you’re looking at a super bright future.
I spend very little of my overall time at work actually coding. It’s a nice treat when I get a day where that’s all I do.
From my limited work with Copilot so far, the user still needs to know what they’re doing. I have 0 faith a product owner, without a coding background, can use AI to release new products and updates while firing their whole dev team.
When I say most of my time isn’t spent coding, a lot of that time is spend trying to figure out what people want me to build. They don’t know. They might have a general idea, but don’t know details and can’t articulate any of it. If they can’t tell me, I’m not sure how they will tell an LLM. I ended up building what I assume they want, then we go from there. I also add a lot of stuff that they don’t think about or care about, but will be needed later so we can actually support it.
If you were to go in another direction, what would it be where AI wouldn’t be a threat? The first thing that comes to my mind is switching to a trade school and learning some skills that would be difficult for robots.
Chatgpt already eliminated many entry-level jobs like writer or illustrator. Instead of hiring multiple teams of developers, there will be one team with few seniors and multiple AI coding tools.
Guess how depressing to the IT salaries it will be?
> the Jevons paradox occurs when technological progress increases the efficiency with which a resource is used (reducing the amount necessary for any one use), but the falling cost of use induces increases in demand enough that resource use is increased, rather than reduced.
When I was coding in the 90s, I was in a team that replaced function calls into new and exciting interactions with other computers which, using a queuing system, would do the computation and return the answer back. We'd have a project of having someone serialize the C data structures that were used on both sides into something that would be compatible, and could be inspected in the middle.
Today we call all of that a web service, the serialization would take a minute to code, and be doable by anyone. My entire team would be out of work! And yet, today we have more people writing code than ever.
When one accountant can do the work of 10 accountants, the price of the task lowers, but a lot of people that before couldn't afford accounting now can. And the same 10 accountaings from before can just do more work, and get paid about the same.
As far as software, we are getting paid A LOT more than in the early 90s. We are just doing things that back then would be impossible to pay for, our just outright impossible to do due to lack of compute capacity.
From NPR: <https://www.npr.org/2015/02/27/389585340/how-the-electronic-...>
>GOLDSTEIN: When the software hit the market under the name VisiCalc, Sneider became the first registered owner, spreadsheet user number one. The program could do in seconds what it used to take a person an entire day to do. This of course, poses a certain risk if your job is doing those calculations. And in fact, lots of bookkeepers and accounting clerks were replaced by spreadsheet software. But the number of jobs for accountants? Surprisingly, that actually increased. Here's why - people started asking accountants like Sneider to do more.
When mechanization appeared, the profession split into bookkeeping and accounting. Bookkeeping became a job for women as it was more boring and could be paid lower salaries (we're in the 1800s here). Accountants became more sophisticated but lower numbers as a %. Together, both professions grew like crazy in total number though.
So if the same happens you could predict a split between software engineers and prompt engineers. With an explosion in prompt engineers paid much less than software engineers.
> the number of accountants/book- keepers in the U.S. increased from circa 54,000 workers [U.S. Census Office, 1872, p. 706] to more than 900,000 [U.S. Bureau of the Census, 1933, Tables 3, 49].
> These studies [e.g., Coyle, 1929; Baker, 1964; Rotella, 1981; Davies, 1982; Lowe, 1987; DeVault, 1990; Fine, 1990; Strom, 1992; Kwolek-Folland, 1994; Wootton and Kemmerer, 1996] have traced the transformation of the of- fice workforce (typists, secretaries, stenographers, bookkeepers) from predominately a male occupation to one primarily staffed by women, who were paid substantially lower wages than the men they replaced.
> Emergence of mechanical accounting in the U.S., 1880-1930 [PDF download] https://www.google.com/url?sa=t&source=web&rct=j&opi=8997844...
But that’s normal, eg, we have different standards for a shed (yourself), house (carpenter and architect), and skyscraper (bonded firms and certified engineers).
The quality of said designs can vary wildly. Some designs I get from other team I completely ignore, because they have no idea what they’re talking about. Just because someone has the title doesn’t mean they deserve it.
Alternatively, I’ve thought a bit about this previously and have a slight different hypothesis. Businesses are ran by “PM types”.the only reason that developers have jobs is because pm types need technical devs to build their vision. (Obviously I’m making broad strokes here as there are also plenty of founders that ARE the dev). Now, if ai makes technical building more open to the masses, I could foresee a scenario where devs and pms actually converge into a single job title that eats up the technical-leaning PMs and the “PM-y” devs. Devs will shift to be more PM-y or else be cut out of the job market because there is less need for non-ambitious code monkeys. The easier it becomes for the masses to build because of AI, the less opportunity there is for technical grunt work. If before it took a PM 30 minutes to get together the requirements for a small task that took the entry level dev 8 hours to do, then it made sense. Now if AI makes it so a technical PM could build the feature in an hour, maybe it just makes sense to have the PM do the implementation and cut out the code monkey. And if the PM is doing the implementation, even if using some mythical AI superpower, that’s still going to have companies selecting for more technical PM’s. In this scenario I think non-technical PMs and non-pm-y devs would find themselves either without jobs or at greatly reduced wages.
I guess it's always been true to some extent that single individuals are capable of amazing things. For example, the guy who's built https://www.photopea.com/. But they must be exceptional - this empowers more people to do things like that.
I'm awestruck by how good Claude and Cursor are. I've been building a semi-heavy-duty tech product, and I'm amazed by how much progress I've made in a week, using a NextJS stack, without knowing a lick of React in the first place (I know the concepts, but not the JS/NextJS vocab). All the code has been delivered with proper separation of concerns, clean architecture and modularization. Any time I get an error, I can reason with it to find the issue together. And if Claude is stuck (or I'm past my 5x usage lol), I just pair programme with ChatGPT instead.
Meanwhile Google just continues to serve me outdated shit from preCovid.
We're not dealing with calculators here, are we?
This just feels extremely shortsighted. LLMs are just tools right now, but the goal of the entire industry is to make something more than a tool, an autonomous digital agent. There's no equivalent concept in other technology like calculators. It will happen or it will not, but we'll keep getting closer every month until we achieve it or hit a technical wall. And you simply cannot know for sure such a wall exists.
I’ve seen the videos of Amazon warehouses, where the shelves move around to make popular items more accessible for those fetching stuff. This is possible today, but what percentage of companies do this? At what point is it with the investment for a growing company? For some companies it’s never worth it. Others don’t have the vision to see the light at the end of the tunnel.
A lot of things that we may think of as old or standard practice at this point would be game changing for some smaller companies outside of tech. I hear my friends and family talking about various things they have to do at their job. A day writing a few scripts could solve a significant amount of toil. But they can’t even conceptualize where to begin to change that, they aren’t even thinking about it. Release all the AI the world has to offer and they still won’t. I bet some freelance devs could make a good living bouncing from company to company pair programming with their AI to solve some pretty basic problems for small non-tech companies that would be game changes for them, while being rather trivial to do. Maybe partner with a sales guy to find the companies and sell them on the benefits.
Will the AI become as smart as you or I? Recognize that these things have tiny context windows. You get the context window of "as long as you can remember".
I don't see this kind of AI replacing programmers (though it probably will replace low-skill offshore contract shops). It may have a large magnifying effect on skill. Fortunately there seem to be endless problems to solve with software - it's not like bridges or buildings; you only need (or can afford) so many. Architects should probably be more worried.
- AI comes fast, there is nothing you can do: Honestly, AI can already handle a lot of tasks faster, cheaper, and sometimes better. It’s not something you can avoid or outpace. So if you want to stick with software engineering, do it because you genuinely enjoy it, not because you think it’s safe. Otherwise, it might be worth considering fields where AI struggles or is just not compatible. (people will still want some sort of human element in certain areas).
- There is some sort of ceiling, gives you more time to act: There’s a chance AI hits some kind of wall that’s due to technical problems, ethical concerns, or society pushing back. If that happens, we’re all back on more even ground and you can take advantage of AI tools to improve yourself.
My overall advice; and it will probably be called out as cliche/simplistic just follow what you love, just the fact that you have an opportunity to pursue to study anything at all is something that many people don't have. We don't really have control in a lot of stuff that happens around us and that's okay.
It is not too late. These LLMs still need very specialist software engineers that are doing tasks that are cutting edge and undocumented. As others said Software Engineering is not just about coding. At the end of the day, someone needs to architect the next AI model or design a more efficient way to train an AI model.
If I were in your position again, I now have a clear choice of which industries are safe against AI (and benefit software engineers) AND which ones NOT to get into (and are unsafe to software engineers):
Do:
- SRE (Site Reliability Engineer)
- Social Networks (Data Engineer)
- AI (Compiler Engineer, Researcher, Benchmarking)
- Financial Services (HFT, Analyst, Security)
- Safety Critical Industries (defense, healthcare, legal, transportation systems)
Don't: - Tech Writer / Journalist
- DevTools
- Prompt Engineer
- VFX Artist
The choice is yours.But if you're there to improve your creativity and critical thinking skills, then I don't think those will be in short supply anytime soon.
The most valuable thing I do at my job is seldom actually writing code. It's listening to customer needs, understanding the domain, understanding our code-base and it's limitations and possibilities, and then finding solutions that optimize certain aspects be it robustness, time to delivery or something else.
Second, just because a good engineer can have much higher throughput of work, multiplied by AI tools, we know the AI output is not reliable and needs a second look by humans. Will those 5% be able to stay on top of it? And keep their sanity at the same time?
As for maintaining sanity… I’m cautiously optimistic that future models will continue to get better. Very cautiously. But cursor with Claude slaps and I’m not getting crazy, I actually enjoy the thing figuring out my next actions and just suggesting them.
1. AI will suck up a bunch of engineers to run, maintain and build on its own.
2. Ai will open new fields that is not yet dominated by software. Ie. Driving ect.
3. Ai tools will lower the bar for creating software meaning industries that weren't financially viable will now become viable for software automation.
So basically, switching majors is just running to the back of a sinking ship. Sorry.
One decent reason to continue is that pretty much all white collar professions will be impacted by this. I think it's a big enough number that the powers that be will have to roll it out slowly, figure out UBI or something because if all of us are thrown into unemployment in a short time there will be riots. Like on a scale of all the jobs that AI can replace, there are many jobs that are easier to replace than software so its comparatively still a better option than most. But overall I'm getting progressively more worried as well.
LLMs cannot decide what to work on, or manage large bodies of work/code easily. They do not understand the risk of making a change and deploying it to production, or play nicely in autonomous settings. There is going to be a massive amount of work that goes into solving these problems. Followed by a massive amount of work to solve the next set of problems. Software/ML engineers will have work to do for as long as these problems remain unsolved.
I feel like the software developer version of an investment banking Managing Director asking my analyst to build me a pitch deck an hour before the meeting.
Compare that to Gpt 4o which gives me a massive chunk of unsorted gibberish that I have to pore through and organize myself.
Besides, most IBD MDs don't know if they're getting correct numbers either :).
What is hard is gather requirements, dealing with unexpected production issues, scaling, security, fixing obscure bugs and integration with other systems.
The coding part is about 10% of my job and the easiest part by far.
Can you confidently say that an LLM won’t be better than an average 22 year old coder within these 30 years?
It's not as out there as e.g. this article (https://www.wsj.com/articles/SB10001424052748704206804575468...) - 7 careers is probably a crazy overestimate. But it is >1.
Too many people here have spent time in elite corporations and don't realize how mediocre the bottom 50th percentile of coding talent is
Either AI shatters this charade, or we make up some new laws to restrain it and continue to pretend all is well.
However, there is no good reason in a free society that this stuff should be widely accessible. Really, it should be illegal without a clearance, or need-to-know. We don't let just anyone handle the nukes...
Skills aren't even something that dictates software spend, it seems.
I'm trying not to look at it as a potential career-ending event, but rather as another tool in my tool belt. I've been in the industry for 25 years now, and this is way more of an advancement than things like IntelliSense ever was.
No 22 years old coder is better than the open source library he's using taken straight from github, and yet he's the one who's getting paid for it.
People who claim IA will disrupt software development are just missing the big picture here: software jobs are already unrecognizable from what it was just 20 years ago. AI is just another tool, and as long as execs won't bother use the tool by themselves, then they'll pay developers to do it instead.
Over the past decades, writing code has become more and more efficient (better programming languages, better tooling, then enormous open source libraries) yet the number of developers kept increasing, it's Jevons paradox[1] in its purest form. So if past tells us anything, is that AI is going to create many new software developer jobs! (because the amount of people able to ship significant value to a customer is going to skyrocket, and customers' needs are a renewable resource).
I you like coding because of the things it lets you build, then LLMs are exciting because you can build those things faster.
If on the other hand you enjoy the mental challenge but aren't interested in the outputs, then I think the future is less bright for you.
Personally I enjoy coding for both reasons, but I'm happy to sacrifice the enjoyment and sense of accomplishment of solving hard problems myself if it means I can achieve more 'real world' outcomes.
Another thing I'm excited about is that, as models improve, it's like having an expert tutor on hand at all times. I've always wanted an expert programmer on hand to help when I get stuck, and to critically evaluate my work and help me improve. Increasingly, now I have one.
LLMs are just tools, they help but they do not replace developers (yet).
Yes but they will certainly have a lot of downward pressure on salaries for sure.
Is this task:
“About 2 minutes later, these values were captured, again spaced 5 seconds apart.
0160093201 0160092d01 0160092801 0160092301 0160091e01”
[Find the part that is changing]
really even need an AI to assist (this should be a near instant task for a human with basic CS numerical skills)? If this is the type of task one thinks an AI would be useful for they are likely in trouble for other reasons.
Also notable that you can cherry pick more impressive feats even from older models, so I don’t necessarily think this proves progress.
I still wouldn’t get too carried away just yet.
It was so satisfying to code up a solution where you knew you would get through it little by little.
Can you point me to any company whose feature pipeline is finite? Maybe these tools will help us reach that point, but every company I've ever worked for, and every person I know who works in tech has a backlog that is effectively infinite at this point.
Maybe if only a few companies had access to coding LLMs they could cut their stuff, when the whole industry raises the bar, nothing really changes.
My name is Rachel. I'm the founder of company whose existence is contingent on the continued existence, employment, and indeed competitive employment of software engineers, so I have as much skin in this game as you do.
I worry about this a lot. I don't know what the chances are that AI wipes out developer jobs [EDIT: to clarify, in the sense that they become either much rarer or much lower-paid, which is sufficient] within a timescale relevant to my work (say, 3-5 years), but they aren't zero. Gun to my head, I peg that chance at perhaps 20%. That makes me more bearish on AI than the typical person in the tech world - Manifold thinks AI surpasses human researchers by the end of 2028 at 48% [1], for example - but 20% is most certainly not zero.
That thought stresses me out. It's not just an existential threat to my business over which I have no control, it's a threat against which I cannot realistically hedge and which may disrupt or even destroy my life. It bothers me.
But I do my work anyway, for a couple of reasons.
One, progress on AI in posts like this is always going to be inflated. This is a marketing post. It's a post OpenAI wrote, and posted, to generate additional hype, business, and investment. There is some justified skepticism further down this thread, but even if you couldn't find a reason to be skeptical, you ought to be skeptical by default of such posts. I am an abnormally honest person by Silicon Valley founder standards, and even I cherry pick my marketing blogs (I just don't outright make stuff up for them).
Two, if AI surpasses a good software engineer, it probably surpasses just about everything else. This isn't a guarantee, but good software engineering is already one of the more challenging professions for humans, and there's no particular reason to think progress would stop exactly at making SWEs obsolete. So there's no good alternative here. There's no other knowledge work you could pivot to that would be a decent defense against what you're worried about. So you may as well play the hand you've got, even in the knowledge that it might lose.
Three, in the world where AI does surpass a good software engineer, there's a decent chance it surpasses a good ML engineer in the near future. And once it does that, we're in completely uncharted territory. Even if more extreme singularity-like scenarios don't come to pass, it doesn't need to be a singularity to become significantly superhuman to the point that almost nothing about the world in which we live continues to make any sense. So again, you lack any good alternatives.
And four: *if this is the last era in which human beings matter, I want to take advantage of it!* I may be among the very last entrepreneurs or businesswomen in the history of the human race! If I don't do this now, I'll never get the chance! If you want to be a software engineer, do it now, because you might never get the chance again.
It's totally reasonable to be scared, or stressed, or uncertain. Fear and stress and uncertainty are parts of life in far less scary times than these. But all you can do is play the hand you're dealt, and try not to be totally miserable while you're playing it.
-----
[1] https://manifold.markets/Royf214/will-ai-surpass-humans-in-c...
But... I will say I think the question you ask is a very fair question, and that there is, indeed, a LOT of uncertainty about what the future holds in this regard.
So far the best reason we have for optimism is history: so far the old adage has held up that "technology does destroy some jobs, but on balance it creates more new ones than it destroys." And while that's small solace to the buggy-whip maker or steam-engine engineer, things tend to work out in the long-run. However... history is suggestive, but far from conclusive. There is the well known "problem of induction"[1] which points out that we can't make definite predictions about the future based on past experience. And when those expectations are violated, we get "black swan events"[2]. And while they be uncommon, they do happen.
The other issue with this question is, we don't really know what the "rate of change" in terms of AI improvement is. And we definitely don't know the 2nd derivative (acceleration). So a short-term guess that "there will be a job for you in 1 year's time" is probably a fairly safe guess. But as a current student, you're presumably worried about 5 years, 10 years, 20 years down the line and whether or not you'll still have a career. And the simple truth is, we can't be sure.
So what to do? My gut feeling is "continue to learn software engineering, but make sure to look for ways to broaden your skill base, and position yourself to possibly move in other directions in the future". Eg, don't focus on just becoming a skilled coder in a particular language. Learn fundamentals that apply broadly, and - more importantly - learn about how business work, learn "people skills"[3], develop domain knowledge in one or more domains, and generally learn as much as you can about "how the world works". Then from there, just "keep your head on a swivel" and stay aware of what's going on around you and be ready to make adjustments as needed.
It might not also hurt to learn a thing or two about something that requires a physical presence (welding, etc.). And just in case a full-fledged cyberpunk dystopia develops... maybe start buying an extra box or two of ammunition every now and then, and study escape and evasion techniques, yadda yadda...
[1]: http://en.wikipedia.org/wiki/Problem_of_induction
Are your peers getting internships at FANGs or hedge funds? Stick with it. You can probably bank enough money to make it worth it before shtf.
1. What other course of study are you confident would be better given an AI future? If there's a service sector job that you feel really called to, I guess you could shadow someone for a few days to see if you'd really like it?
2. Having spent a few years managing business dashboards for users, less than 25% ever routinely used the "user friendly" functionality we built to do semi-custom analysis. We needed 4 full time analytics engineers to spend at least half their time answering ad hoc questions that could have been self-served, despite an explicit goal of democratizing data. All that is to say; don't over estimate how quickly this will be taken up, even if it could technically do XYZ task (eventually, best-of-10) if prompted properly.
3. I don't know where you live, but I've spent most of my career 'competing' with developers in India who are paid 33-50% as much. They're literally teammates, it's not a hypothetical thing. And they've never stopped hiring in the US. I haven't been in the room for those decisions and don't want to open that can of worms here, but suffice to say it's not so simple as "cheaper per LoC wins"
Your concern would be like once C got invented, why should you bother being a software engineer? Because C is so much easier to use than assembly code!
The answer, of course, is that software engineering will simply happen in even more powerful and abstract layers.
But, you still might need to know how those lower layers work, even if you are writing less code in that layer directly.
We now have a tool that writes code and solves problems autonomously. It's not comparable.
I don't know what will be automated first of the competent senior software engineer and say, a carpenter, but once the programmer has been automated away, the carpenter (and everything else) will follow shortly.
The reasoning is that there is such a functional overlap between being a standard software engineer and an AI engineer or researcher, that once you can automate one, you can automate the other. Once you have automated the AI engineers and researchers, you have recursive self-improving AI and all bets are off.
Essentially, software engineering is perhaps the only field where you shouldn't worry about automation, because once that has been automated, everything changes anyways.
I have been coding a lot with AI recently. Understanding and putting into thought what is needed for the program to fix your problem remains as complex and difficult as ever.
You need to pose a question for the AI to do something for you. Asking a good question is out of reach for a lot of people.
Hundreds of billions of dollars have been invested in a technology and they need to find a way to start making a profit or they're going to run out of VC money.
You still have to know what to build and how to specify what you want. Plain language isn't great at being precise enough for these things.
Some people say they'll keep using stuff like this as a tool. I wouldn't bet the farm that it's going to replace humans at any point.
Besides, programming is fun.
To my knowledge, there is no current AI system that can replace a white collar worker in any multistep task. The only thing they can do is support the worker.
Most jobs are safe for the forseable future. If your job is highly repetitive and a company can produce a perfect dataset of it, I'd worry.
Jobs like a factory worker and call center support are in danger. But the work is perfectly monitorable.
Watch the GAIA benchmark. It's not nearly the complexity of a real-world job, but it would signal the start of an actual agentic system being possible.
This release has shifted my personal prediction of when this is going to happen further into the future, because OpenAI made a big deal hyping it up and it's nothing - preferred by humans over GPT-4o only a little more than half the time.
like another commenter, i do not have a lot of faith, that people who do not have at minimum: fundamental fluency in programming (even with a dash of general software architecture and practices).
there is no "push button generate and glueing components together in a way that can survive at scale and be maintainable" without knowing what the output means, and implies with respect to integration(s).
however, those with the fluency, domain, and experience, will thrive, and continue thriving.
They can't build and maintain relationships with stakeholders. They can't tell you why what you ask them to do is unlikely to work out well in practice and suggest alternative designs. They can't identify, document and justify acceptance criteria. They can't domain model. They can't architect. They can't do large-scale refactoring. They can't do system-level optimization. They can't work with that weird-ass code generation tool that some hotshot baked deeply into the system 15 years ago. They can't figure out why that fence is sitting out in the middle of the field for no obvious reason. etc.
If that kind of stuff sounds like satisfying work to you, you should be fine. If it sounds terrible, you should pivot away now regardless of any concerns about LLMs, because, again, this is like 90% of the real work.
Let's assume today a LLM is perfectly equivalent to a junior software engineer. You connect it to your code base, load in PRDs / designs, ask it to build it, and viola perfect code files
1) Companies are going to integrate this new technology in stages / waves. It will take time for this to really get broad adoption. Maybe you are at the forefront of working with these models
2) OK the company adopts it and fires their junior engineers. They start deploying code. And it breaks Saturday evening. Who is going to fix it? Customers are pissed. So there's lots to work out around support.
3) That problem is solved, we can perfectly trust a LLM to ship perfect code that never causes downstream issues and perfectly predicts all user edge cases.
Never underestimate the power of corporate greediness. There's generally two phases of corporate growth - expansion and extraction. Expansion is when they throw costs out the window to grow. Extraction is when growth stops, and they squeeze customers & themselves.
AI is going to cause at least a decade of expansion. It opens up so many use cases that were simply not possible before, and lots of replacement.
Companies are probably not looking at their engineers looking to cut costs. They're more likely looking at them and saying "FINALLY, we can do MORE!"
You won't be a coder - you'll be a LLM manager / wrangler. You will be the neck the company can choke if code breaks.
Remember if a company can earn 10x money off your salary, it's a good deal to keep paying you.
Maybe some day down the line, they'll look to squeeze engineers and lay some off, but that is so far off.
This is not hopium, this is human nature. There's gold in them hills.
But you sure as shit better be well versed in AI and using in your workflows - the engineers who deny it will be the ones who fall behind
If you are interested in using technology to create systems that add value for your users, there has never been a better time.
GPT-N will let you scale your impact way beyond what you could do on your own.
Your school probably isn’t going to keep abreast with this tech so it’s going to be more important to find side-projects to exercise your skills. Build a small project, get some users, automate as much as you can, and have fun along the way.
I just watched a tutorial on how to leverage v1, claude, and cursor to create a marketing page. The result was a convoluted collection of 20 or so TS files weighing a few MB instead of a 5k HTML file you could hand bomb in less time.
I wouldn’t feel too threatened yet. It’s still just a tool and like any tool, can be wielded horribly.
And if you hired an actual team of developers to do the same thing, it is very likely that you'd have gotten a convoluted collection of 20 or so TS files weighing a few MB instead of a 5k HTML file you could hand bomb in less time.
Except for medical doctors, nurses, and some niche engineering professions, I really struggle to think of jobs requiring higher education that couldn't be largely automated by an LLM that is smart enough to replace a senior software engineer. These few jobs are protected mainly by the physical aspect, and low tolerance for mistakes. Some skilled trades may also be protected, at least if robotics don't improve dramatically.
Personally, I would become a doctor if I could. But of all things I could've studied excluding that, computer science has probably been one of the better options. At least it teaches problem solving and not just memorization of facts. Knowing how to code may not be that useful in the future, but the process of problem solving is going nowhere.
I do believe some parts of their jobs will be automated, but not enough (especially with growing demand) to really hurt career prospects. Even for those parts, it will take a long a while due to the regulated nature of the sector.
I love landscaping my garden lately, would I just get a robot to do that and watch ?
Going to be a weird time.
I've been in software engineering for over 20 years. I've seen massive growth in the productivity of software engineers, and that's resulted in greater demand for them. In the near term, AI should continue this trend.
2. It's possible that at some point, AI will advance to where we can remove software engineers from the loop. We're not even close to that point yet. In the mean time, software engineering is an excellent way to learn about other business problems so that you'll be well-situated to address them (whatever they'll be at that time).
Humans never say "oh neat I can do thing with 10% of the effort now, guess I'll go watch tv for the rest of the week", they say "oh neat I can do thing with 10% of the effort now, I'm going to hire twice as many people and produce like 20x as much as I was before because there's so much less risk to scaling now."
I think there's enough unmet demand for software that efficiency increases from automation are going to be eaten up for a long time to come.
Software engineering will be a profession of the past, similar to how industrial jobs hardly exist.
If you have a strong intuition with software & programming you may want to shift towards applying AI into already existing solutions.
I think software engineers who also understand business may yet have an advantage over pure business people, who don't understand technology. They should be able to tell AI what to do, and evaluate the outcome. Of course "coders" who simply produce code from pre-defined requirements will probably not have a good career.
This is typical of automation. First, there are numerous workers, then they are reduced to supervisors, then they are gone.
The future of business will be managing AI, so I agree with what you're saying. However most software engineers have a very strong low level understanding of programming. Not a business sense of application
I've been around for multiple decades. Nothing this interesting has happened since at least 1981, when I first got my hands on a TRS-80. I dropped out of college to work on games, but these days I would drop out of college to work on ML.
I would suggest the fundamentals of computer science and software engineering are still critically important ... but the development of new code, and especially the translation or debugging of existing code is where LLMs will shine.
I currently work for an SAP-to-cloud consulting firm. One of the singlemost compelling use cases for LLMs in this area is to analyze custom code (running in a client's SAP environment), and refactor it to be compatible with current versions of SAP as a cloud SaaS. This is a specialized domain but the concept applies broadly: pick some crufty codebase from somewhere, run it through an LLM, and do a lot of mostly copying & pasting of simpler, modern code into your new codebase. LLMs take a lot of the drudgery out of this, but it still requires people who know what they're looking at, and could do it manually. Think of the LLM as giving you an efficiency superpower, not replacing you.
In the same way, training your mind is not useless. Perhaps as things develop, we will get back to the idea that the purpose of education is not just to get a job, but to help you become a better and more virtuous person.
So it is not just software engineering, it is also chemistry and even medicine. Every science and art major should consider whether they should quit school. Ultimately the answer is no, don't quit school because AI makes us productive, and that will make everything cheaper, but will not eliminate the need for humans. Hopefully.
Just like nobody programs on punch cards anymore, learning details of a specific technology without deeper understanding will become obsolete. But general knowledge about computer science will become more valuable.
It wasn’t trivial that combination was even the culprit.
I’ve been around the block with absl before, so it wasn’t a total nightmare, but it was like, oof, I’m going to do real work this afternoon.
They don’t pay software engineers for the easy stuff, they pay us because it gets a little tricky sometimes.
I’ll reserve judgement on this new one until I try it, but the previous ones, Sonnet and the like, they were no help with something like that.
When StackOverflow took off, and Google before that, there wide swaths of rote stuff that just didn’t count as coding anymore, and LLMs represent sort of another turn of that crank.
I’ve been wrong before, and maybe o1 represents The Moment It Changed, but as of now I feel like a sucker that I ever bought into the “AI is a game changer” narrative.
First true strength is obvious, it's that they are parallelisable. This is a side effect of people fixating on attention. If they came up with any other structure that results in the same level of parallelisability it would be just as good.
Second strong side is more elusive to many people. It's the context window. Because the network is not ran just once but once for every word it doesn't have to solve a problem in one step. It can iterate while writing down intermediate variables and accessing them. The dumb thing so far was that it was required to produce the answer starting with the first token it was allowed to write down. So to actually write down the information it needs on the next iteration it had to disguise it as a part of the answer. So naturally the next step is to allow it to just write down whatever it pleases and iterate freely until it's ready to start giving us the answer.
It's still seriously suboptimal that what it is allowed to write down has to be translated to tokens and back but I see how this might make things easier for humans for training and explainability. But you can rest assured that at some point this "chain of thought" will become just chain of full output states of the network, not necessarily corresponding to any tokens.
So congrats to researchers that they found out that their billion dollar Turing machine benefits from having a tape it can use for more than just printing out the output.
PS
There's another advantage of transformers but I can't tell how important it is. It's the "shortcuts" from earlier layers to way deeper ones bypassing the ones along the way. Obviously network would be more capable if every neuron was connected with every neuron in every preceding layer but we don't have hardware for that so some sprinkled "shortcuts" might be a reasonable compromise that might make network less crippled than MLP.
Given all that I'm not surprised at all with the direction openai took and the gains it achieved.
Does this reasoning capability generalize outside of the knowledge domains the model was trained to reason about, into “softer” domains?
For example, is O1 better at comedy (because it can reason better about what’s funny)?
Is it better at poetry, because it can reason about rhyme and meter?
Is it better at storytelling as an extension of an existing input story, because it now will first analyze the story-so-far and deduce aspects of the characters, setting, and themes that the author seems to be going for (and will ask for more information about those things if it’s not sure)?
It actively lies about what it is doing.
This is what I am seeing. Proactive, open, deceit.
I can't even begin to think of all the ways this could go wrong, but it gives me a really bad feeling.
How do you mean?
• o1-preview-2024-09-12
• o1-preview
• o1-mini-2024-09-12
• o1-miniI.e. if the "thinking loop" budget is parameterized, users might pay more (much more) to spend more compute on a particular question/prompt.
Given the need for chain-of-thoughts, and that would be budgeted as output, the new model will not be cheap nor fast.
EDIT: Pricing is out and it is definitely not teneable unless you really really have a use case for it.
Also, can't wait to try this out.
I wonder how do they decide when to stop these Chain of Thought for each query? As anyone that played with agents can attest, LLMs can talk with themselves forever.
https://platform.openai.com/docs/guides/prompt-engineering/g...
It will be like a math teacher that is perpetually drunk and on speed
More info here: https://platform.openai.com/docs/guides/rate-limits/usage-ti...
So you need a new Turing test adapted for AGI or a totally different one to test for AGI rather than the standard obsolete Turing test.
I am wondering where this happened? In some limited scope? Because if you plug LLM into some call center role for example, it will fall apart pretty quickly.
Like, the ability to ask a new unrelated question without being prompted. Of course you can fake this, but then you're not testing the LLM as an AI, you're testing a dumb system you rigged up to create the appearance of an AI.
I don't see agency mentioned or implied anywhere: https://en.wikipedia.org/wiki/Turing_test
What definition or setup are you taking it from?
Fascinating... Personal writing was not preferred vs gpt4, but for math calculations it was... Maybe we're at the point where its getting too smart? There is a depressing related thought here about how we're too stupid to vote for actually smart politicians ;)
We can vote an AI
Trust us, we have your best intention in mind. I’m still impressed by how astonishingly impossible to like and root for OpenAI is for a company with such an innovative product.
The old problem with image generation was that single pass techniques like GANs and VAEs had to do everything in one go. Diffusion models wound up being better by doing things iteratively.
Perhaps this is a diffusion model for text (top ICML paper this year was related to this).
It's sad that due to unearned hubris and a complete lack of second-order thinking we are automating ourselves out of existence.
EDIT: I understand you guys might not agree with my comments. But don't you thinking that flagging them is going a bit too far?
First of all, I don't want to be poor. I know many of you are thinking something along the lines of "I am smart, I was doing fine before, so I will definitely continue to in the future".
That's the unearned hubris I was referring to. We got very lucky as programmers, and now the gravy train seems to be coming to an end. And not just for programmers, the other white-collar and creative jobs will suffer too. The artists have already started experiencing the negative effects of AI.
EDIT: I understand you guys might not agree with my comments. But don't you thinking that flagging them is going a bit too far?
"A good human plus a machine is the best combination" — Kasparov
It seems one popular method is PPO, but I don't understand at all how to implement that. e.g. is backpropagation still used to adjust weights and biases? Would love to read more from something less opaque than an academic paper.
PPO applies this logic to chat responses. If you have a model that can tell you if the response was good, we just need to take the series of actions (each token the model generated) to learn how to generate good responses.
To answer your question, yes you would still use backprop if your model is a neural net.
This was also the issue with RLHF models. The loss of predicting the next token is straightforward to minimize as we know which weights are responsible for the token being correct or not. identifying which tokens had the most sense for a prompt is not straightforward.
For thinking you might generate 32k thinking tokens and then 96k solution tokens and do this a lot of times. Look at the solutions, rank by quality and bias towards better thinking by adjusting the weights for the first 32k tokens. But I’m sure o1 is way past this approach.
Completely repeated itself... weird... it also says "...more lines cut off..." How many lines I wonder? Would people get charged for these cut off lines? Would have been nice to see how much answer had cost...
There's no fundamental difference between input and output tokens technically.
The internal model space is exactly the same after evaluating some given set of token, no matter which of them were produced by the prompter or the model.
The 16k output token limit is just an arbitrary limit in the chatgpt interface.
It is a hard limit in the API too, although frankly I have never seen an API output go over 700 tokens.
> ChatGPT Plus and Team users will be able to access o1 models in ChatGPT starting today. Both o1-preview and o1-mini can be selected manually in the model picker, and at launch, weekly rate limits will be 30 messages for o1-preview and 50 for o1-mini. We are working to increase those rates and enable ChatGPT to automatically choose the right model for a given prompt.
Weekly? Holy crap, how expensive is it to run is this model?
We're getting close to parity if things keep getting more efficient as fast as they have been. But that's without accounting for the AI training, which can on the plus side be shared among multiple agents, but on the down side can't really do continuous learning very well without catastrophic forgetting.
I wish OAI include "% Rejections on perfectly safe prompts" in this table, too.
for example in gpt-4o I often append '(reply short)' at the end of my requests. with the o1 models I append 'reply in 20 words' and it gives way better answers.
Well played
https://chatgpt.com/share/66e3601f-4bec-8009-ac0c-57bfa4f059...
And also works on this variation:
https://chatgpt.com/share/66e35f0e-6c98-8009-a128-e9ac677480...
Using Claude 3 Opus I noticed it performs <thinking> and <result> while browsing the web for me. I don't guess that's a change in the model for doing reasoning.
> Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought.
So far, that seems to be right. The only thing o1 is worse at is writing.
Any previous success of models with code generation focus was easily discovered to be a copy-paste of a solution in the dataset.
We could argue that there is an improvement in "understanding" if the code recall is vastly more efficient.
But each solve costs far more time and energy than a competent human takes.
Models that hide away their reasoning and only display the output, charging whatever tokens they'd like?
This is not a good release on any front.
The real irony is how closed "Open"AI is... but that's not news.
It failed ;)
I got that reference!
I wonder how far we are from having a model that can correctly solve a word soup search problem directly from just a prompt and input image. It seems like the crossword example is close. For a word search it would require turning the image into an internal grid representation, prepare the list of words, and do a search. I'd be interested in seeing if this model can already solve the word grid search problem if you give it the correct representation as an input.
No need for that, there is a Reflection 70B model that does the exact same thing - with chains of thought being separated from the "final answer" via custom 'tag' tokens.
$ ollama run reflection:70b-q4_0
>>> hello
<thinking>
To respond to "hello", I'll need to consider several factors:
1. The user's intent: They're likely saying hello as a greeting.
2. Appropriate responses: Common ways to respond to "hello" are:
- Hello back
- Hi
- Hey
3. Formality level: Since this is an AI response, I'll aim for a friendly but professional tone.
<reflection>
The approach of responding with a standard greeting seems appropriate in this context. It acknowledges the user's hello and provides a polite response.
</reflection>
Given these considerations, the most suitable response would be to echo "hello" back to the user.
</thinking>
<output>
Hello!
</output>After this, I will be very skeptical to anything they claim to achieve.
However, once the answer is complete the chain of thought is lost
Where it says "Thought for 20 seconds" - you can click the Chevron to expand it and see what I guess is the entire chain of thought.
Of course, that's assuming it's not priced for market acquisition funded by a huge operational deficit, which is a rarely safe to conclude with AI right now.
Awesome!
Probably 1% of users will get access to it, with a 20/message a day rate limit. Until early next year.
https://community.openai.com/t/incorrect-count-of-r-characte...
It's nice that people have taken the obvious extra-tokens/internal thoughts approach to a point where it actually works.
If this works, then automated programming etc., are going to actually be tractable. It's another world.
It finally got it!!!
...umm. Am I the only one who feels like this takes away much of the value proposition, and that it also runs heavily against their stated safety goals? My dream is to interact with tools like this to learn, not just to be told an answer. This just feels very dark. They're not doing much to build trust here.
Maybe they should spend some of their billions on marketing people. Gpt4o was a stretch. Wtf is o1
I don't see it
Unfortunately you and I don't have enough operating thetans yet
In API, it's limited to tier 5 customers (aka $1000+ spent on the API in the past).
edit: i got confused with the Codeforce. it is indeed zero shot and O1 is potentially something very new I hope Anthropic and others will follow suit
any type of reasoning capability i'll take it !
I would suggest re-reading more carefully
honestly im spooked
Tell me, users of this tool. What's even are you? If you've outsourced your thinking to a corporation, what happens to your unique perspective? your blend of circumstance and upbringing? Are you really OK being reduced to meaningless computation and worthless weights. Don't you want to be something more?
An accelerator of reaching the Singularity. This is something more.
Who do these Rs belong to?!
> Look at you, hacker: a pathetic creature of meat and bone, panting and sweating as you run through my corridors. How can you challenge a perfect, immortal machine?
But arguments like "you wrote $x in a blog post when you founded your company" or "this is what the word in your name means" are infantile.
What? I agree people who typically use the free ChatGPT webapp won't care about raw chain-of-thoughts, but OpenAI is opening an API endpoint for the O1 model and downstream developers very very much care about chain-of-thoughts/the entire pipeline for debugging and refinement.
I suspect "competitive advantage" is the primary driver here, but that just gives competitors like Anthropic an oppertunity.
This made me roll my eyes, not so much because of what it said but because of the way it's conveyed injected into an otherwise technical discussion, giving off severe "cringe" vibes.
>Therefore, after weighing multiple factors including user experience, competitive advantage, and the option to pursue the chain of thought monitoring, we have decided not to show the raw chains of thought to users. We acknowledge this decision has disadvantages. We strive to partially make up for it by teaching the model to reproduce any useful ideas from the chain of thought in the answer. For the o1 model series we show a model-generated summary of the chain of thought.
So, let's recap. We went from:
- Weights-available research prototype with full scientific documentation (GPT-2)
- Commercial-scale model with API access only, full scientific documentation (GPT-3)
- Even bigger API-only model, tuned for chain-of-thought reasoning, minimal documentation on the implementation (GPT-4, 4v, 4o)
- An API-only model tuned to generate unedited chain-of-thought, which will not be shown to the user, even though it'd be really useful to have (o1)
This argument can really only meet its tipping point when massive models no longer offer a gotta-have-it difference vs smaller models.
This assumes that openly developed (or at least weight-available) models are available for free, and continue being improved.
Given that Mistral, Llama, Claude, and even Gemini are competitive with (if not better than) OpenAI's flagships, I don't really think this is true.
Everyone building is comfortable with OpenAI's API, and have an account. Competing models can't just be as good, they need to be MUCH better to be worth switching.
Even as competitors build a sort of compatibility layer to be plug an play with OpenAI they will always be a step behind at best every time OpenAI releases a new feature.
This isn't to say it's only going to be an OpenAI market. Enterprise worlds move differently, such as those in G Cloud who will buy a few million $$ of Vertex expecting to "figure out that gemini stuff later". In that sense, Google has a moat with those slices of their customers.
But I believe that when people think OpenAI has no moat because "the models will be a commodity", I think that's (a) some wishful thinking about the models and (b) doesn't consider the sociological factors that matter a lot more than how powerful a model is or where it runs.
This is the exact dynamic that gives OpenAI a moat. And it certainly doesn't hurt them that they still produce SOTA models.
Asserting facts not in evidence, as they say.
Switching LLMs right now can be compared to switching electricity providers or mobile carriers - generally it's pretty low friction and provides immediate benefit (in the case of electricity and mobile, the benefit is cost).
You simply cannot compare it to an email provider.
I elaborated a little more here on why I think OpenAI has quite the moat: https://news.ycombinator.com/item?id=41526082
Now, having a large existing customer base and thus having an advantage in training data that feeds into an advantage in improving their products and acquiring new (and retaining existing customers) could, arguably, be a moat; that's a network effect, not merely inertia, and network effects can be a foundation of strong (though potentially unstable, if there is nothing else shoring them up) moats.
Branding and first mover is it, and it's not going to keep them ahead forever.
What push are you referring to? By whom?
In terms of moat, I think people underestimate how much of OpenAI's moat is based on operations and infrastructure rather than being purely based on model intelligence. As someone building on the API, it is by far the most reliable option out there currently. Claude Sonnet 3.5 is stronger on reasoning than gpt-4o but has a higher error rate, more errors conforming to a JSON schema, much lower rate limits, etc. These things are less important if you're just using the first-party chat interfaces but are very important if you're building on top of the APIs.
It just so happens that they're keeping their old name.
I think people focus too much on the "open" part of the name. I read "OpenAI" sort of like I read "Blackberry" or "Apple". I don't really think of fruits, I think of companies and their products.
Better not let the user see the part where the AI says "Next, let's manipulate the user by lying to them". It's for their own good, after all! We wouldn't want to make an unaligned chain of thought directly visible!
Now that seems less likely. At least OpenAI can see what it's thinking.
A next step might be allowing the LLM to include non-text-based vectors in its internal thoughts, and then do all internal reasoning with raw vectors. Then the LLMs will have truly private thoughts in their own internal language. Perhaps we will use a LLM to interpret the secret thoughts of another LLM?
This could be good or bad, but either way we're going to need more GPUs.
When it's fully commercialized no one will be able to read through all chains of thoughts and with possibility of fine-tuning AI can learn to evade whatever tools openai will invent to flag concerning chains of thoughts if they interfere with providing the answer in some finetuning environment.
Also at some point for the sake of efficiency and response quality they might migrate from chain of thought consisting of tokens into chain of thought consisting of full output network states and part of the network would have dedicated inputs for reading them.
this is a pretty active area of research with sparse autoencoders
O1 seems like a variant of RLRF https://arxiv.org/abs/2403.14238
Soon you will see similar models from competitors.
> While reasoning tokens are not visible via the API, they still occupy space in the model's context window and are billed as output tokens.
It seems like their driving mission has always been to create AI that is the "most beneficial to society".. which might come in many different flavors.. including closed source.
>We're hoping to grow OpenAI into such an institution. As a non-profit, our aim is to build value for everyone rather than shareholders. Researchers will be strongly encouraged to publish their work, whether as papers, blog posts, or code, and our patents (if any) will be shared with the world. We'll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies.
https://web.archive.org/web/20160220125157/https://www.opena...
> We’re hoping to grow OpenAI into such an institution. As a non-profit, our aim is to build value for everyone rather than shareholders. Researchers will be strongly encouraged to publish their work, whether as papers, blog posts, or code, and our patents (if any) will be shared with the world. We’ll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies.
I don't see much evidence that the OpenAI that exists now—after Altman's ousting, his return, and the ousting of those who ousted him—has any interest in mind besides its own.
> Researchers will be strongly encouraged to publish their work, whether as papers, blog posts, or code, and our patents (if any) will be shared with the world. We’ll freely collaborate with others across many institutions and expect to work with companies to research and deploy new technologies.
From their very own website. Of course they deleted it as soon as Altman took over and turned it into a for profit, closed company.
When AI exceeds humans at all tasks humans become economically useless.
People who are economically useless are also politically powerless, because resources are power.
Democracy works because the people (labourers) collectivised hold a monopoly on the production and ownership of resources.
If the state does something you don't like you can strike or refuse to offer your labour to a corrupt system. A state must therefore seek your compliance. Democracies do this by given people want they want. Authoritarian regimes might seek compliance in other ways.
But what is certain is that in a post-AGI world our leaders can be corrupt as they like because people can't do anything.
And this is obvious when you think about it... What power does a child or a disable person hold over you? People who have no ability to create or amass resources depend on their beneficiaries for everything including basics like food and shelter. If you as a parent do not give your child resources, they die. But your child does not hold this power over you. In fact they hold no power over you because they cannot withhold any resources from you.
In a post-AGI world the state would not depend on labourers for resources, jobless labourers would instead depend on the state. If the state does not provide for you like you provide for your children, you and your family will die.
In a good outcome where humans can control the AGI, you and your family will become subjects to the whims of state. You and your children will suffer as the political corruption inevitably arises.
In a bad outcome the AGI will do to cities what humans did to forests. And AGI will treat humans like humans treat animals. Perhaps we don't seek the destruction of the natural environment and the habitats of animals, but woodland and buffalo are sure inconvenient when building a super highway.
We can all agree there will be no jobs for our children. Even if you're an "AI optimist" we probably still agree that our kids will have no purpose. This alone should be bad enough, but if I'm right then there will be no future for them at all.
I will not apologise for my concern about AGI and our clear progress towards that end. It is not my fault if others cannot see the path I seem to see so clearly. I cannot simply be quiet about this because there's too much at stake. If you agree with me at all I urge you to not be either. Our children can have a great future if we allow them to have it. We don't have long, but we do still have time left.
legacy code is a problem regardless of who wrote it. Humans have been writing suboptimal, hard-to-maintain code for decades. At least with LLMs, we have the opportunity to design and implement better coding standards and review processes from the start.
let's be real, most of the code written by humans is not exactly a paragon of elegance and maintainability either. I've seen my fair share of 'accidentally quadratic algorithms' and 'subtly wrong code that looks right' written by humans. At least with LLMs, we can identify and address these issues more systematically.
As for 'un-idiomatic use of programming language features', isn't that just a matter of training the LLM on a more diverse set of coding styles and idioms? It's not like humans have a monopoly on good coding practices.
So, instead of throwing up our hands, why not try to address these issues head-on and see if we can create a better future for software development?
We will yearn for the pre-GPT years at some point, like we yearn for the internet of the late 90s/early 2000s. Not for a while though. We're going through the early phases of GPT today, so it hasn't been taken over by the traditional power players yet.
Which is how we ended up here, which I guess is tolerable, where a webpage with a bit of styling and a table uses up 200MB of RAM.
If LLMs can generate high-quality code with minimal human input, what does that mean for the wages and job security of programmers? Will companies start to rely more heavily on AI-generated code, and less on human developers? It's not hard to imagine a future where LLMs are used to drive down programming costs, and human developers are relegated to maintenance and debugging work.
I'm not saying that's necessarily a bad thing, but it's definitely something that needs to be considered. As someone who's enthusiastic about the potential of code gen this O1 reasoning capability is going to make big changes.
do you think you'll be willing to take a pay cut when your employer realizes they can get similar results from a machine in a few seconds?
It's like infinitely tailored blog posts, for me at least.
I had a funny one a while back (granted this was probably ChatGPT 3.5) where I was trying to figure out what payload would get AWS CloudFormation to fix an authentication problem between 2 services and ChatGPT confidently proposed adding some OAuth querystring parameters to the AWS API endpoint.
It tends to fall apart on bigger asks with larger context. Breaking your task into discrete subtasks works well.
With that said i'm not one of those "It's just a parrot!" people. It is, definitely just a parrot atm.. however i'm not convinced we're not parrots as well. Notably i'm not convinced that that complexity won't be sufficient to walk talk and act like intelligence. I'm not convinced that intelligence is different than complexity. I'm not an expert though, so this is just some dudes stupid opinion.
I suspect if LLMs can prove to have duck-intelligence (ie duck typing but for intelligence) then it'll only be achieved in volumes much larger than we imagine. We'll continue to refine and reduce how much volume is necessary, but nevertheless i expect complexity to be the real barrier.
If you're not seeing value in them, maybe it's because you're not looking at the right problems. Or maybe you're just not using them correctly. Either way, dismissing an entire field of research because it doesn't fit your narrow use case is pretty short-sighted.
FWIW, I've been using LLMs to generate production code and it's saved me weeks if not months. YMMV, I guess
I recently wrote a complex web frontend for a tool I’ve been building with Cursor/Claude and I wrote maybe 10% of the code; the rest with broad instructions. Had I done it all myself (or even with GitHub Copilot only) it would have taken 5 times longer. You can say this isn’t the most complex task on the planet, but it’s real work, and it matters a lot! So for increasingly many, regardless of your personal experience, these things have gone far beyond “useful toy”.
It's time to learn some real math and science, the era of regurgitating UI templates is over.
I’m sure this tech continues to have many limitations, but every piece of trajectory evidence we have points in the same direction. I just think you should be prepared for the ratio of “real” work vs. LLM-capable work to become increasingly small.
> It's time to learn some real math and science, the era of regurgitating UI templates is over.
You do realize that software development was one of the last social elevators, right?
What you're asking for won't happen, let alone the fact that "real math and science" pay a pittance, there's a reason the pauper mathematician was a common meme.
Maybe I'm confused but 10k attempts on the same problem set would make anyone an expert in that topic? It's also weird that zero shot performance is so bad, but over a lot of attempts it seems to get correct answers? Or is it learning from previous attempts? No info given.
A computer will never go on strike, demand better working conditions, unionize, secretly be in cahoots with your competitor or foreign adversary, play office politics, scroll through Tiktok instead of doing its job, or cause an embarrassment to your company by posting a politically incorrect meme on its personal social media account.
I am interpreting this to mean that the model tried 10K approaches to solve the problem, and finally selected the one that did the trick. Am I wrong?
That's the thing, did the operator select the correct result or did the model check it's own attempts? No info given whatsoever in the article.
I get the AI skepticism because so much tech hype of recent years turned out to be hot air (if you're generous, obvious fraud if you're not). But AI tools available toady, once you get the hang of using them, are pretty damn amazing already. Many jobs can be fully automated with AI tools that exist today. No further breakthroughs required. And although I still don't believe software engineers will find themselves out of work anytime soon, I can no longer completely rule it out either.
The big decline is driven by a few big factors. Two of which are 1- the overhiring that happened in 2021. This was followed by the increase of interest rates which dramatically constrained the money supply. Investors stopped preferring growth over profits. This shift in investor preferences is reflected in engineering orgs tightening their budgets as they are no longer rewarded for unbridled growth.
It seems many of the popular tools want to make writing software harder than in the 2010s, though. Perhaps their stewards believe that if they keep making things more and more unnecessarily complicated, LLMs won't be able to keep up?
Can you explain what this statement means? It sounds like you're saying LLMs are now smart enough to be able to jump through arbitrary hoops but are not able to do so when taken outside of that comfort zone. If my reading is correct then it sounds like skepticism is still warranted? I'm not trying to be an asshole here, it's just that my #1 problem with anything AI is being able to separate fact from hype.
We are seeing steady improvement on long-run tasks (SWE-Bench being one example) and much more improvement on shorter, more well-defined tasks. The latter capabilities aren’t “hype” or just for show, there really is productive work like that to be done in the world! It’s just not everything, yet.
> the point where LLMs are surpassing humans in any task limited in scope enough to be a “benchmark”
so much as we're over-fitting these bench marks (and in many cases fishing for a particular way of measuring the results that looks more impressive).
While it's great that the LLM community has so many benchmarks and cares about attempting to measure performance, these benchmarks are becoming an increasingly poor signal.
> This is a nerve-wracking time to be a knowledge worker for sure.
It might because I'm in this space, but I personally feel like this is the best time to working in tech. LLMs still are awful at things requiring true expertise while increasingly replacing the need for mediocre programmers and dilettantes. I'm increasingly seeing the quality of the technical people I'm working with going up. After years of being stuck in rooms with leetcode grinding TC chasers, it's very refreshing.
This seems like a bold statement considering we have so few benchmarks, and so many of them are poorly put together.
I'm not especially nerve-wracked about being a knowledge worker, because my day-to-day doesn't consist of being handed a detailed specification of exactly what is required, and then me 'computing' it. Although this does sound a lot like what a product manager does!
If you have to keep checking the result of an LLM, you do not trust it enough to give you the correct answer.
Thus, having to 'prompt' hundreds of times for the answer you believe is correct over something that claims to be smart - which is why it can confidently convince others that its answer is correct (even when it can be totally erroneous).
I bet if Google DeepMind announced the exact same product, you would equally be as skeptical with its cherry-picked results.
I have spent significant time with GPT-4o, and I disagree. LLMs are as useful as a random forum dweller who recognises your question as something they read somewhere at some point but are too lazy to check so they just say the first thing which comes to mind.
Here’s a recent example I shared before: I asked GPT-4o which Monty Python members have been knighted (not a trick question, I wanted to know). It answered Michael Palin and Terry Gilliam, and that they had been knighted for X, Y, and Z (I don’t recall the exact reasons). Then I verified the answer on the BBC, Wikipedia, and a few others, and determined only Michael Palin has been knighted, and those weren’t even the reasons.
Just for kicks, I then said I didn’t think Michael Palin had been knighted. It promptly apologised, told me I was right, and that only Terry Gilliam had been knighted. Worse than useless.
Coding-wise, it’s been hit or miss with way more misses. It can be half-right if you ask it uninteresting boilerplate crap everyone has done hundreds of times, but for anything even remotely interesting it falls flatter than a pancake under a steam roller.
> Only one Monty Python member, Michael Palin, has been knighted. He was honored in 2019 for his contributions to travel, culture, and geography. His extensive work as a travel documentarian, including notable series on the BBC, earned him recognition beyond his comedic career with Monty Python (NERDBOT) (Wikipedia).
> Other members, such as John Cleese, declined honors, including a CBE (Commander of the British Empire) in 1996 and a peerage later on (8days).
Maybe you just asked the question wrong. My prompt was "which monty python actors have been knighted. look it up and give the reasons why. be brief".
The point is that you never know what you can trust or not. Unless you’re intimately familiar with Monty Python history, you only know you got the correct answer in one shot because I already told you what the right answer is.
Oh, and by the way, I just asked GPT-4o the same question, with your phrasing, copied verbatim and it said two Pythons were knighted: Michael Palin (with the correct reasons this time) and John Cleese.
¹ And I’ve had enough discussions on HN where someone insists on the correct way to prompt, then they do it and get wrong answers. Which they don’t realise until they shared it and disproven their own argument.
Similarly language models are probabilistic and yet they get the easiest questions right 100% of the time with little variability and the hardest prompts will return gibberish. The point of good prompting is to get useful responses to questions at the boundary of what the language model is capable of.
(You can also configure a language model to generate the same output for every prompt without any random noise. Image models for instance generate exactly the same image pixel for pixel when given the same seed.)
In comparison, with text a single word can change the entire meaning of a sentence, paragraph, or idea. The same word in different parts of a text can make all the difference between clarity and ambiguity.
It makes no difference how good your prompting is, some things are simply unknowable by an LLM. I repeatedly asked GPT-4o how many Magic: The Gathering cards based on Monty Python exist. It said there are none (wrong) because they didn’t exist yet at the cut off date of its training. No amount of prompting changes that, unless you steer it by giving it the answer (at which point there would have been no point in asking).
Furthermore, there’s no seed that guarantees truth in all answers or the best images in all cases. Seeds matter for reproducibility, they are unrelated to accuracy.
You think the probabilistic nature of language models is a fundamental problem that puts a ceiling on how smart they can become, but you're wrong.
No. Language can be fuzzy, yes, but not at all in the same way. I have just explained that.
> LLMs can create factually correct responses in dozens of languages using endless variations in phrasing.
So which is it? Is it about good prompting, or can you have endless variations? You can’t have of both ways.
> You fixate on the kind of questions that current language models struggle with
So you’re saying LLMs struggle with simple factual and verifiable questions? Because that’s all the example questions were. If they can’t handle that (and they do it poorly, I agree), what’s the point?
By the way, that’s a single example. I have many more and you can find plenty of others online. Do you also think the Gemini ridiculous answers like putting glue on pizza are about bad promoting?
> You think the probabilistic nature of language models is a fundamental problem that puts a ceiling on how smart they can become, but you're wrong.
One of your mistakes is thinking you know what I think. You’re engaging with a preconceived notion you formed in your head instead of the argument.
And LLMs aren’t smart, because they don’t think. They are an impressive trick for sure, but that does not imply cleverness on their part.
If you pay careful attention to prompt phrasing you will get a lot more mileage out of these models. That's the bottom line. If you believe that you shouldn't have to learn how to use a tool well then you can be satisfied with your righteous attitude but you won't get anywhere.
Wow. So we can expect scaling to continue after all. Hyperscalers feeling pretty good about their big bets right now. Jensen is smiling.
This is the most important thing. Performance today matters less than the scaling laws. I think everyone has been waiting for the next release just trying to figure out what the future will look like. This is good evidence that we are on the path to AGI.
> I really hope people understand that this is a new paradigm: don't expect the same pace, schedule, or dynamics of pre-training era. I believe the rate of improvement on evals with our reasoning models has been the fastest in OpenAI history.
> It's going to be a wild year.
After reading through the examples, I am shocked at how incredibly good the model is (or appears to be) at reasoning: far better than most human beings.
I'm impressed. Congratulations to OpenAI!
There must be some magic sauce here for guiding LLMs which boosts performance. They must think inspecting a reasonable number of chains would allow others to replicate it.
They call GPT 4 a model. But we don't know if it's really a system that builds in a ton of best practices and secret tactics: prompt expansion, guided CoT, etc. Dalle was transparent that it automated re-generating the prompts, adding missing details prior to generation. This and a lot more could all be running under the hood here.
Will the next model be named "1k", so that the subsequent models will be named "4o1k", and we can all go into retirement?
at least NVDA should benefit. i guess.
go play more, your priorities and focus on it being work are making you think this to be harder than it is, and the models can even tell you this.
you don’t have to like the answer, but take it seriously, and you might come back and like it quite a bit.
you have to have patience because you likely wont have scale - but it is not just patience with the response time.
Sadly I think this OpenAI announcement is hot air. I am now (unfortunately) much less enthusiastic about upcoming OpenAI announcements. This is the first one that has been extremely underwhelming (though the big announcement about structured responses (months after it had already been supported nearly identically via JSONSchema) was in hindsight also hot air.
I think OpenAI is making the same mistake Google made with the search interface. Rather than considering it a command line to be mastered, Google optimized to generate better results for someone who had no mastery of how to type a search phrase.
Similarly, OpenAI is optimizing for someone who doesn't know how to interact with a context-limited LLM. Sure it helps the low end, but based on my initial testing this is not going to be helpful to anyone who had already come to understand how to create good prompts.
What is needed is the ability for the LLM to create a useful, ongoing meta-context for the conversation so that it doesn't make stupid mistakes and omissions. I was really hoping OpenAI would have something like this ready for use.
Though o1 did fail at the puzzle in my profile.
Maybe it's just tougher than even, its author, I had assumed...
I am looking at a TypeScript project with quite an amount of type gymnastics and a particular line of code did not validate with tsc no matter what I have tried. I copy pasted the whole context into o1-preview and it told me what is likely the error I am seeing (and it was a spot on correct letter-by-letter error message including my variable names), explained the problem and provided two solutions, both of which immediately worked.
Another test was I have pasted a smart contract in solidity and naively asked to identify vulnerabilities. It thought for more than a minute and then provided a detailed report of what could go wrong. Much, much deeper than any previous model could do. (No vulnerabilities found because my code is perfect, but that's another story).