OpenAI's o1-pro now available via API
platform.openai.com
platform.openai.com
Very expensive, but I've been using it with my ChatGPT Pro subscription and it's remarkably capable. I'll give it 100,000 token codebases and it'll find nuanced bugs I completely overlooked.
(Now I almost feel bad considering the API price vs. the price I pay for the subscription.)
Look carefully through my codebase and identify any bugs/issues, or refactors that could improve it.
<codebase>
…
</codebase>
Doesn't have to be anything overly complicated to get good results. It also does well if you give it a git diff. Do blah blah blah while taking blah and blah into account. Here is my current code:
File `file1.js`:
```javascript
console.log('I am number one!')
```
File `file2.js`:
```javascript
console.log("I am number two :(")
```
Not sure if I'm imagining, but when I tried with/without the markdown code blocks, it seems to do better when I used markdown code blocks, so wrote a quick CLI that takes a directory path + prompt and creates something like that automatically for me. Often times I send identical prompts to ChatGPT+DeepThink+Claude, compare the approaches and continue with the one that works best for that particular problem, so having something reusable really saved time for this.Edit: fuck it, in case people are curious how my little CLI works, I threw it up here: https://github.com/victorb/prompta (beware of bugs and whatnot, I've quite literally hacked this together without much thought)
What I end up with, is one .md file that uses variables like "$SRC", "$TESTS" and "$DOCS" inside of it, that gets replaced when you run `prompta output`, and then there is also a JSON file that defines what those variables actually get replaced with.
Bit off-topic, but curious how your repository ends up having 8023 lines of something for concatenating files, while my own CLI sits on 687 lines (500 of those are Rust) but has a lot more functionality :)
https://marketplace.visualstudio.com/items?itemName=DVYIO.co...
When you say: But is that really a bug?
GPT: That's right. Now that I see it again this is not a bug….and a lot of blah blah.
What difference are you seeing from these models that makes it better?
I say this as someone considering shelling out $200 for ChatGPT pro for this.
I also frequently download all the source code of libraries I am debugging, and when running into issues, pass that code in along with my own broken code. Its very good
It may also have a larger usable context window, not totally sure about that.
Can you provide an example of what you mean by this? I provide very verbose prompts where I know what needs to be done and just let AI “do” the work. I’m curious how this is different?
And partly it can actually execute more at the same time without starting to make mistakes
I routinely put about 100K + of context into Sonnet 3.7 in the form of source code, and in the Extended mode, given the right prompt, it will output perhaps 20 large source files before having to make a "continue" request (for example if it's asked to convert a web app from templates to React).
I'm curious whether O1 Pro actually exceeds Sonnet 3.7 in Extended mode for coding or not. Looking forward to seeing some benchmarks.
> We evaluate 12 popular LLMs that claim to support contexts of at least 128K tokens. While they perform well in short contexts (<1K), performance degrades significantly as context length increases. At 32K, for instance, 10 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almost-perfect baseline of 99.3% to 69.7%.
Have a look at the RULER benchmark for a bit more detail.
The naming would suggest that o1-pro is just o1 with more time to reason. The API pricing makes that less obvious. Are they charging for the thinking tokens? If so, why is it so much more expensive if there are just more thinking tokens anyways?
Shameless plug: One of the reasons I wrote my AI coding assistant is to make it easier to get problems into o1pro. https://github.com/jbellis/brokk
AGI means the value of a human is the same as an LLM, but the energy requirements of a human are higher than those of an LLM, so humans won't be economical any more.
Some quick back of the envelope says that it would take around 35 MWh to get to 40 years old (2000 kcal per day)
(And way more sense than how the power of love was supposed to be a nearly magical power source in #4. Boo. Some of the ideas in that film were interesting, but that bit was exceptionally cliché.)
Two axies: Quality and length.
They're good quality. Not award winning, but significantly better than e.g. even good Reddit fiction.
But they still struggle with length, despite what the specs say about context length. You might manage the script length needed for a kid's cartoon, but not yet a film.
I'll see if I can find another copy of the script; what I saw was long enough ago my computer had a PPC chip in it.
Pizza box? I loved the 6100.
Looks like there's many different old scripts, no idea which, if any, was what I read back in the day: https://old.reddit.com/r/matrix/comments/rb4x93/early_draft_...
I miss those days. Even software development back then was more fun with REALbasic than today with SwiftUI.
“AGI” is also an under-specified term. It will start (maybe is already there) equivalent to, say, a human in an overseas call center, but over time improve to the equivalent of a Fortune 500 CEO or Nobel prize winner.
“ASI”, on the other hand, will just recreate entire businesses from scratch.
Could take me a while to add support for it to my LLM tool: https://github.com/simonw/llm/issues/839
This is not a chat model and thus not supported in the
v1/chat/completions endpoint. Did you mean to use v1/completions?Notes and SVG output here: https://simonwillison.net/2025/Mar/19/o1-pro/
I don't know what I was expecting when I clicked the link but it definitely wasn't this: https://simonwillison.net/tags/pelican-riding-a-bicycle/
OpenAI is now within an order of magnitude of a highly skilled humans with their frontier model pricing. o3 pro may change this but at the same time I don’t think they would have shipped this if o3 was right around the corner.
If you attach a credit card to o3 and give it some onboarding docs, it'll give you a nice summary of your onboarding docs that you didn't need.
We're a long way from a model doing arbitrary roles. Currently at the very minimum, you need a competent office worker to run the model, filter its output through their judgement, and act on it.
Yet somehow the AI knows a treatment?
LLM’s are good at a class of tasks that humans aren’t.
Every time I try to get this thing to read my codebase and onboarding docs (about 40k line angular codebase) it is "pull your hair out" failing leading to frustration.
I guess...if by office worker you mean a manager that does nothing but attend meetings and otherwise talk to people. For other workers you probably want to count the token equivalent of their actual work output and not just the chatting.
My guess is that this is better than a human who would cost $16k/year to hire. But with the logarithmic improvements in quality for linear price increases, I'm not sure it would be good enough to replace a $160k/year worker.
> 180x60x5x48 (working weeks/year) = 20,736,000 tokens/year
Ironic since I was getting ready to cancel my Pro subscription, but 4.5 is too nice for non-coding/math tasks.
God I can't wait for o3 pro.
To me it seems like o1-pro would be to be used as a switch-in tool or to double-check your codebase, than a constant coding assistant? (Even with lower price), as I assume I would need to get done a tremendous amount of work including domain knowledge done to come up for the 10x more speed (estimated) of Sonnet?
I pay for it and will probably keep doing so, but I find that I use it only as a last resort.
...and then ask you for a refund or service credit.
What use case could possibly justify this price?
What you cannot do. You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not: Use Output to develop models that compete with OpenAI.
Presumably you run the risk of getting banned if they realize what you're doing.
OpenAI trained on the world's data. Data they didn't license.
Anyone should be able to "rip them off" and copy their capabilities on the cheap.
This reads as if they consider developing models that compete with OpenAI as illegal, harmful or abusive. Which is crazy. (The other dot points in their list in the linked terms seem better).
1) Why wasn't OpenAI doing it themselves?
2) This means we've reached technological singularity if AI models can improve themselves (as in getting a smarter model, not just compressing existing ones like Deepseek)
As far as I know, OpenAI has been doing this, using both experts and Kenyan workers as well as their own discriminator models. Unfiltered synthetic data is generally used more for distilled models and fine tunes for a specific use case.
Now to have that delivered to you in less than an hour? That’s a huge win.
IR35 rules largely closed down a lot of the high-value contracting business and pushed day rates down heavily.
So why are you saying we'd get a dev for a week for $600? :)
Even we take that at face value, now we're claiming devs are $30K/year. (This was sort of baseline full-time pay when I was in high school, 20 years ago)
Why do I need 1,000,000 tokens for 500 loc?
I don't think further futzing with the numbers makes this work, it's off by multiple OOMs while barely credible
As we all know it's never just 500 lines of code, and 500 lines can take up quite a lot of tokens, especially on a reasoning model like o1-pro which we know can spend quite a hefty chunk of tokens on thinking.
Once you add in iteration and testing you can find yourself racking up quite the bill even on small projects.
It's non-trivial work that requires domain knowledge, i.e. it would have been a 4-6 week project for a noogler, and I would have been impressed if they did it without heavy guidance. (Emergent requirements around purposely using an older API, stuff like that).
- 80 chars per line, 30 occupied (avg'd across 300 KLOC in codebase)
- 500 lines of code
- 15000 characters
- 4 chars / token
- 3750 tokens output
- 10 full iterations, and don't apply cached token pricing that's 90% off
- 37,500 tokens req'd in output
- $600 / 1M tokens
- $0.60 / 1K tokens
- $18
Further we also aren't counting input here which can get long since it includes the previous output, which for the last request will be 33,750 + reasoning + any prompts, which will increase your cost quite a bit there.
But yes that is more reasonable than I'd expect I must admit, but I still think it needs to be at least an order of magnitude cheaper to compete against the other models out there.
I'm not sure I know a lot of employees who would allocate that sort of constant funding to a consumables tool for an employee, given that's your usual monthly cost for a typical saas product.
You generally can't see this due to all the middle men and bloat, and that no corp really wants to measure dev productivity against revenue like this as it'd raise a lot of uncomfortable questions by both devs and shareholders alike.
YMMV
Go check out flutter_pcm_sound_fork, find me even one package with the same streaming PCM => speakers functionality, and I'll give you $500. All I ask is, as a personal favor to me, you read the part in the Hacker News FAQ about "coming with curiosity"
I think you can probably get similar results for a much lower price using llm-consortium. This lets you prompt as many models as you can afford and then chooses or synthesises the best response from all of them. And it can loop until a confidence threshold is reached.