That said, I'd like this quality with a relatively quick tool using model; I'm not sure what else I'd want to call it "AGI" at that point.
That said, I'd like this quality with a relatively quick tool using model; I'm not sure what else I'd want to call it "AGI" at that point.
With coding using anything is always a hit and miss, so I prefer to have faster models where I can throw away the chat if it turns into an idiot.
Would I wait 15 minutes for a transcription from Python to Rust if I don't know what the result will be? No.
Would I wait 15 minutes if I'd be a mathematician working on some kind of proof? Probably yes.
It’s the progressive jpg download of 2025. You can short circuit after the first model which gives a good enough response.
I feel like we are in awkward phase of: "We know this has severe environmental impact - but we need to know if these tools are actually going to be useful and worth adopting..." - so it seems like just keeping the environmental question at the forefront will be important as things progress.
This is a rhetorical question.
Sure we aren’t capturing every last externality, but optimization of large systems should be pushed toward the creators and operators of those systems. Customers shouldn’t have to validate environmental impact every time they spend 0.05 dollars to use a machine.
I don't really understand the critique of GPT-4 in particular. GPT-4 cost >$100 Million to train. But likely less than 1 billion. Even if they pissed out $100 million in pure greenhouse gases, that'd be a drop in the bucket compared to, say 1/1000 of the US military's contributions
Does that "hundreds" include the cost of training one human to do the work, or enough humans to do the full range of tasks that an LLM can do? It's not like-for-like unless it's the full range of capabilities.
Given the training gets amortised over all uses until the model becomes obsolete (IDK, let's say 9 months?), I'd say details like this do matter — while I want the creation to be climate friendly just in its own right anyway, once it's made, greater or lesser use does very little:
As a rough guess, let's say that any given extra use of a model is roughly equivalent to turning API costs into kWh of electricity. So, at energy cost of $0.1/kWh, GPT-4.1-mini is currently about 62,500 tokens per kWh.
IDK the typical speed of human thought (and it probably doesn't map well to tokens), but for the sake of a rough guide, I think most people reading a book of that length would take something around 3 hours? Which means if the models burn electricity at about 333 W, they equal the performance (speed) of a human, whose biological requirements are on average 100 W… except 100 W is what you get from dividing 2065 kcal by 24h, and humans not only sleep, but object to working all waking hours 7 days a week, so those 3 hours of wall-clock time come with about 9 hours of down-time (40 hour work week/(7 days times 24 hours/day) ~= 1/4), making the requirements for 3 hours work into 12 hours of calories, or the equivalent of 400 W.
But that's for reading a book. Humans could easily spend months writing a book that size, so an AI model good enough to write 62,500 useful tokens could easily be (2 months * 2065 kcal/day = 144 kWh), at $0.1/kWh around $14.4, or $230/megatoken price range, and still more energy efficient than a human doing the same task.
I've not tried o3*, but I have tried o1, and I don't think o1 can write a book-sized artefact that's worth reading. But well architected code isn't a single monolith function with global state like a book can be, you can break everything down usefully and if one piece doesn't fit the style of the rest it isn't the end of the world, so it may be fine for code.
* I need to "verify my organisation", but also I'm a solo nerd right now, not an organisation… if they'd say I'm good, then that verification seems not very important?
Transporting something with a car using fossil fuel usually uses less energy than if a human did the same thing by hand, that doesnt mean fossil fuel is environmentally friendly. LLM:s does not decrease the population even if it can do human tasks. If the LLM is used for the good of humanity it is probably a win, but I mean obviously a lot of the use of AI is not.
I use LLM:s as well, I'm just saying, I dont think it is a totally strange question to ponder over the energy use of different use cases with LLM:s.
Until then, the choice is being made by the entities funding all of this.
It's just the number of tokens it's willing to expend in the little internal dialogue before it has to spit out an answer.
Am I the only one who is looking around at the AI industry and seeing only developers and artists being replaced with AI?
It's hardly AGI when it can't replace a salesperson, or an accountant, or a lawyer, or a teacher, or ...
All the headlines I am seeing are software development related. In this post, you yourself are using s/ware development as a measure for how good/bad AI is.
I use AI a lot to double check my code via a code review what I've found is
Gemini - really good at contextual reasoning. Doesn't confabulate bugs that don't exist. Is really good at finding issues related to large context. (this method calls this method, and it does it with a value that could be this)
Sonnet/Opus - Seems to be the more creative. More likely to confabulate bugs that don't exist, but also most likely to catch a bug o3 and gemini missed.
o3 - Somewhere in the middle
the only example uses I see written about on HN appear to basically be Substack users asking o3 marketing questions and then writing substack posts about it, and a smattering of vague posts about debugging.
Example: Pull together a list of the top 20 startups funded in Germany this year, valuation, founder and business model. Estimate which is most likely to want to take on private equity investment from a lower mid market US PE fund, as well as which would be most suitable taking into consideration their business model, founders and market; write an approach letter in english and in german aimed at getting a meeting. make sure that it's culturally appropriate for german startup founders.
I have no idea what the output of this query would be by the way, but it's one I would trust to get right on
* the list of startups
* the letter and its cultural sensitivity
* broad strokes of what the startup is doing
Stuff I'd "trust but verify" would be
* Names of the founders
* Size of company and target market
Stuff I'd double check / keep my own counsel on
* Suitability and why (note that o3 pro is def. better at this than o3 which is already not bad; it has some genuinely novel and good ideas, but often misses things.)
Or you could just cut out the middleman(bot) and just do the search yourself, since you're going to have to anyway to verify what the "AI" wrote. It's just all so stupid that society is rushing towards this iffy-at-best technology when we still need to do the same work anyway to verify it isn't bullshitting us. Ugh, I hate this timeline.
With no deep research - agreed; too recent to believe info is accurately stored in the model weights.
how do you validate all of that is actually correct?
Like how there's a ton of psychics, tarot and palm readers around Wall St.
If OP had suggested that they were just medium-quality nonsense generators I would have just agreed and not replied.
Then I have it take those matches and try and chase down the hiring manager based on public info.
I did it at first just to see if it was possible, but I am getting direct emails that have been accurate a handful of times and I never would have gotten that on my own
Thank you!
I had an example where o1 really wowed me - something I don't want to post on the internet because I want to use it to test models. In that case I was thinking through a problem where I had made an incorrect mathematical assumption. I explained my reasoning to o1 and it was able to point out the flaw in my reasoning, along with some examples mathematical expressions that disproved my thinking.
The funny thing in this case it basically functioned as a rubber duck. When it started producing a response I had deduced essentially what it told me - but it was pretty nice to see the detailed reasoning with examples that might've taken me a few more minutes to work out. And I never would've produced a little report explaining in detail why I was wrong, I would've just adjusted my thinking. Having the report was helpful.