HNHacker News
TopNewBestAskShowJobs

zzleeper

3,122 karma · joined July 31, 2009

submissionscomments
zzleeper··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
A bit tired of spending $200 out-of-pocket for openai. What do you use as harness? (for me the harness if half of the benefit... controlling my PC, working from phone, etc.)
zzleeper··on South Africa is at risk of becoming a mafia state
Far left? Most economists consider it center-right, and it has been like that for more than 100 years.
zzleeper··on How to Write with an LLM
I thought that was a joke.. Like how I used "delve into" back in 2024 or so lol
zzleeper··on The Microeconomics of Artificial Intelligence (2025)
Honestly, it's a bit of a disappointment

- Many more mediocre papers written (mediocre ideas, implementation, claude-isms everywhere)

- Much easier to try every possible combination of a regression in order to show the result you want (same for theorists).

The one thing I'm happy about is it's now much easier to extract historical data from old documents from Google Books. Still not perfect, but takes you 95% there. And creating plots and datavis just for quick exploration is super fast.

zzleeper··on OpenAI begins rolling out GPT-6 Astra
Also very puzzling to me. And the jargon-speak, albeit is more of an issue for Claude, is still puzzling. Wonder what part of RL led to this.
zzleeper··on OpenAI begins rolling out GPT-6 Astra
I just went to bed and left it running; was expecting maybe 20 minutes :)

And I did gave the program a bunch of code guides -- this [1] for instance -- which included quotes like "Prefer straightforward code over clever code." but somehow that didn't matter.

[1] https://github.com/sergiocorreia/overengineered-rand-mcnally...

zzleeper··on OpenAI begins rolling out GPT-6 Astra
Sure, why not: https://github.com/sergiocorreia/overengineered-rand-mcnally

The original script was mostly very simple python:

1. Download some public PDFs. 2. Have a double for-loop (over PDFs and pages within PDF), 3. Use a library to call gemini-3.7-flash and ask it to run some OCR 4. Save JSON outputs, save a csv with results, validate with some Stata code

New code folder was 189 files. Just the PDF download folder is now 7 files involving an adapter, a source manager, an acquisition manager, etc.

Every instance of saving a file involves saving a temporary copy and then moving it, so e.g. I lose power, we minimize the risk of corrupted files.

And so on!

zzleeper··on OpenAI begins rolling out GPT-6 Astra
I wonder if I would need a non-openai agent to enforce it.. I have tried so far with skills and agents.md and code stills end up over engineered to the moon.

Will ask OpenAI to write me that agent! Hope the agent is not over engineered or else unsure how to solve the bootstrap puzzle :D

zzleeper··on "Welcome to the AGI era," OpenAI says as GPT-6 Astra debuts
noted, thanks dang!
zzleeper··on "Welcome to the AGI era," OpenAI says as GPT-6 Astra debuts
(cross posted from the other announcment thread)

A big problem I have with OpenAI's models (and of course Claude) is that they tend to write the most over-engineered pieces of code, beyond the imagination of any architecture's astronaut.

Just this week I asked 5.6-sol-ultra to update a 1000 LOC python script I had, to "incorporate the key lessons learned when using it for another project".

I left it overnight and went to sleep. In the morning I realized it had created a monstruosity of 180 PYTHON SCRIPTS, with maybe 100,000 lines of code, each more crazy than the other. It took me minutes even to track where a single action took place, due to all the crazy imports, defensive coding, and premature optimization.

Similarly, anything they write is riddled with jargon that almost feel like they want me to give up trying to understand. Made up phrases that ended up with me having no idea of what was going on.

So now to my assessment: The reason why " Nobody Has Actually Built a Software Factory" [1], and why even SOTA LLMs struggle so much with open-ended unsupervised tasks is precisely this. They somehow let complexity explode, and unless it's also accompanied with an explosion in e.g. the number of agents, the amount of processing time, etc. then projects become broken/unmanageable.

Sure, LLMs are great at producing code that can be thrown out, so they are amazing when searching for exploits, for instance. But as of 5.6 they still lack either a better harness that encourages KISS principles, or a better RL step.

(And not sure why, but doubt Astra will fix this.. they seem to be aiming for AGI and for beating crazy benchmarks, which is not very aligned with KISS)

[1] https://news.ycombinator.com/item?id=49510843

zzleeper··on OpenAI begins rolling out GPT-6 Astra
(Posting partly so I can revisit my predictions when they open access more widely)

A big problem I have with OpenAI's models (and of course Claude) is that they tend to write the most over-engineered pieces of code, beyond the imagination of any architecture's astronaut.

Just this week I asked 5.6-sol-ultra to update a 1000 LOC python script I had, to "incorporate the key lessons learned when using it for another project".

I left it overnight and went to sleep. In the morning I realized it had created a monstruosity of 180 PYTHON SCRIPTS, with maybe 100,000 lines of code, each more crazy than the other. It took me minutes even to track where a single action took place, due to all the crazy imports, defensive coding, and premature optimization.

Similarly, anything they write is riddled with jargon that almost feel like they want me to give up trying to understand. Made up phrases that ended up with me having no idea of what was going on.

So now to my assessment: The reason why " Nobody Has Actually Built a Software Factory" [1], and why even SOTA LLMs struggle so much with open-ended unsupervised tasks is precisely this. They somehow let complexity explode, and unless it's also accompanied with an explosion in e.g. the number of agents, the amount of processing time, etc. then projects become broken/unmanageable.

Sure, LLMs are great at producing code that can be thrown out, so they are amazing when searching for exploits, for instance. But as of 5.6 they still lack either a better harness that encourages KISS principles, or a better RL step.

(And not sure why, but doubt Astra will fix this.. they seem to be aiming for AGI and for beating crazy benchmarks, which is not very aligned with KISS)

[1] https://news.ycombinator.com/item?id=49510843

zzleeper··on OCR It – pull text out of un-copyable documents for your LLM
Definitely not. Even Chrome has a built in OCR that performs amazingly. I got an LLM to write a quick python wrapper to it [1], so I'm sure you should be able to access it from an extension

[1] https://github.com/sergiocorreia/clv-locro

zzleeper··on Mistral OCR 4.1
Same here. Maybe Fable is better but in terms of cost effectiveness it wouldn't even make sense to test it
zzleeper··on Advancing the price-performance frontier with GPT‑5.6
Had to ctrl+f for someone saying this.

I typically do lots of mini calls for research (100s of millions or something in that ball park). Newer models made that absolutely impossible, and the fact that the older ones are starting to get deprecated made me switch to e.g. deepseek for some of my runs. We'll see if I move back after this.

zzleeper··on Qwen 3.8
yes you are!
zzleeper··on Pacing the frontier
It means only those vetted will be allowed to use frontier models (i.e. let's pace ourselves and not share the frontier broadly)
zzleeper··on Private Claude Chats Exposed in Google and Bing Search Results
Does Chrome use URLs you click on to help its indexer? EG if someone sends you a link to www.example.com/mysecretpage and somehow it appears in Google later.

That might be a case where what you expect is private is leaked by the browser

zzleeper··on ARC-AGI Leaderboard
How believable is this benchmark? EG maybe opus was training on this? (You can try to identify the IP of wherever previous ARC questions came from)
zzleeper··on Country went 100% electric vehicles overnight with a drastic approach
> But there’s a catch. Actually, several.

> Laos isn’t doing this for climate headlines. The logic is economic.

> The headline writes itself. A country went nearly 100% EV overnight. But the mechanism matters..

I mean, come on... if it quacks like a duck and it em-dashes like a duck...

zzleeper··on Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge
Random question, but has there been any improvement in OCR/document understanding in these newer models? Last time I checked (1mo ago) SOTA was still sadly Gemini, unless you wanted to pay $$$ for e.g. Sol
zzleeper··on Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?
> You don’t have access to this conversation. Make sure you’re logged in to the right account, or ask the conversation owner to send you a share link.
zzleeper··on Please don't discontinue Gemini 2.5 Flash
gemini 2 and 2.5 were great models for quick-and-dirty OCR

It was fine to lose 2, but 2.5 will be dearly missed as it hit the sweet spot in terms of cost-performance :/

zzleeper··on ArXiv's Next Chapter
At least in economics it can easily be 1-5 years until you go from draft to journal. In the meantime, you want a way for others to easily cite your paper, to make different revisions available, for you to post it in a way that's stable (people's websites change all the time, etc.)

Also, because most folks don't want to deal with paywalls, it's standard practice to put the last version of your draft before conditional acceptance on an online repository. It used to be SSRN for econ/finance, but they sold out to Elsevier, so now arxiv is increasingly being used.

zzleeper··on Asian AI startups launch Mythos-like models
I tested Fable through Cursor; asked for ideas on how to make a data website I have less "Claude-like" (IYKYK what are the usual tells), and it spun out the most useless, Claude-like CSS styling ever, wasting $40 in 10 minutes.

The website was created through Opus, so you could also say the results were worse than Opus. (This is just to say that I had the same experience using the US models, so perhaps those Asian models are Mythos-like lol)

zzleeper··on Haystack: Open-Source AI Framework for Production Ready Agents, RAG
Would you have used something else without that constraint?
zzleeper··on FDA advisors unanimously vote to approve Moderna's mRNA after agency drama
I really wonder what's up with that. Also remember the crazy Stanford guys.. did something flip in their brain or were they just always like that?
zzleeper··on SpaceX to buy Cursor for $60B
Same path as you. Went from $60 cursor plan (often exceeding it which costed more in API) to a limitless $100 codex plan where I basically say "read the markdown and implement the instructions". Deepseek also works quite well, surprisingly!

(FWIW Im mostly using python for OCR, LLM calls, data analysis..)

zzleeper··on New pancreatic cancer drug might open the door to much longer survival times
I'm pretty sure only a small fraction of grants gave this issue, and the cuts have meanwhile being very wide, without any sort of intelligent approach (I know ppl doing stuff like material science at nasa that now have nothing to do because they cut costs of various inputs, while the very expensive lab equipment is sitting there now unused)
zzleeper··on Statement on US government directive to suspend access to Fable 5 and Mythos 5
I asked it to tweak the fonts/colors of a very very simple static page and it blew through $35 (which is a lot for me lol; it's 10 days of my monthly codex plan).
zzleeper··on Show HN: FablePool – pool money behind a prompt, and Fable builds it in public
I managed to write one that at least didnt had the font and colors (using 4.5)

Yesterday, I prompted Fable to improve the frontend to make it look different from Claude style, gave detailed examples etc. 15 minutes and $32 dollars (!) later (used cursor lol) it gave me the shittiest more claudiest website ever, basically ignoring everything I asked

Page 1 of 33Next →