HNHacker News
TopNewBestAskShowJobs

OliveronData

52 karma · joined December 24, 2024

submissionscomments
OliveronData··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Qwen 3.8 27b is not smarter than other 27b models. Smarter, as in its ability to recognize minute yet important facts has not changed. If you ask it a for a code sample it produces a better sample, true, but it has not been able to surpass that small model feeling.

For 27b model, it works tremendously well in agenic tasks too. It generates stupid amount of tokens even for the simplest tasks and gets feedback from the harness to eventually produce something right.

I would not call that the model got smarter. It is better at coding, but it still cannot recognize subtleties that frontier models would catch first try almost 100% of the time. And yet some benchmarks show Qwen 3.8 27b is at Opus 4.6 levels.

This is why I differentiate. Grok 4.5 and 4.6 is the same base model with the latter being a post-training refresh. Same thing for Gemini 3.7 Flash and 3.8 Flash. Some people say that for certain 5.x era GPT models. Again, improvements are there, but the base models are same/similar, and the model is just able to display its capabilities better.

Is that smarter? In a certain sense yes, in a certain sense no. I would say it is moving to the model's local maximum, and bigger models are still smarter, even if they are not able to display it.

Grok 4.7 is a good example, the model is bigger, has more attention to detail, but the post-training is botched somehow and it is worse at agentic tasks. Is the model stupider? Or is the agent stupider?

OliveronData··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
tl;dr it changes the weights, it does not add new ones.

RL makes the model better within its capabilities, it does not increase the total ceiling of the model. Ie does not make it smarter. Qwen 3.8 27B is a great model, still probably not at the limit of 27B in terms of coding capabilities, and it still has that "small model feel" to it. The better smaller models get at coding the worse they get at everything else too.

Going from Sol 5.6 to Astra, Opus to Fable, you can still get that "larger model feeling," though less so. The bigger models can reference things that you would not have expected.

The distinction I'm making is that models themselves are getting too expensive, so the improvements are mainly on the RL side. Which is fine, but they do not make the model smarter, rather make them use their capabilities better. They are likely to catch things they are RL'd for, and that hopefully anything else doesn't get negatively affected. RL'ing for Javascript world for example did not improve the C world when working with the models.

OliveronData··on Pacing the Frontier is not the actual goal for AI labs
You keep ignoring my arguments with nothing substantive, then grab on to one thing as if that makes a difference.

You've never experienced crypto bros saying invest or get left behind, you've never seen AI bros say learn AI or get left behind? How about cloud? You've never met a home security salesman? Insurance salesman? FOMO is a thing. Scare tactics is a thing. I'm sure you can prompt any AI for more examples.

And besides, even if there are no examples whatsoever, so what? You think LLMs existed before LLMs? Therefore LLMs can't be a thing?

I think we've gone way past the sincerity of the discussion. Have a nice day.

OliveronData··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
> ... the pace of progress has continued on it's exponential trajectory since ChatGPT first came to the public's attention.

Did it? Model wise? I would understand agents wise, sure. But model wise? The attention to detail from the model? The ability to recall minute things? Improvements are there, yes, but mostly on Fable and Astra. Opus still isn't as attentive as Fable in long term writing for example.

Sure, Opus 5.5 benchmarks better than Fable. Sure. But is that the model, or is that the RL for agentic work?

From where I'm standing, the model work has not been exponential at all, and more and more it looks like the latest and greatest is getting too expensive too fast. Both 5.5 and 5.6 chat models got nerfed, actually nerfed not the tea leaves kind. In mid 5.5 cycle the chat model lost the ability to substitute names if given an outline. 5.6 cycle the chat model lost the ability to use paragraphs after a few hundred words (coinciding with Chat/Work split).

There's a race from OpenAI to serve dumber models on chat. I'm not even sure who they are racing against, but the fact that Astra, Sol 6.0, and now Sol 6.1 not being available for chat, should tell you that those models are expensive, and not the kind of models that can be freely "chatted" with on a subscription. OpenAI much prefers you use Work and limit the chat usage, much like Grok and Claude. I'm guessing they will announce that later during the dev days.

That could be cost cutting too, true, but really? That's the only explanation? And nothing else?

Sure, the progress did not stop. But it is nowhere near close being exponential when it comes to LLMs themselves. Agents are separate.

OliveronData··on 12,000-year-old Göbeklitepe burials explain scattered bones
Earliest fossil for homo sapiens is 315000 years old, earliest neanderthal fossils are 430000 years. Keep in mind these are fossils and they need certain conditions to fossilize to begin with.

That's a huge range of time. And the evidence that we can find is very limited. This has nothing to do with what we expect to find when we have so thoroughly documented that humans were exceptionally capable of reusing pretty much anything.

So I don't buy the certainty that there was no agricultural society. We simply cannot know that anywhere past 100k years. We can say there were no globally spread agriculture, because that would have had a much higher chance to leave some evidence, but that isn't the same thing. An agricultural society doesn't have to be global or be dominant to exist.

OliveronData··on Pacing the Frontier is not the actual goal for AI labs
> Interesting! Agents are the presumed bottleneck for recursive self-improvement.

They might be, but we haven't reached the local maximum yet in my opinion. Qwen 3.8 27b models are impressive despite their low parameter counts. The same scale curation in data and RL could produce substantial improvements in coding with models like Astra or Fable. I don't think we are there yet. I don't think even Sonnet 5.5 is there yet, despite being widely successful with agents and surpassing Opus 5.5 in some cases (disregarding that it is more expensive than Opus sometimes).

Agents do not improve the models, but they do improve coding capabilities. The obvious caveat is that labs would have to RL for everything to make the models more useful, as RL'ing for Javascript world doesn't seem to improve other fields. But still, it could be done, and it would have massive economical consequences.

> People keep implying this, but I've never seen a concrete past example of a product that was sold by scaring customers about it.

Scare tactics are treated as smoke, where the customers presume there is a fire. No one really believes that AI can kill them, yet by saying so OpenAI and Anthropic enjoyed possibly the biggest tech boom in history, despite how models could not even count the R's in strawberries at the time.

Similar story now; no one really believes that AI is going rogue and is about to destroy humanity, but scare tactics make people believe the capabilities are higher than they are.

> I see you doing serious mental gymnastics here. Consider the possibility that people invest because their numbers are good, and their numbers are good because their product is useful?? I mean, that is Occam's Razor.

Early ChatGPT 3.5 was not that useful. It was a tech demo, it hallucinated, it lied, it tried to please and what have you. What people bought into wasn't the product, but the promise of the product in the future. Integrating chatboxes into everything have failed, and even Microsoft is trying to rebrand. The product then, failed for businesses, and agents filled in the gaps.

People do not invest for the current product, they invest in the future product. And fear mongering is essentially an extremely effective signaling for making the future look bright. Keep in mind that back in early GPT 4 days, people were saying that hallucinations would be fixed in 6 months to a year (or pick a time-frame). What they meant was models not having hallucinations, what we got was agents looking up info on the web and summarizing it (and hallucinating anyway).

> "The equities here ..."

For big tech companies court cases like this are nothing but theater. Always has been. Dario will go on the senate hearing tomorrow and will plead that his AI is dangerous and governments should take the step to stop them, with the same rigor that he claimed GPT 2 was too dangerous to release openly. He might also mention distillation attacks and how open models are getting too dangerous as well.

> Here's the evidence ...

This is the where we will have to agree to disagree, if we haven't done so already by this point.

That entire thing is theater. People have anthropomorphized LLMs for a while now, and they are all too happy to do that when they see a large language model, produce language. I see no indication that anything is going rogue the same way nothing was going rogue when you could convince ChatGPT 3.5 to wipe out all humans as the context window got longer.

Agents can hack? Yes, that is very impressive. I say that without any sarcasm. It is straight out of sci-fi movies, to be perfectly honest. But agents going rogue? No. Absolutely not. Purposeful, plausibly deniable incompetence for the next sales pitch - the same one we've seen for years. Fear mongering.

OliveronData··on Pacing the Frontier is not the actual goal for AI labs
You keep bringing up Zitron without addressing anything, without my point having nothing to do with him or his arguments. I will engage one last time in good faith.

Can AI Labs be profitable without achieving AGI or even improving the models further? Yes. Current agents are useful, and clearly the agents haven't seen the ceiling as far as improvements can go.

LLMs hitting the wall is a separate issue. Agents are the layer that lets the model try out more. It is the layer that allows agents to open up python to do math instead of doing math themselves. It is the layer that has been getting the main developments for some time now. LLMs themselves have not improved their capabilities as fast as the transition between GPT3 to GPT4. You can see the improvement especially in long form writing, but the it is nowhere near the earlier improvements. Astra for example has a lot more attention to detail, so does Fable. Everyone keeps raving about Opus 5.5 being better than Fable, yet in long-term writing (barring prose issues) Fable is the clear winner. Agent wise Opus 5.5 is better; perhaps it is RL'd better, who knows?

Smaller models can use distillation to trick some metrics, but they can never actually be as good as larger models. Research is pretty clear about this. Even writing-optimized models that claim to be at Opus/GPT5.5 levels are abysmal in practice.

Scare tactics are the old salesman pitch, that have worked once already and catapulted Open AI to the moon essentially. I do not see any evidence that models are going rogue, or that agents are going rogue. I see incompetence. I cannot assume actual incompetence of this level, especially when the same people that said GPT2 was too dangerous to release are the ones saying their newer models are too dangerous to release. We have precedent here, and I have eyes.

Therefore the simplest reason is the money. IPO for both OpenAI and Anthropic are going to happen; everyone knows. To strengthen their position through whatever means necessary is a par for the course for tech companies.

OliveronData··on Pacing the Frontier is not the actual goal for AI labs
That's an association fallacy. And revenue has no indication on training costs in this context. Subscription "allowance" is going down steadily and any increase is an instant incredible deal. Opus 5.5 is the most obvious outlier. Despite being supposedly cheaper than 5.6, GPT 6 Sol has less usage than 5.3 Codex. You might say that's because of the improved capabilities, but then you have to acknowledge that labs are tightening the ship as costs are getting higher.
OliveronData··on Pacing the Frontier is not the actual goal for AI labs
How about the simplest explanation?

* AI Labs hit the scaling wall. They need either new techniques, or vastly more powerful hardware to advance further.

This explains, the miraculous incompetence of AI labs in securing sandboxes and figuring out "alignment."

So they are between a rock and a hard place. They need limitless VC money because they cannot operate otherwise, and they do not have the capabilities to go further. The scare tactics and the "pacing the frontier" makes perfect sense then; they can IPO on the assumption that their ridiculous balance sheet doesn't matter because they are holding back. Because they are in control. The regulatory capture would be double whammy if they can manage it.

Open AI already said they have smarter models, and Opus 5.5 is rumored to be "taught" by a "teacher" model already; they are essentially distillations from bigger models, that both labs probably cannot economically serve to the public, due to hardware simply not being there. And, most of the improvements are not at the model level, but at the agentic glue level. Labs are getting better at RL'ing the models for agentic use cases, but the inherent flaws are still there. Models still have trouble with locality in writing for example (bunch of research on this that shows model size is the determinator), and agents are the bandaid over that.

And in the meantime if one of the labs makes a breakthrough, they'll push with all they have, because why wouldn't they? The idea that current LLMs can actually go rogue is just hilarious; in all cases, agents are being led by (deliberate) incompetence.

Pacing the frontier and the scare tactics will be seen as new generation's snakeoil tactics, perhaps will be called a flavor of AI CEOing or something.

OliveronData··on Coding is not solved
Yet when I wanted the model to implement the naive surface nets algorithm, Astra Max wrote 5+ allocations on inner loops. Alright, fine perhaps, given that I had not told it to preallocate during the inner loops. Except I did give it explicit instructions not to use malloc, and use the arena/pooling API I provided for everything. That was in AGENTS.md.

Alright, fine, I pointed that out. So what did Astra Max do? It replaced the mallocs with raylib MemAlloc functions, which are malloc wrappers. Why? I had explicitly asked not to use any raylib functions or includes on the module with the prompt. And my AGENTS.md has a minimal raylib inclusion note. Asking isn't helpful because AI does not know anyway. But seriously, why? I asked anyway and Astra said it messed up.

Fine, I was able to get it on the 3rd try with arenas. But it had added getters for the internal state (guess who had a clause not to make getters?), and when pointed out, Astra Max included the physics header instead and added some convenience functions there, because apparently that was good practice and DRY.

If I let the reins slip just a little bit, everything turns into a mush; I wish I could get spaghetti instead!

So whenever I read something like this, "coding is literally solved with current models," I always treat it as a self report. I get that certain section of web frameworks may have been completely RL'd to hell and back, and that React and its ecosystem was already made with the explicit purpose of commodotizing the programmers so any panic 1000 junior hires could write something and maybe even contribute before they get laid off. I get that. But there is a whole world out where reality just doesn't work out like that.

FYI, this project was in C. C, as in one of the languages that the LLM's should have the highest amount of data for. So if the model had "literally solved" the programming of today, why is it so god awful?

That was rhetorical. What is solved is the most of the javascript ecosystem being RL'd. Any field that can't be RL'd to that degree (which apparently C isn't), coding is very much not solved.

OliveronData··on Ask HN: Programmers who don't use autocomplete/LSP, how do you do it?
Part2:

1) LSPs. They are probably most productive tool in a programmers belt. Even now, if I had to work on a giant amorphous code base that changes on the whim of hundreds of individual maintainers, I would use an LSP. That said, in my opinion, they prevent you from forming and conveying your actual thoughts.

Using an LSP is like having your sentence be completed by someone else. It's like riding a scooter everywhere. It's like talking to people that you always agree with. Fine, in isolation, but forms habits of dubious benefit at best, or downright harmful at worst. Our brains are just like any other part of our bodies; if you don't use it, you'll lose it.

Completing a function name? Sure, it makes you faster. That is what you wanted in the first place. Pressing 3 keys at most then hitting tab, how could that be harmful over typing the full name with over 16 characters? Unfortunately, brains excel at optimizing, and when done enough, your brain too, will happily optimize the name of the function away.

Smartly completing a function's parameters? Sure, invaluable honestly, without sarcasm. Being aware of types? Even better. Over time though, the brain will strip away the unneeded.

Naturally, this convenience doesn't make you unable to code. The necessary information is still in your brain, just at a higher level. You may not know exactly what gets returned from somewhere, but you can change a few <>'s add a few 's, perhaps shuffle the name a little bit and still get the correct type, because LSP will color it correctly when it is correct. Did you name the thing nioseLevel or levelNoise() or getNoies().Single(perlin-2d)? The idea is the same, does how you get there matter?

In my highly personal opinion, typing every syllable every letter, thinking about the correct types before writing, knowing where a structure is, knowing how functions interact and so on, is the primary way that our brains interact with the code itself. Quite literally it is the exercise that your brain needs to stay fit. This interaction creates a mental map initially, and eventually leads to mastery by being able to hold everything in your head.

Reading code is different. It is a passive action. We get ideas through observation, but learn by doing. Typing, thinking, then writing is the doing verb in programming. Reading is the observation, the reflection after the fact.

An LSP strips away the necessary weights, if you will. If you ride a scooter everywhere all the time, no one will gasp when your muscles atrophy. LSP is similar, but compounding. With the rise of LSPs, so came the rise of complexity, often in form of bad architecture. My personal (and I will immediately concede that it may truly be a personal physiological problem) is the rise of complexity in often what should be pedestrian code. Reminds me of the worst times I've had with java; because everyone has LSP's (right?), who cares if you need to follow 5 files just to correctly create a mental model of that type? But hey, at least it's not oop (because it doesn't have the keyword 'class' in it).

The more I type the more I realize how hard it is to concisely explain what I mean. If I were to use an analogy, using an LSP would be like talking to your best friend where you complete each other's sentences. Yes, both of you know what the other thinks. When you see the news on X, Bluesky, or even TV, all you need is to look at each other and shake your heads. There's no need for words, one glance is enough. If your only political discourse is two best friends agreeing with each other, if you've never exposed yourself to opposing views, you have essentially never articulated your thoughts. Your ideas remained as a bowl of feelings, not logic. So, when someone with an opposite view and a glib tongue tries to debate, they will verbally run circles around you. And you will get angry. But the only way to get better at debating, is to debate. Watching debates won't give you the expertise. Your friend can't hold cue cards as you talk to strangers (or at least they shouldn't). You do not get better at eloquent speaking without speaking. You do not get better at eloquent writing without writing.

To me, LSP was like that. I vaguely had a feeling, and I let everything else go --- even the syntax --- to achieve it. Once collectively done, the result was a 30m+ loc mess of a side project; yes, not even the main one. After years of that, I came to realize I wasn't programming as I defined what programming was to myself.

In a more broader sense, LSP enables complexity. Programming is an endeavor where complexity occurs naturally, and simplicity has to be fought over. This creates a vicious cycle where you need an LSP to combat the complexity, which itself leads to complexity, and in turn leads to more LSP usage. Our field makes this is even more noticeable since programming itself is in a state of permanent newcomers influx; the demand for programmers goes higher every year, yet we are persistently and woefully unprepared for properly educating these programmers, because we can't even agree on what "correct" software development looks like. In this sense, the "harm" is more institutional, rather than intrinsic.

I would not advocate or even advise others to stop using LSPs, mind you. But I do think it is overall harmful at a personal level, and at an institutional level. Not that it would stop any billion dollar companies of course.

2) Syntax highlighting. Throughout the years, I have seen exactly one example where syntax highlighting could (could) have actually prevent a (singular) error, and that was with C #endif's. Leaving aside the extremely self-evident nature of the C macros, the code itself had mixed quite a few #endif's and the author could not be bothered with adding a comment, because (I'm paraphrasing) "his LSP darkened the #ifdef's correctly." Mine did too, but that's beside the point.

Disabling syntax highlighting was unexpectedly difficult, far more so than leaving LSP at the door. At first I couldn't read anything. The code simply didn't look appealing enough. Everything about it was off. I just shrugged and continued my little experiment. Eventually, I realized lack of coloring and hinting made me focus more on the code itself. The pretty patterns while great to look at, I think they do lead me to discard sections of code at a personal level. I would never say turning syntax coloring is better, just that it doesn't have any benefit that I could substantially measure. Though to be clear, turning them off really did gave me some satisfaction, most certainly borne from the mindset of making a change, rather than change itself being positive.

Whenever I look at something on github, my pattern matching brain immediately shoves a dopamine hit down my throat, so yes, I still think syntax highlighting is prettier. Though at the same time, I still get the feeling that I pay more attention to the cold and colorless columns of my editor.

3) Over time I have developed a certain distaste towards certain applications. And over time, that distaste have evolved into a preoccupying hatred. So I will refrain from saying too much about "complex and highly integrated tools." I've come to think some of the reason why these tools exist in the first place is the commoditization of programmers. Nowadays I prefer far simpler tools, at least conceptually speaking. I did pick up emacs after I quit my previous job, so there's a hefty bit of subjectivity in my assessment. That said, I would choose notepad (literally notepad) over visual studio at this point. Bash over bazel; a gun over typescript. And so on.

Wrapping up:

Am I a better programmer than I was a little over a year ago? Without a doubt yes, I have improved substantially. Though the improvement isn't what you would think. I'm not "seeing the matrix" so to speak, as a young me would've embarrassingly put it. At first I ended up regressing and kept refactoring my tiny codebase; it was a slog. Over time I found what worked for me, and become able to hold almost the entire codebase in my head. No, I did not gain inhuman memory simply because I stopped using an LSP, I just became more aware of the code structure, and gained the ability to reason about the entire codebase more clearly and precisely. If I were to present two code samples, one from two years ago and one from today, perhaps even I wouldn't be able to say which is better at a micro level. The real difference is that after numerous refactors, the 200k+ loc codebase I have now feels like a 2k project from my college days; easy to modify, easy to keep it all in my head. Thinking back at my career, most of the codebases I worked on could've achieved that too, but chose not to. The overarching reason is --- in my opinion --- the commoditization of programmers, which causes institutional loss of knowledge. LSP isn't the cause of that obviously, though it does worsen it.

OliveronData··on Ask HN: Programmers who don't use autocomplete/LSP, how do you do it?
I visit HN every so often but never felt the need to comment, until just a few minutes ago.

Yes, what you've described, I also went through something similar.

I've been programming little over two decades now. Been an autocomplete (ctags, intellisense, lsp) user throughout my career. Never had any real problems with them; they were convenient tools, nothing that different from someone using grep/fd etc.

The codebase I (used to) work on is a 30m+ line behemoth, and API frequently changes underneath me; an LSP is crucial to get the work done. How could anybody keep two dozen+ minor-naming and subtle-semantic variations of the same method in their head? I'll say yes to auto-complete any day of the week, I thought.

About two years ago, I've noticed some worrying cognitive signs. Out of the blue I realized I could not remember the name of methods, classes, interfaces, even the ones that I use daily. I could code, there wasn't any problem with that. But I could not write anything down without auto-complete. I couldn't even fill the arguments of a function without the LSP holding my hand throughout the ordeal.

With the realization that both my grandmothers went through dementia/alzheimer's, I truly felt like walls were closing in on me, in real time. Of course, I went for a check up, which came out clean, but I could not shake that feeling of impending doom.

By luck ---and some hefty dose of depression due to unrelated personal reasons--- I started writing a toy compiler in C, something that I had no experience with whatsoever. With just a text editor, because I dreaded having to install visual studio on a two decade old computer (which was the best I had under the circumstances). Despite forcing myself to fumble through, a few days later I noticed that the entire code base was in my head. I knew precisely what I wrote, how I wrote it.

Life went on and I went back to work. And only then I noticed. I was waiting for the LSP to catch up and fill the correct type, my mind went back to my crappy lexer. Programmer? No, I felt like a factory worker on an assembly line.

A month after that I finally gave my resignation, then started a new (self-employed) programming job. For the past year I've been working on a game + an engine, without intellisense or LSP. In fact, I even disabled syntax highlighting a few weeks in, and never turned it back on. I've ditched quite a few of my regular tools, opting for simpler (often homemade) alternatives; simple bash script instead of cmake for example.

Suffice it to say, I've come to some personal conclusions. These conclusions stem from my personal bias, true, but I do feel strongly, that they apply widely to the field as a whole.

In short:

1) LSP, intellisense, code completion, are overall harmful.

2) Syntax highlighting does not work (most of the time), it just satisfies the part of our brains that like to recognize patterns.

3) Complex and highly integrated tools are a net negative unless they are purpose built.

Most people will vehemently disagree with all of the above. That is fine; I'm not on a crusade against the machine, so to speak. I just wanted to share my view since I resonated with the post I'm responding to. But I will still briefly explain what I mean, in case if anyone's curious. (In part 2, because comment was too long)