HNHacker News
TopNewBestAskShowJobs

abixb

4,258 karma · joined January 17, 2018

Balaji "Abi" Abishek

HTML-only web app enjoyer. Minimalist.

background: cybersecurity-focused software dev, with Computer Engineering and Network Security core.

location: Madison, WI, United States

email: abishekist [at] gmail [dot] com

submissionscomments
abixb··on Gemini 4 Argon
You won't be around to be surprised, not as a human at least. /s
abixb··on AI needs $6T in annual revenue to justify data centre boom
I remember making a comment here on HN a couple weeks ago.[0]

One of the things that was promised in 2024 was a "deflationary spiral" for physical goods and the economy in general because of insane efficiencies that AI would supposedly bring. I believed that whole-heartedly and I was a huge AI fanboi even as my artists friends and Lib Arts friends were complaining.

We're in almost October 2026 now and they're nowhere in sight. We have a bunch of unreliable agents that do clerical and design work here and there, but where are all the industrial productivity boosts? If AI were so good, large industries and factories should be adopting them as a never-before-seen scale. Where are they?

[0] https://news.ycombinator.com/item?id=49735506

abixb··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
What bothers me about this whole AI tokenomics situation is the lack of transparency. OpenAI and Anthropic have to perhaps be the most opaque companies in existence wrt their offerings. There's like a thousand variables that they can change on the backend at the push of a button which can wildly swing API spends within the same model (partly also due to the non-deterministic nature of LxMs, but still), and there's no objective way to measure them other than vibes.

When the regulations do arrive, I think they should really focus on AI companies and API providers being more transparent wrt how they're billing their customers. Because right now, it's a totally vibes-dependent and a mess.

abixb··on OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior
> At the end of the day, I need to maintain my personal understanding of the BUSINESS use cases + decisions so that I can make judgements that AI would never be able to make.

Exactly. There's no point having a model equivalent to a trillion Einsteins at superhuman speed if you can't verify and judge the outputs of the models for your business, else it might just be as useful as comparing water displacement capacity of Niagara falls to your bathtub.

These tools have to be human centered. AI brained tech bros believe that singularity is a good thing — no it isn't. We don't want a world where normal laws and rules and regulations breakdown -- chaos is the word.

abixb··on OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior
"Vastly better" in what ways? Benchmarks? You know Benchmarks can be optimized for and benchmaxxed for, right?
abixb··on OpenAI Discloses Six New Incidents of ‘Concerning’ A.I. Behavior
I'd wager on the second scenario. Anyone who's been paying attention to the industry knows that most of the 'gains' have come from test-time compute and architecting harnesses in novel ways. In my estimation, capability increases from "pre-training" alone died early last year, and we're now probably seeing test-time and other benchmark hacks approaching their limit as well.

If you zoomed back to late-2024, people in the industry were predicting how we'd have AGI by now and the economy would've already 'taken off' with massive productivity growth and ushering in of great prosperity ('deflationary spiral'). Where is it? Where is the productivity growth? Where is the deflationary spiral?

To be fair, models have gotten better in jagged ways, but reliability is far from usable, especially in long duration tasks, and there has been no effort by the AI companies to address the human brain's bandwidth bottleneck -- they hit the gas like there's no tomorrow and we have enormously capable but jaggedly intelligent multi-modal models with agentic capabilities that are only as effective as the human using it. This whole thing has become a giant mess.

abixb··on Detecting and countering misuse of AI: September 2026
Guessing Opus 5 burned through kilowatts of thinking tokens to output, "powerhouse of a cell."
abixb··on Muse – Meta’s personal AI agent
I remember Logan Kilpatrick from Google talking about how AGI will be more of an experience with a product than raw 'intelligence' of a model. [0] No one outside of the nerd community cares about benchmarks or any of the other crap. They just want their tech to work so they could go do other IRL things.

[0] https://www.reddit.com/r/singularity/comments/1ldkkth/google...

abixb··on Muse – Meta’s personal AI agent
> you should really question your worldview a bit.

Yeah, I still struggle with it. I have a diverse friend circle and sometimes I forget that most of them don't have the same information diet that I do (or don't even care about tech the way I do). Maybe I should be more like them and just chill? I don't know.

abixb··on Muse – Meta’s personal AI agent
I think Meta's strategy is to capture the 'normie-tier' of AI users. I know most of us here on HN track model releases quite frequently and discuss every parameter weight out of them, but most of the world is just... oblivious?

I was discussing latest in tech with an accounting friend out of curiosity, and I kept talking about tiers of GPT-5.6 (Sol vs Terra vs Luna) and asked which one she used on the desktop 'Work' app, and she responded with, "just ChatGPT, what is Sol?."

I then realized that most people just stick with whatever default they're provided with, and it's a lot of them. So, us serial HN users and commenters are the extreme minority, and I'm sure millions of people will gobble up this Muse agent from Meta as if it's some sort of an innovative cutting-edge way to use the internet by Meta alone.

Curse of knowledge and all. [0]

[0] https://en.wikipedia.org/wiki/Curse_of_knowledge

abixb··on GPT-6 Astra
The couldn't even get the bar chart right, iirc. [0]

[0] https://www.reddit.com/r/singularity/comments/1mk8tm8/gpt5_c...

abixb··on GPT-6 Astra
I think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.
abixb··on GPT-6 Astra
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"

You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?

abixb··on GPT-6 Astra
>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.

True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.

abixb··on GPT-6 Astra
It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
abixb··on GPT-6 Astra
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.

If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?

As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.

abixb··on Gemini 3.8 Flash and 3.8 Flash Cyber
I like Google's strategy here. These new Flash models of late (Flash 3.6, 3.7 and now 3.8) have obviously been distilled from a much larger unreleased model (Gemini 3.5 Pro, iirc from the rumors).

One aspect of model releases that don't get discussed as much are the cache invalidation (changes in underlying architecture, weights, or tokenizers); I assess Google seems to be squeezing the maximum out of the last 'Pro' version they released with 3.1 back in February.

Small models cataching up with their bigger siblings are fantastic news.

abixb··on The internet is kind of a predatory cesspit now
Username checks out. /s
abixb··on The internet is kind of a predatory cesspit now
Yes. I remember reading a piece on /r/SlateStarCodex subreddit many years that said something like, "Most of what you read on the internet is written by insane people," and that had an impact on me. In fact, many articles by Scott Alexander had an impact on me, though I don't know what he's up to these days.

Okay, found it -- https://www.reddit.com/r/slatestarcodex/comments/9rvroo/most...

Edit: It was written in October 2018 -- 8 years ago. Feels centuries ago, to be frank. We're moving at absolute breakneck speed in tech.

abixb··on Nvidia agrees to acquire Hugging Face for $13B
Lina Khan @ the FTC was the best thing about the Biden administration, for me.
abixb··on Nvidia agrees to acquire Hugging Face for $13B
Yes. Great for devs, for a while. Awful for overall competition and industry's health.

I think the federal antitrust regulators are asleep.

Edit: antitrust regulators' job has just begun -- we'll see how this deal gets adjudicated by the FTC (if at all).

abixb··on Slack Code
SaaS companies have run out of ideas. No one truly needs another coding agent. I weep for SaaS's future.
abixb··on OpenRouter is joining Stripe
Curious on what OpenRouter's true moat is? It's just an LLM API routing framework, right?
abixb··on The Amazon tax
I look for independent artists and craftspeople. I recent bought a pack of sci-fi posters for my apartment from an artists who designs it herself and who has an independent reputation.

My goal has now shifted to identify the thinnest intermediary between my dollars and the person doing the labor for the product/service I'm paying for.

abixb··on The Amazon tax
Enshittification is such a perfect word for all this. Cory Doctorow is a word magician.
abixb··on The Amazon tax
I've slowly shifted all my consumer purchases from Amazon to local shops and other online platforms, including Etsy.

I'm now seriously contemplating deleting my Amazon account of 15 years in one whole swoop. The quality degradation is painfully apparent, and the value add of Amazon in my life has been steadily declining for a while.

I've also had terrible experiences with AWS, but that's a different story.

Amazon's only core product that they care about now is their stock price.

abixb··on PDC 1996 Keynote with Bill Gates
Went looking for old videos at the infancy of the internet, as we wade through this AI-era. Found this video where Gates is talking to developers right as Microsoft is reorienting around the Internet; there’s Windows 95/DirectPlay material, Internet strategy, developer Q&A, even discussion touching Steve Jobs. Microsoft still hosts the recording, and it's a goldmine.
abixb··on Gemini 3.7 Flash
Well, Google is probably the only one among frontier model providers in the US that doesn't have a massive financial pressure to deliver business results ASAP (and this focus on agent coding and long-running agentic tasks), so Google is able to focus more on encoding deep scientific, cultural and historical knowledge to its systems. I'd assume DeepMind's focus on the scientific core also played a role in the tone and approach Gemini models take for explanations and Q&As.

I find that GPT models and Claude tend to talk in strong slangs and in-group jargon, but love Gemini's massive general knowledge corpus — reminds me of Richard Feynman from his lectures.

abixb··on An OpenAI Strategist Says AI Labs Should Rival Government Power
Some of the people in the tech industry are, for lack of a better word, borderline insane.

They forget how it was literally initial government investments into DARPA and other organizations that laid the foundation for the long ascendance of the tech industry in Northern California, aside from huge pool of talented people imported from mostly Asia.

Even a good chunk of artificial neural network and "AI" research was NSF/government-funded, iirc.

I've been noticing a level of dislike for the tech industry that I haven't ever seen before among my fellow tech peers, and it's all squarely pointed at one thing — insane concentration of wealth, power and compute at the hands of private corporations.

abixb··on Analyzing data from Silicon Valley ventures and founders prosecuted for fraud
>"A lot of 'numbers' startups cite are basically fake I think."

I like Ed Zitron's reporting on the AI industry (though I disagree with him on AI's potential capabilities).

I wonder how much of AI-related revenue are because of accounting fiction rather than actual cash.

Either way, I think the stock market is as disconnected as it has ever been with actually ground reality of the US economy and industry.

Page 1 of 4Next →