HNHacker News
TopNewBestAskShowJobs

eis

4,444 karma · joined November 22, 2010

submissionscomments
eis··on Gemini 3.8 text-to-speech
Until the end of the year, then double that. And that's only the audi output though the text input shouldn't cost much in comparison.

Official pricing can be seen here: https://ai.google.dev/gemini-api/docs/pricing#gemini-3.8-fla...

eis··on Gemini 3.8 text-to-speech
Nice but in general I wouldn't put my API key into some third party website, no matter if it claims to not store it. It's not something personal, you are a bit of a celeb here so I wouldn't think you'd save the keys, it's just good general data hygiene. Especially not if a leaked API key can rack up thousands of dollars in fees quickly.
eis··on Stripe's Knowledge AI Platform
Hope this comes over as constructive criticism:

  1. The whole interface feels way too vibe coded with tons of unneeded stuff. Why does it have a console and snake game built in? Why does it have annoying sounds? Why is it full of AI slop writing? The latter is especially confusing because the first person listed under authors is a "Technical Writer". I guess the interface is the general stripe.dev page not exactly related to Kai but the post definitely is pure AI slop writing.
  2. It seems to claim things that might not be substantiated like the following: "When Account Executives use Kai, they produce 2x the sales activity, create 17% more opportunities, generate 26% more revenue opportunities, and close 39% more deals when compared to the same sellers in weeks they don't use it." Correlation is not causation. If sales people have less activity (vacations or sick days etc) then they also wont use this tool much, it doesn't mean all the increase in sales is because of the tool.
  3. the post talks a lot about how great this tool is but it describes nearly nothing of value to the outsider who can't access it. What were the valuable lessons learned? What's neat about it? There is not much meat imho.
It doesn't live up to my usual expectations from Stripe.
eis··on GPT-6 Astra Solves a WWI German Radio Cipher
The question is how was it able to cite the pages if it doesn't know what the content is and can't access it either?
eis··on GPT-6 Astra Solves a WWI German Radio Cipher
I asked Astra to describe the content of pages 214-215 of the source it cited in this article (J. Rives Childs's The History and Principles of German Military Ciphers). It said it can't and it can't find this book online either.
eis··on Mistral X Mozilla: Private, Multilingual AI Browsing
Mozilla lost the tech race and with it a lot of users. Firing a lot of their talented developers didn't help. Now Mozilla is giving up the only reason why they kept a core audience of privacy focused users by sending our browsing data to third party servers for inference just at the time when small local models are getting good enough for many/most use cases like summarization or translation. Chrome provides an API for local LLMs. Apple is trying to run as much interference locally as possible. Mozilla is worse than that. Claiming private browsing while sending data to third parties.

I don't care what their data retention policies are, I don't want my data to be sent to others full stop.

Private means it's mine, it's under my control. Handing it to third parties is not private.

eis··on Sean Carroll explains the biggest ideas in the universe – Full Interview [video] (2025)
He actually publishes a new episode every Monday! Once a month he publishes a special AMA episode which covers a lot of listener (Patreon subscribers) questions and usually is 3-4 hours long! I'm always looking forward to these especially :)
eis··on Sean Carroll explains the biggest ideas in the universe – Full Interview [video] (2025)
I can wholeheartedly recommend Sean Carroll's Mindscape podcast if you want to listen to interesting topics related to Physics, Philosophy, Quantum Mechanics and Science in general: https://www.youtube.com/@seancarroll/videos

He has a talent for explaining complex topics in easy to understand terms and avoids the trap of outrage and controversy driven social media tricks like some others in the field.

eis··on Samsung Debuts zHBM Prototype, Stacking Memory Directly on AI Accelerators
I know consumers hate the situation with ram and storage prices right now, as do I. But at least on the bright side all this AI investment has unlocked a lot of progress in a space that didn't see huge advancements in a good while. All these 10-20% improvements gen-on-gen have resulted in upgrade cycles of well over 5 years for many use cases in order to really feel like it's worth it. RAM capacities especially have felt near stagnant for a decade.
eis··on bzip3
The benchmark is very rudimentary. It does not test different levels/settings apart from its own -b 256/512 (does it affect decompression?), it doesn't measure compression time and memory usage. It does not specify parallel vs single-threaded (it mentions parallel on the one decoding number but what about the others?).

The lrzip test is interesting but it omits for example zstd and doesn't even have (de-)compression timings.

A lot more numbers are needed to present a fair and informative comparison.

I don't want this to be a swipe against bzip3, I only want to point out the presented benchmarks could be a lot better.

eis··on Keep Our Servers Running
This is completely normal. You can't just issue a chargeback because you feel like you want your money back after the fact. They want to know what went wrong. It could be fraud but it could be also misleading checkout experience and other reasons. The merchant is charged often a fee for a chargeback on the order of $15 and it might result in penalties.

So, if you feel like you've been charged unfairly (e.g. product not what was expected), in error, fraudulently etc. then chargeback is the route and you should provide a reason. It's a mechanism to protect consumers. If on the other hand you just want your money back without good reason then chargebacks are not the right tool. The consumer also has an obligation to pay if there was no fault on the merchant side.

Honestly, sounds like everything is as it should be?

eis··on Keep Our Servers Running
Making a wire transfer is easy, but can it be deducted from taxes? That's the tricky part.
eis··on Keep Our Servers Running
Debit cards can do chargebacks just as credit cards can do. That's a feature of Visa and Mastercard (and others). If your bank refuses, complain to the card network. Oh and consider changing banks, that's really unacceptable in 2026 and shows an anti customer stance.
eis··on GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.

Am I missing something or is this not looking too... stellar?

eis··on Elevated Errors for Multiple Models
That's a good point. Seems like Bedrock offers the same pricing while also providing an uptime SLA.
eis··on Elevated Errors for Multiple Models
Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa depending on use case). Personally I like Claude and have the 20x Max plan but even there I burned through the whole weekly quota with 3 prompts in less than a day using the new Fable 5.1 which is crazy. Now Opus 5 is down. These two issues are really testing my patience.
eis··on Gemini 3.8 Flash and 3.8 Flash Cyber
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...

eis··on Gemini 3.8 Flash and 3.8 Flash Cyber
3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...

3.7 used 64M on high: https://artificialanalysis.ai/models/gemini-3-7-flash 3.8 used 120M on high: https://artificialanalysis.ai/models/gemini-3-8-flash

Even their own chart showed more than 2x higher cost compared to 3.7: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...

eis··on Claude Fable 5.1 and Claude Mythos 5.1
According to Artificial Analysis, 5.1 cost 56% MORE than 5, $8523 vs $5455. Yes cache cost is lower but it was MUCH more verbose: 140M vs 83M output tokens.

This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency.

Fable 5: https://artificialanalysis.ai/models/claude-fable-5 Fable 5.1: https://artificialanalysis.ai/models/claude-fable-5-1

eis··on Claude Fable 5.1 and Claude Mythos 5.1
I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big differences? I did notice Opus maybe making more mistakes repeatedly but I don't have hard numbers on this. I hope Fable 5.1 brings noticeable improvements. I am giving it a go now on my 20x Max plan on a problem that Opus 5 has struggled for more than week now and has made very slow progress with regular regressions on the way.
eis··on I trained a small transformer in 1.5hrs and it beats many LLMs
> Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning

Agreed and that's for any benchmark. Private tests are better but you still have to trust the provider to not log and use them for training.

That's why I like when a new set of tests like a new ARC-AGI version is published, that's where you can see which of the models abstracted to more general capabilities instead of being focused on the previous tasks. Most models completely fail new ARC-AGI tests.

The "67 cents" part though is misleading imho. You can't extrapolate from there and think that investing say $100 will get you a lot better results. You hit a ceiling very fast and investing into more compute will give you diminishing results. So yes, you can train a custom model to do somewhat decently on a specific set of tasks but then what?

eis··on Konrad Zuse Museum shutting down due to lack of funding
The person I replied to compared religion with physics (god vs deterministic computable universe). I said those are not remotely equally defensible theories.
eis··on Konrad Zuse Museum shutting down due to lack of funding
> Everything leads to paradoxes or unanswerable questions when you think it through. If our universe is computable then what "computer" is it running on? Why is our universe as computationally strong as it is and not more or less? Is it a simulation and if yes who is simulating it? Is there an infinite tower of universes each running progressively weaker (as in they can do less complex computation) versions of themselves? And if not then where does it stop and why?

None of the things you listed are paradoxes. A paradox is something that is self-contradictory.

> Even the deterministic part has problems, like why does quantum mechanics appear random when it's not or why we see ourselves as having free will.

Something can easily appear random when it is in fact not. Any random number your computer gives you is not truly random. Any hash looks random but is completely deterministic.

> A god is no better or worse than alien simulations, or an infinite multiverse, or the anthropic principle (aka giving up), or any other possible explanation for why the universe is the way it is.

Hard disagree. Not all attempts at explanations are equally valid or invalid. Some are more "out there" than others. And it should not stop us from trying to understand the universe more and more.

eis··on Konrad Zuse Museum shutting down due to lack of funding
The notion of god leads immediately to paradoxes and logical contradictions when thinking it through a little bit. It's fine if people have certain believes but let's not put religion on the same level as physics. No such paradoxes exist with the notion of a deterministic computable universe.

I concur with the OP when they said Zuse is being waaay underappreciated. He is one of the fathers of the modern computational age.

eis··on NSA and IETF, Part 9
I'm confused by your messages linked by DJB. You say that better cryptographers would not choose hybrids, which seems to say that you should indeed think that hybrids are not a good choice. Then you say you are not such a good cryptographer and would choose a hybrid. But if you know that more senior cryptographers think they are not the right choice then why choose them anyways? Or am I misreading "cryptography-literate" here?

Can you explain a bit more regarding your statement that DJB's POV on the matter has no broad support amongst his peers? I'm not in the field but Bernstein seemed like a highly respected member with a long track record in the crypto community, at least from the outside. Do you think the community is wrong or is it DJB who's wrong and why? There's also a good chance that I totally missed the argument being made.

eis··on Gemini 3.7 Flash
Sure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling into that pattern? Earning reports are not to come until end of October, that's not it.
eis··on Gemini 3.7 Flash
3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.
eis··on Gemini 3.7 Flash
Grok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just the release season and/or everyone is benchmaxxing?
eis··on When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
You post your benchmark on every other AI article, I've seen you do this by now more than a dozen times. It's a bit much. I don't want to be too harsh but your benchmark is obviously flawed when the top 3 models for Typescript (Combined) are Grok 4.5, Muse Spark 1.1 (lol), Gemini 3.5! Flash and then followed by Luna, beating Opus 5, Fable, 5.6 Sol etc by quite some margin. In fact 5.6 Sol ranks lower than Kimi K2.7 Code and even Grok Build 0.1. There are so many entries in your rankings that don't make any sense whatsoever that I can't take this benchmark serious and I have not seen it gaining traction. Please stop spamming it?
eis··on Grok 4.5
Google wanted to release 3.5 Pro last month but because of the trouble Anthropic got with Fable they might have wanted to wait a bit for the dust to settle I could imagine. And now there is quite some competition. 3.5 Flash for me is a replacement to 3.1 Pro. It's more like a 3.2 Pro. It costs about the same (or more!) than 3.1 Pro, is a little bit smarter in many cases and a little bit faster. 3.5 Pro will be a lot more expensive and I expect it to juuuust be able to hang with Opus 4.8 and GPT-5.5.

I wish Google was able to actually push the industry further, either in terms of quality (intelligence) or quantity (price) but they've been playing catch up a lot.

They are playing the game a bit differently than all the others. The others have useable IDEs etc. while Google has a boatload of half-assed products.

Google better come out with a banger 3.5 Pro because who would have thought that Grok and GLM would be beating them?

Page 1 of 19Next →