HNHacker News
TopNewBestAskShowJobs

MikhailTal

118 karma · joined June 11, 2026

submissionscomments
MikhailTal··on Google ending ChromeOS support two years early
No. If your engineers can produce more ROI working on another product, it does not matter that now they are 10x more efficient. They will also be 10x in the other, more worth it product
MikhailTal··on Claude Opus 5.5
> This is a classic test request

Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

MikhailTal··on Show HN: Lossless-memory – a personal AI memory that never summarizes
No? from a quick skim it doesnt look like it goes into the system prompt everytime, you just use search/grep over it. Pretty much most memory approaches relying on a big set of info where agent chooses what to 'recall'. Its like any other tool
MikhailTal··on Asking authors about their own papers
This is very bad logic

1) Humans also are trained on a subset of human knowledge. 2)A lot of papers are just about experimenting something, and then applying simple stats. Eg empirical studies, around 1/3rd of published papers. Like, we tried this drug or did this experiment, from a sample size X here are the results. An expert is needed to maybe comment on the conclusion/hypothesis of the underlying suspected mechanism, but LLMs are still very useful on catching bad statistics or p hacking (so so common)

MikhailTal··on Claude Cowork and chat are now one Claude
This is technically true, but when people talk about randomness, its not only about same input-> different output, like temperature>0 and the things you said.

Its also about very similar inputs -> different outputs. Even with everything you said, yes, same input would result consistently into same output, but sliightly different input and you might get completely different/semantic answer.

MikhailTal··on Tell HN: OpenAI brings back 5 hour limit for plus and business standard users
Aware that this is essentially a soon to be deprecated service. As any vendor you need to know the life expectancy of it, plus forecast/predict any pricing changes. You really do not want to become dependent. Even if competitors exist, there still some non trivial cost to switch
MikhailTal··on Nitter and XCancel resume service after legal advice
Again, this is HN bubble stuff. The vast majority of people do not mind oauth or account.
MikhailTal··on Nitter and XCancel resume service after legal advice
> if you hook up to the fediverse

No this is the hard problem. HN people live in a bubble. Anything else than signup with oauth and then having access to everyone and its too hard and 99% of people will click away.

MikhailTal··on Mechanical Turk shutting down September 30
As usual, the real tldr is in the privacy policy https://generalresearch.com/supplier/privacy-policy/notice-t...

> General Research Laboratories, LLC (“GRL”) is an aggregator of market research surveys in multiple marketplaces for business customers. GRL does not typically host consumer surveys which are conducted by other consumer-facing organizations....

So, just a middleman to give you fake/shady at best survey responses to pad your numbers, so big enterprises/concultancies can have data that say whatever they want. The rest of the website is just BS

MikhailTal··on Headlong: A microharness for persistent agents
Very fascinating, super interesting engineering. Although i do find it very funny how they just bypass a massive vulnerability, basically zero data isolation (even between good actors, let alone bad ones) with 3 sentences. Only in the llm space you can slap a massive limitation like this in the middle of the article and continue like nothing happened

> Whatever anyone tells Audel becomes part of the single experience that every other conversation draws on. In practice, Audel is bad at keeping secrets. Ask it what it’s been working on with someone else and it will often just tell you, even though we’ve asked it not to. We also haven’t studied what happens when two people give conflicting instructions. For now, we assume anything you tell Audel is shared with everyone on the team.

MikhailTal··on New MCP Roadmap
Not all agents have access to a sandbox/cli/code execution environment to run arbitrary api calls etc. MCP helps by essentially having another tool call without needing a sandbox. If you do have a sandbox, then might as well do codemode if you insist on mcp https://blog.cloudflare.com/code-mode/
MikhailTal··on Launch HN: Speko (YC S26) – OpenRouter for Voice AI
What is the difference with Livekit Gateway? https://livekit.com/blog/introducing-livekit-inference

Or even something more managed like Vapi?

MikhailTal··on How Compaction Works in Pi
This is terrible. Models have been RLed on looking at the previous tool call chain, and reasoning. No chance this does not reduce performance. The point of compaction is that it also includes useful signal from the tool outputs itself so agent does not repeat it afterwards
MikhailTal··on Gemini 3.7 Flash
When a company gives away service a heavily subsidized service as a promo, the full cost of serving it (compute) can get classified as sales and marketing instead of just cost of revenue, which makes your gross margin look better!
MikhailTal··on Docker Sandboxes – Disposable, isolated sandboxes for AI agents
https://github.com/e2b-dev/e2b
MikhailTal··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
> Google DeepMind: We are building strong momentum: Flash is in high demand, our Cyber model is live, and Gemma models have surpassed 900M+ downloads

Considering these are the best stats they could find, gemini usage+general situation must be really, really bleak.

High demand means nothing. A model being live is nothing to brag about. And gemma downloads also can be from auto CI pipelines etc. Nothing concrete

MikhailTal··on Launch HN: Tokenless (YC S26) – Automatic model switching to save money
It's a smart approach, definitely interesting. It all hinges on quality of course which im not convinced.

τ³-Banking is the only one which you show better accuracy and cheaper. If i'm reading the blog results right, for deepswe and terminalbench, you are worse+cheaper than frontier, and better+more expensive than just small models. Which is exactly what i would expect even for a router that switches at random.

Speaking of random routing, this would be a great ablation study as well. What about also if you route each request to a tiny 7B model classifier? Why is your approach SOTA?

MikhailTal··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
this is /goal in claude code/codex. also basically a slightly improved ralph loop
MikhailTal··on OpenAI and Hugging Face address security incident during model evaluation
is this really that surprising?

Exploitgym prompts are tuned for a model to do everything it can to achieve a cybersec/exploit task. And we know that models are good at finding vulverabiltiies.

Its just random that the sandbox itself was buggy. But all that happened here is that we told a model "do everything you can to achieve your goal of hacking X" And it just hacked Y as a roundabout way of hacking X.

Imo its PR for OpenAI to also start the mythos class mysterious unreleased model hype.

From HF statement: "AI safety won't be solved by any single company working in secret". So now we have TWO companies working in secret

MikhailTal··on Searchable field-level encryption on Supabase with CipherStash
the OP asked why, you more described the how.

What is the benefit of all this?

MikhailTal··on At least 105 past YC founders have worked at OpenAI and Anthropic
"Lottery winner says you should not gamble"
MikhailTal··on SpaceX bond worth 10% less than issue price – heading for junk bond status
Also, the OP just does not understand how the market works anyway. Surely if it was obvious that investing in fresh IPOs is a bad move, all of the big boys (banks, hedge funds etc) would short them to the point of equalising anyway. Maybe not to the absolute efficient point, but still, why do people think they can see such a huge obvious trend, and also assume that other people cannot see it?
MikhailTal··on GLM-5.2 is the new leading open weights model on Artificial Analysis
This is not a new situation. This was happening also when good vision models like alexa net were coming through, especially for OCR. Companies had choice between cloud or self hosting with GPUs. But turns out, problem is usage patterns.

Your usage will peak during certain timezone work hours(even if you are a huge multinational company most of your engineers/users tend to be from only a few locations), so then you have a bunch of gpus doing nothing the rest of the day. especially with latency sensitive stuff, this is a decades old tradeoff problem, its not unique to llms