HNHacker News
TopNewBestAskShowJobs

manquer

4,909 karma · joined December 9, 2015

submissionscomments
manquer··on Sonnet 5.5
Context needs to pre-filled into a GPU memory in a node (usually 8xB300 or 8xH200) so there isn't any context or cache sharing between model families given their different parameter sizes, tokenizers, unlikely they are co-located in the same node.

Sub-agents not sharing context is a useful design-pattern when you want adversarial or independent reviews.

Cache reads could be shared between sub-agents, A single node(8GPU cluster) supports few hundred concurrent user sessions, that all share the same KV cache memory, so it is likely model providers do colocate your sub-agents in one node, it is more efficient , but may not be guaranteed so performance could vary; like we have with elastic compute and storage[1]

This can be cheaper depending on your coding flow i.e. cache hit % and the billing plan - cache reads are basically free or charged very little in subscription plans.

[1] Modern AWS does offer collocation at additional costs for compute but that is not the default and most other clouds do not offer it

manquer··on Japan moves to tighten rules for foreigners
Most immigration seems to me as some form of wealth differential to attract the relatively poor foreigners giving them less rights than local population.

Middle East runs on poor people from the subcontinent and Africa.

Hong Kong depends on live-in foreign labor who have to camp out in public spaces on their day-off.

America is built on poorer labor of every kind - enslaved African cotton farmers , indentured Chinese railway workers, famine stricken Irish, or more recently Latin American or H1B Indians in tech.

Some countries try to attract the richest percent like New Zealand, however it is not economically meaningful nor is major portion of immigrants globally or within their countries even.

Countries may not like foreign poor immigrants, but they certainly need them to work harder for less pay and rights.

manquer··on Italian parliament votes for return to nuclear energy
Overnight costs do not get impacted by guaranteed purchase price or subsidy for it neither does it include interest rates. It is a specific aspect of the TCO.

Nobody is saying nuclear is cheap it is not - we will be building a plant once in while for pure strategic reason no matter the cost anyway. The point is construction cost comes down per MW with higher frequency of plant building, and this is the case SMRs, we can get cheaper construction both Per MW and overall TCO in absolute dollar terms relative to large plants with SMRs.

manquer··on Italian parliament votes for return to nuclear energy
That is why even Flameville 3 was cheaper than Hinkley point C, not that it was cheap.

Flameville 3 is the first reactor for France in 21st century and suffers the frequency problem I am arguing. It was a path finder therefore only one was ever ordered. EPR2 is 6 under construction and France does have a track record of doing it consistently at scale for cheap(by nuclear standards).

The Asian giants all do it for cheaper because they build regularly, Japan (post Fukushima policy direction notwithstanding) and South Korea achieve overnight numbers of $5-6M/MW or less, so does China.

Regular construction is key, to the industry having a chance, without SMR that is never realistic anymore.

manquer··on Italian parliament votes for return to nuclear energy
Vogtle, Hinkley Point C- are all one off constructions and have very expensive regulatory compliance requirements

AP1000 (reactor as Vogtle 3/4) in Taishan is estimated to cost $4M-6M /MW compared to $15M/MW+ at Vogtle. if Chinese safety standards seem lax, South Korea does regularly reactors at $4-6M/MW range.

Even Flameville 3 in France was significantly cheaper than Hinkley point C (same reactor type), which infamously required some quite ineffective Fish speakers. The EPR2 series (6 reactors) in France is expected to cost considerably less than Flameville did.

Key to cheap nuclear - SMR or otherwise is regular construction. SMR makes regular construction more likely as each plant is not a $10-20B+ capex that is a tough sell to tax payers particularly when it inevitably overruns, SMR also potentially have a bit less regulatory hoops to jump through due to their smaller size.

manquer··on Claude Opus 5.5
Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain
manquer··on AI coding has made CI a bottleneck, so we reworked ours to keep up
> calculate the merkle tree forward. he blob itself can stay on the remote cache server.

Not sure how that would work with building say a docker image, reproducible builds are pretty hard problem to solve, and caching intermediate layers is not always simple or even doable, we typically still need to publish to a registry which is not the cache server.

manquer··on Python Workers are now generally available
Have we come back full circle back to GAE [1], launched in 2008 with Python 2.5 support ...

[1] https://googleappengine.blogspot.com/2008/04/introducing-goo...

manquer··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Typically you need the intermediates to compute if you need the next one , so you cannot skip to final until you have the intermediaries.

The final asset/artifact is rarely small either. even best optimized artifacts can be few hundred MB docker image or more commonly multiple image layers running GBs in size .

each step is a network pull then recompute cache if stale and keep going till end .

manquer··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Can't comment on Bazel specifically, but having worked with both nx and turbo, the bottleneck was usually network and disk IOPS rarely compute.

Even fully cached outputs needs to fetched and read from a remote server[1]. A step n-1 outout fetched from remote cache server need to written to disk and then again read by step n[3] - all disk I/O and network bound operations.

10s may be achievable/realistic goal in the Java/C++ world where Bazel normally seen. In TS eco-system most people would be over the moon to get into ballpark of 1-2m for a decently large monorepo.

We should define Build more clearly here, if you mean running just transpile/compile steps or the full series of steps that includes tests (as the linear post here is talking about). It is hard to see even a small sub-set of a large suite of test that require a virtual DOM or a real browser can run in 10s or less.

[1] Typical for say managed CI setup .

[3] Common run-of-the-mill frontend + backend stacks in different languages etc.

manquer··on San Francisco Onion Futures Company
Just like solving math or programming in a general sense is much harder than a specific solution , so is passing a broader law.

We don’t complain about switch cases in code when there is two or three switches we start refactoring once it starts to proliferate.

The law is no different , passing a wider ban would not get the votes easily or quickly and the interested parties the onion industry have no reason to push for it neither does the lawmaker acting on their interests.

If said small markets also had exceptions passed seeing the onion one there could have been case to be broad.

It would premature optimization to otherwise, based on just need for elegance , code or law has to work first even if dirty .

manquer··on Nvidia announces native GPU programming in Rust
I think implication being organizations with 40,000+ employees and even more consultants and contractors plus a lot of budget are also using LLMs to draft public facing content instead of paying for content writers or even just proof readers .

It points to friction rather than cost economics. Same reason we are always surprised why multi billion dollar product companies with millions of install base prefer electron instead of a native app.

manquer··on Why I'm still bearish on LLMs after Navier-Stokes
There are more possible game combinations than atoms in the universe, even those generation of valid game states are as you say pre-defined. that is why models cannot go this route and therefore are poor at chess
manquer··on Why I'm still bearish on LLMs after Navier-Stokes
> People who are good at it rely more on experience and deep domain expertise

People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.

A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.

manquer··on We got admin access to Baseten's production GitHub in 25 minutes
> free work

Don't know if I would call it that ?

This was a potential customer reporting a result of an audit of a tool they are evaluating. This is frequent and normal activity in enterprise deals. Most of the time such reports are not critical vulnerabilities it would things like tenant configuration -what business would like versus what CISO will accept or risk acceptance of the product they are buying with monitoring or other prescription on access restrictions or a DPA and so on.

It would be novel business model to spend ton of money in getting a prospect to late-deal stage where they are ready to do a security audio for you just so that part is "free" .

Most companies wouldn't disclose(to the public) even if it was serious , that is not their job, they will report to internal teams and re-review on fix. Strix.ai has a benefit in doing so as they sell a scanning tool for this purpose so we get to hear of this.

manquer··on We got admin access to Baseten's production GitHub
>Wallets usually belong to real people with lives.

So does data ? it belongs to real people.

I would imagine baseten's customers and eventually their end-users[1] were also grateful that their data was not compromised here and the disclosure was responsible.

[1] There is a pretty good chance you and I could be using services who are using baseten

manquer··on We got admin access to Baseten's production GitHub
Swag packages like these are a token of appreciation not a reward.

The front page post in HN here is worth far more than few thousand dollars , don’t think either organization is operating under purely financial transactional nature .

Most people who find a dropped wallet will return it without evaluating the market value of your compromised identity or the contents of the wallet .

Grateful owners may buy you a beer that doesn’t make them cheap , not everything is evaluated in purely money terms, and that is a good thing ?

manquer··on Reverse-Engineering Claude Web's MicroVM: Uncovering Anthropic's Hidden Antspace
Repetition at the start or end (i.e. Anaphora/Epistrophe) are not used lightly in prose or verse, they serve a rhetorical purpose - usually add rhythm or strengthen the theme.

The type of repetition Claude uses is what Fowler[1] called "elegant variations" and discouraged in modern style guides for a good reason - people find it really annoying.

[1] Henry not Martin

manquer··on We must pace the frontier
As long you are burning more cash than you bring in, you are beholden to investors - whether it is retail, VCs, banks or a government giving you a bailout it is still someone signing you a check.

Golden shares, vetos, PBC, charter are all paper tigers , they only matter if/when the firm is self-sustaining business with no outside capital needed, the alternative to not listening to investors till then is crash and burn.

After that point, you will have to listen to the paying customers (sometimes but not always they are also users ) as they are ones now funding your organization.

Bottom line you are always listening to someone.

manquer··on A Tesla ran a stop sign and killed a man, Full Self-Driving/Autopilot was on
If safety is the objective we should be prioritizing public transit(automated even better) and not discussing agent/human driving a personal vehicle, here we value freedom[1] over safety, my comment is only a reflection of that ethos.

[1]Whether we like it or not individual freedom reflects in everything from gun safety(unique in the developed world) to the state of public transport in the richest country in the world.

manquer··on A Tesla ran a stop sign and killed a man, Full Self-Driving/Autopilot was on
You are talking about enforcement. I am saying there are no AI laws on books, there is no established liability for the self driving software developer, the liability is still to the driver solely and completely.

First lets have those laws and regulations for AI and we can THEN talk about conviction rates for humans(or agents) and effectiveness of enforcement.

Right now human and agents are in two different leagues of enforcement civil and criminal, so they cannot be compared and agents should be held to much higher standard than humans until that changes.

manquer··on Mamdani bans AI in NYC schools
Exactly why a blanket calculator ban for entire education system would be bad. It should be course specific decided by someone with pedagogical understanding of course contents ?
manquer··on A Tesla ran a stop sign and killed a man, Full Self-Driving/Autopilot was on
Humans can go to prison [1], and depending on the state there are specific vehicular homicide laws or they are could be charged under normal homicide laws, specific statutes exist because vehicular laws can have higher punishment than typical negligence/reckless standards.

In California the charge can be anything from vehicular manslaughter to second degree murder, that means 15 years-to-life without parole, depending on other factors such prior(murder) convictions or if the victim is a peace officer and so forth.

In comparison self-driving systems do not have any punishment mechanism in the penal code. There is not even a well established liability standard, or ban/suspension threshold or any regulatory body equipped to audit a self driving system specifically [2]

[1] if found guilty - we may disagree on the mechanisms, effectiveness ,conviction rate and other details, nonetheless there is enforcement and there are going people in prison for this type of crime everyday.

[2] yes NTSB can do investigations but only for accidents (and give recommendations) and also has just 400 employees and $150M budget to cover everything from aviation to marine or pipeline and rail accidents.

manquer··on OpenAI's GPT-6 Astra on ARC-AGI-3
You missed the hard part getting on HN front page , ie. Getting the acceptance of the community / zeitgeist .

There is no incentive for OpenAI to subsidize is you if no one reads /reports on your benchmark . They are only going to fund a few that are currently popular .

Community acceptance doesn’t automatically mean the best , it is combination of some level of technical quality and the ability of the promoter to socially influence or get support of influencers .

manquer··on Mamdani Bans AI in NYC Schools
> calculators are banned from math classes

Not all math classes yes, if your are learning basic arithmetic yes calculators should not be allowed, but if am learning say applied statistics I would be wasting time adding numbers by hand and not learning what I am actually supposed to be taking the course for. A tool is just a tool[1].

Even when it was not allowed, schools would still teach you how to use a calculator correctly and your peers that 80085 looks like BOOBS. Effective AI usage education specifically as an aid to learning is important for children today. There is no avoiding the fact that generative models will part of their whole lives .

The instructor or the school board with inputs from professional experts in learning pedagogy should be taking this kind of call at individual course level. IMO it shouldn't be the mayor doing a blanket ban

[1] There were time when instead of calculators, you would need to carry Clark's Tables booklets you weren't expected to learn numerical analysis to just use value of ln(n) , Clark's tables were a tool, so was the calculator, or a computer or now a machine learning model.

manquer··on Claude Fable 5.1 and Claude Mythos 5.1
The implicit point being adding this type of safeguards to Fable dumbs down the model in measured performance even though it is not fundamentally different.

Note it may not even be actual performance, typically in most benchmarks the model would be scored zero for refusing a task just the same as not completing it, so it could just be the Fable's stronger safeguards is just making it refuse more or perhaps even drop down to Opus.

manquer··on Nvidia projects $673B in sales as AI demand widens
Bubble means inflated not fake, i.e when they make less money their stock will crash and bubble will pop, which is what people buying the stock today are concerned with , is this going to hold
manquer··on Meta reaches $17B settlement over social media harms to children
It is over 10 years, so cash flow impact is ~1.8b per year if it was equal installments. Given that 30% is conditional on TikTok and YouTube agreeing to similar controls and paying similar amounts and other payment terms the number might be even lower for 2026.

Share prices are weakly correlated to news events anyway, if you ask the professionals they would say it was already "priced in" that is why a news event is not moving the stock when it doesn't move, and when it does move they will it is news event causing the move.

manquer··on New Mac mini, featuring M6 and M5 Pro
Sadly maybe 20 years ago $17K was price of a car. Median car price these days is close to $50K and there are hardly any new cars less than 20
manquer··on OpenAI Jalapeño: Better than Nvidia Blackwell
ASICs always do better than general purpose chips. General purpose chips is turtles and turtles of virtualization and have to consider 4+ decades of backward compatible instructions set support.

ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.

General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.

New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.

Page 1 of 34Next →