HNHacker News
TopNewBestAskShowJobs

Majromax

3,544 karma · joined February 1, 2019

submissionscomments
Majromax··on Owed a billion dollars in Nvidia stock
> You would need sufficient evidence to support a claim that NVIDIA intentionally lied, which you obviously don't have otherwise you would have mentioned it in your post.

Would even an intentional lie act to to reset the limitation period here? The hypothetical lie wasn't a deep secret exposed by some whistleblower, it came to light by... reading the vesting agreement. Since AFAIK limitation periods run from "know or ought to have known," I can't see a viable construction to keep the dispute live after 30 years.

Majromax··on Owed a billion dollars in Nvidia stock
> They should just offer to settle at a reasonable value

Since litigation is costly, the acceptable range for a settlement is centered around the expected outcome of a trial, plus or minus each party's cost of litigation (including opportunity cost).

In this case, "the claim is barred by the statute of limitations" implies that the expected outcome of litigation would be approximately $0. The net range for a settlement is then the 'nuisance value' of a lawsuit including any PR damage for airing the case publicly; that would be orders of magnitude below the $1bn claim.

Majromax··on Yes, Claude can do nine loops
> This feels like an obvious jab at Navier-Stokes, but where we get to shift the numbers around to hide where the compute actually is being spent ... compute is being spent. It's either being spent in amortization to make the search smarter ahead of time, during training, or its being spent after.

I think that argument is recursive? These posts aren't very complicated for either of us, but they're written on devices that are fabricated with billions of dollars of semiconductor equipment. At what point do we just acknowledge that we stand on the shoulders of giants?

To me, the distinguishing factor is that the expense not special-purpose but upfront. The model here is trained without foreknowledge of what problems it will solve. Solutions like nine loops are genuine expressions of a pre-existing model capability, even if that capability has not pre-existed for very long.

Majromax··on A custom virtual machine for the Stars 4X game
I can't speak for the authors' mindsets, but the Stars! map had pointwise movement rather than via strict stellar/planetary waypoints. When fleet interceptions can happen midcourse[1], there's not a whole lot of design room within the Win16 style.

[1] — If I remember correctly, there even was some special-case code for ship pursuit cycles.

Majromax··on A custom virtual machine for the Stars 4X game
Wasn't the shareware version fully featured, such that only the key was necessary?
Majromax··on WeatherNext 3
> The cyclone prediction thing is very interesting to me in particular (not quite sure how you go from the ML matrices to "here's a path the cyclone might take")

In a high-level view, it's the result of specialized decoding heads.

Traditionally one would take gridded forecast outputs, then process those with comprehensible actions like "find all local pressure minima in the ocean, then filter to ones which correspond to warm cores, etc." to infer (diagnose) the presence of a cyclone.

One problem with this is that gridded forecasts suffer from known biases and tradeoffs. For example, a forecast on a ~25km grid is just on the edge of being able to represent the eye of a hurricane (50km scales), and it certainly can't accurately represent the sharp transition of wind in the eyewall. That means that the forecast winds are almost certainly a smoothed (and therefore less intense) version of what observers would see.

The WN2 approach (paper: https://www.nature.com/articles/s41586-026-10953-2) adds a direct readout head to the model: given latent-space access to the full forecast, it tries to predict the bona-fide cyclone observations (https://www.ncei.noaa.gov/products/international-best-track-...).

It's kind of like a post-processing or bias correction (see for example https://www.ecmwf.int/en/about/media-centre/aifs-blog/2026/a..., which applies in physical space), but by having access to the model latent space and by being included in model training it is (probably!) higher-quality than a pure, after-the-fact approach.

Majromax··on WeatherNext 3
> It’s crazy we’ll never have forecasts as good as dark sky again.

'Nowcasting' is an area of active research, both with machine learning and with physics-informed or visual flow approaches.

Part of the problem from the machine learning side is that these are _huge_ problems. NVidia's StormCast (https://research.nvidia.com/publication/2024-08_kilometer-sc...) works globally at kilometer scales, and you can imagine how big those grids are. Even with patch training, you're dealing with very large datasets.

At the same time, this is not exactly a high-profile area of research. National weather centres focus on actionable medium-range weather predictions, and meteorologists can look at radar themselves and perform mark-one-eyeball predictions for very short-range watches and warnings. Some private-sector actors will pay for short range predictions, but they're often looking for something hyperlocalized (e.g. weather at this particular construction site, for crane safety) or specialized (near-real-time cloud and wind predictions for renewable energy).

Most public-accessible weather predictions are downstream of either a public-sector effort (which doesn't internalize benefits, leading to under-resourcing) or a byproduct of another private-sector offering.

Majromax··on WeatherNext 3
A barometer will give you surface pressure, but that's a field that tends to vary relatively slowly over the surface of the Earth. The calibrated weather stations that exist at every airstrip do a reasonable job of providing these conditions over land, and the residual of "pressure from phones" probably won't help all that much.

The data that would be most valuable to initial conditions is upper-atmosphere winds -- this is the kind of data given by weather balloons. In clear air there's no great way to measure this from either the ground or from space.

One important supplemental data source here are aviation reports, from planes flying at altitude and particularly trans-oceanic routes. When air traffic was largely curtailed during the early phase of the Covid pandemic, weather forecasting suffered a bit for the lack of data (see eg https://www.ecmwf.int/en/about/media-centre/news/2020/drop-a...).

Majromax··on Google Antigravity TOS: 3rd party usage can get Google account suspended
That line of reasoning has no end. If you use Antigravity on anything other than a Google Chromebook or Pixel, the hardware is a 'product not provided by them'. Is that a TOS violation?
Majromax··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
If your maximum addressable market is “the whole economy,” as seen in SpaceX filings, then a city-sized call centre (distributed, of course) really is ‘t that much of an ask.
Majromax··on I'm becoming AI-blind
> But it’s just as likely to make an output better.

No, for any particular output token the model's true logits are definitionally the 'best' that the model can achieve.

This is inherently probabilistic. The model's top-1 guess is not guaranteed to be optimal, but it should be so a proportionate fraction of the time. Same with the top-2, top-3, etc.

Watermarking necessarily alters the output distribution away from the model-set distribution, and that alteration is inherently 'worse' in expectation.

You can liken this to a weather forecast. If there's a 25% chance of rain, the forecast should say so (or a 'sampled' deterministic forecast should predict rain 25% of the time). If the forecast is 'watermarked' and predicts rain 27% of the time under identical circumstances, it's a worse forecast.

That being said, this is a case of hiding a message in a noisy channel. Watermarking only needs to communicate one bit ('yes watermark'), so the effects can be arbitrarily small provided one is willing to tolerate an increase to the text size needed for reliable detection.

Majromax··on If this is true, the hyperscalers are toast
> Hyperscalers don't run computing at some multiple more efficient than on prem.

I'd disagree here. I see two avenues for an efficiency multiple, albeit a single-digit multiple:

* Client aggregation allows a hyperscaler to average out demand spikes from uncorrelated clients, reducing the peak:average demand ratio and allowing better budgeting of compute.

* Dynamic batching allows typical requests to run in batches of more-than-1 and/or overlap, offering better internal compute utilization ratios (e.g. interleaving output and input streams). The small limit of on-device LLMs will run with batch sizes of one with strong memory bandwidth bottlenecks.

For an example of these factors in action, see the API cost differential between batch, standard, and 'fast' processing. OpenAI prices these tiers at a 1:2:4 ratio.

Majromax··on OpenRouter is joining Stripe
> As things settle down and commoditize, the value of switching on a dime diminishes as people lock into their favorite models

I can imagine just the opposite outcome from the same scenario: as people settle into their favorite but commoditized models, competition for marginal inference cost will take over. A company like OpenRouter that promises the cheapest tokens by the minute becomes essential on the low-cost margin.

I think that OpenRouter and equivalents get pushed out of the market only if the froth calms down (as you posit) and winning models stay proprietary, perhaps with their own unique API surfaces.

Majromax··on AI usage patterns in software teams
You could say the same thing about compilers versus assemblers, high-level languages versus low-level ones, and services and libraries versus monolithic programs.

All other things being equal, increasing the speed of some part of the development process will increase the overall pace of development. However, By Amdahl's law that increase will be sublinear, and that is why we should take "pull requests" as an imperfect metric.

We also don't get to pick the form that 'better technology products' take. While we'd probably like to keep cost(/effort) and complexity constant and increase robustness and performance, the market equilibrium might be 'worse is better' and reward whiz-bang features and lower effort.

Majromax··on AI usage patterns in software teams
> Is the ROI there to pay for the trillions in commitments that have been bet on that ROI? That looks like a clear no at this point.

That's only a potential crisis for those who have made concrete investments.

On the use side, the 'cost' of AI spans more than two orders of magnitude. Looking at recent models (<6mo) with reasonable performance (intelligence index >= 45) on OpenRouter, the output cost ranges from $50/MTok (Fable) to $0.153/MTok (DeepSeek Flash 0731).

From the perspective of a user of LLM/agent assistance, there's very likely a range where the benefits outweigh the costs.

If the ROI for the model developers isn't there, then that just impairs the future trajectory of the field. Current models are just bits that aren't going anywhere, and as long as they can be served (in inference) above their marginal cost they will continue to be so-delivered.

Majromax··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
A reasonable guess about the algorithm is 'A Watermark for Large Language Models' (https://arxiv.org/abs/2301.10226). The idea is that each generated token (or bigram) seeds a strong PRNG that splits the vocabulary into a 'green' and 'red' set. The sampler then tries to select a 'green' next-token for generation.

After-the-fact checking only needs the vocabulary splitter, which is independent of the LLM. Over a sufficiently large text non-watermarked text would expect to use green and red tokens with the baseline probability, and that difference can easily become statistically significant over sufficiently long texts.

The basic algorithm has obvious knobs to tune, among them the initial ratio of red to green tokens and how hard the sampler tries to pick a green token. These would balance fidelity to the original distribution against watermark detectability (minimum required content length for statistical power).

Majromax··on Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing
> But unless you know the prompt, you don't fully know which choice the model faces. Surely a lot of coding space is wasted compensating for that uncertainty.

When you only need to encode one bit, the signal to noise ratio can be very low. If I try to write my own human words under the policy of "try a little bit to avoid the letter 'e' in every fifth word," then a sufficiently long text would still be 'watermarked' even if I only succeed in this dictum (e.g.) 10% more often than the baseline.

Majromax··on Going Dark, and the era of law enforcement hacking
> Wouldn't one AI or another detect this deliberate backdoor and report it, as it'll look just like any other security vulnerability, the only difference being the intention?

That's precisely the author's point: deliberate backdoors will be more adversary-exploitable than ever before, but the demand for such from law enforcement agencies is likely to ratchet upwards.

Majromax··on Don't classify, hallucinate!
> In the notebook, I compute a MiniLM embedding of every real Wayfair classification. I compute the embedding of the fake, hypothetical embedding from the LLM. I then dot product the fake embedding into the real ones to find the most similar. Producing: [the right answer]

Isn't this begging the question that the hallucinated classification will be more selective with respect to the real schema than the query itself? What would the dot product of <E(search query), E(schema)> have given?

Even if that is too vague, smaller LLMs are capable rerankers; return the top N matching true categories and ask for a contextual ordering.

Majromax··on GLM-5.3: Frontier coding with emergent cyber capabilities
> [A]ll of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.

This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots. Frontier models should large dominate their competitors on a capability basis, but if GLM 5.2 (now 5.3) is routinely finding bugs / vulnerabilities missed by Fable and Sol then GLM might be genuinely a frontier-grade model by itself.

Majromax··on Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space
In my view, it's not so much the writing style itself as the lack of 'taste'. Text that clearly seems AI-written has a flat level of exuberance that's just exhausting, kind of like a written version of the 'loudness war'.

Without some kind of dynamic range, I find myself having to do a lot of work to infer what points are truly important versus what are at best interesting implementation details.

Majromax··on Our position on open-weights models
> The attacker just needs one exploit chain, whereas the defender needs to block every avenue. Open access to models with no guardrails greatly benefits the attackers more than the defenders.

I see it as the opposite, where the attacker needs to find an exploit chain whereas the defender can block any link.

In this model, the balance of convenience favours the defender. The defender presumably has access to the source code and configuration, so their scope of action is much larger than the attacker that must find vulnerabilities in a particular configuration.

I think that the different views might relate to different prior assumptions. If we assume that each layer is mostly secure but may have a small number of latent vulnerabilities, then it should be relatively easy to find and fix those to create a perfectly secure layer. If instead we assume that each layer is mostly insecure but chaining vulnerabilities is time-consuming then the land favours better-resourced attackers.

> Or find a remote exploit in Tesla cars and make their autopilot go on murdering rampages. (that one is from a movie)

In the worst case, air gaps and fixed contracts for information handling cover that. Like any other domain, a car can be remotely exploitable only when untrusted information can influence behaviour inside the secured region. Unfortunately, the convenience of OTA updates and 'cars as tech' rewards velocity at the expense of defensive design.

Majromax··on Our position on open-weights models
> The saying that stuck with me was "defenders have to be right 100% of the time, while attackers only have to be right once".

> You are suggesting this isn't correct?

The intuition behind that is applicable only when correctness is stochastic. If you need to be waved in by a security guard, then one fake mustache might be the difference between being granted or denied entry. However, a keypad either works or it doesn't; entering the wrong PIN is guaranteed refusal.

The other breach of that intuition is defense in depth. Secure systems don't generally rely on a single binary trusted/untrusted status; the classified building still locks its interior doors. This is the part that has – in my view temporarily – changed most with frontier models, in that they are much more skilled at chaining together vulnerabilities than previous models (and much faster about it than human experts, even if potentially less skilled). If a system has a latent (0-day) vulnerability 50% of the time, then 10 independent layers would imply a ≈ 1/1000 chance that a critical compromise is possible.

However, these independent layers don't currently happen in practice because it's easier to write insecure code than secure code. With luck, modest discipline, and defensive use of frontier models I think that this gap will narrow with time, in much the same way that it would be plainly crazy to deploy root access via telnet today.

Majromax··on Our position on open-weights models
> Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?

In the specific case of cybersecurity, this is a reasonable medium-term outcome. IMO, the cybersecurity risk is akin to the spread of a disease among an 'immune-naive' group: we can suddenly deploy much stronger attack-finding tools against large, established codebases created with much weaker security designs. The path from here to there will be rough, but it's still fundamentally easier to write secure code than it is to exploit vulnerabilities. (It's just easier yet to write insecure code, giving our status quo problem.)

For other 'safety' matters, defense isn't so easy because the attack and target are so different. An AI propaganda bot or catfisher 'attacks' slowly-evolving human culture; one that instructs on explosives or bioterrorism directly interacts with an accomplice and not a victim. If you believe that knowledge on how to build a pipe-bomb must be restricted, then giving everyone access to Fable does not mitigate the risk.

The controversial limit of this attitude is recursive self improvement and an AI singularity with potentially destructive results. Proponents of this view think that sufficiently powerful AI is risky in nearly unimaginable ways such that the capability itself is harmful. This is part (but not all) of why Fable (originally?) degraded itself when apparently assisting with AI research.

Majromax··on Our position on open-weights models
> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.

Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.

The 'universal evaluation' criterion then has three outcomes:

* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models. Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.

* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they cannot be made capable of abusive behaviours.

* It could be security theatre.

The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.

Majromax··on Benchmarking Opus 5 on SlopCodeBench
> I may be crucified for asking this but: is there any proof that slop matters beyond our sensibilities as developers?

That's precisely what this benchmark tries to quantify. Since the benchmark incrementally expands the scope of each problem, 'sloppy' code is code that is hard to later modify.

I know this firsthand: the dumbest coder I've ever worked with was 'myself six months ago'. That jackass never keeps the documentation up to date and hard-codes things that ought to be exposed as configuration.

The SlopCodeBench is an important but early-stage probe in this direction.

Majromax··on The Economics of Recursive Self-Improvement [pdf]
> because people have other ideas https://en.wikipedia.org/wiki/Technological_singularity

In a weak sense, singularities are common and should be expected every so often. In a mathematical sense, a singularity is where a model 'blows up' and neglected terms become part of the dominant balance.

Industrialization hit a point of diminishing returns, but the industrial revolution was nonetheless a 'singularity,' where life afterwards was qualitatively unpredictable to people who lived before. Likewise, agriculture was such a technological singularity to hunter-gatherer ancestors.

I could even make a decent argument that writing and literacy were such a singularity, making inconceivable social organizations routine.

In that weak sense I expect AI to be a singularity, recursive self improvement or no. Life in 2050 may be completely unpredictable to someone who was taken out of time in the year 2000.

The remaining questions are speed and intensity, and both of these questions are related to RSI. If RSI works, then the 'fast takeoff' visions become more plausible where society transforms over months to a few years – at least locally where the enabling technologies have diffused. If not, it might take a couple of decades.

Majromax··on Codex starts encrypting sub-agent prompts
No, you'd still care. YOLO mode is about instantaneous permissions and access control, inspection of subagent prompts is about retrospective quality control. If the main model is instructing subagents to do a subtly wrong thing, the overall process quality will degrade in ways that might be very hard to detect or fix without deep inspection of the middle stages.
Majromax··on A 1969 camera operators' strike created Upstairs Downstairs multiverse
If you're deliberately displaying the image in black and white, the colour pattern is interference that should be suppressed.

However, this practice was not universal, and archivists have now recreated colour copies (https://en.wikipedia.org/wiki/Colour_recovery) of some shows where only black and white recordings survived by reconstructing the colour from these interference signals.

Majromax··on Claude Fable is relentlessly proactive
> I haven't yet had an agent rm -rf files.

That happened to me once; I was running one of a few free-tier models in a pi-coding-agent session. The bash tool there is stateless and always begins from the launch directory, but the agent assumed state and executed `rm -rf .` intending to remove a build directory. Instead it removed the whole project tree, including session logs and notes.

This was mostly a matter of amusement for me since I was running the agent inside a bubblewrap sandbox for that very reason, and the project itself was not very important.

Page 1 of 27Next →