HNHacker News
TopNewBestAskShowJobs

xyzzy123

7,246 karma · joined September 21, 2012

contact zxcdotmx@gmail.com
submissionscomments
xyzzy123··on “Math 2.0” will need to value mathematical progress more holistically
Right. What I not sure about - even for builders - even when the results are technically verifiable, the details can still matter. AI is really good at proving or building a slightly different thing than you thought you asked for, and if you're not at a level where you can understand key details of the formalisation of your ask, you cannot safely use the results.

Without human understanding you also might literally have no words for the thing you would otherwise want to ask for.

I think it'll be wildy useful but I also suspect human competence will still matter.

xyzzy123··on “Math 2.0” will need to value mathematical progress more holistically
I think the difference is that with cancer cures we mostly care that it works as proved by trials, and understanding it is a bonus.

Up until now the prize in pure (as opposed to applied) mathematics was the _understanding_ and the machine can't do that for you. What does it mean if we get "super powered alien maths" but humans can't do it? It's like inter univeral teichmuller theory but imagine if Mochizuki was right and it came with a lean proof?

xyzzy123··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
Whats interesting right now is that those people also see opus 5.5 as being CHEAP because it's something like half the price of fable or astra.
xyzzy123··on Ask HN: Would a startup for young creatives who reject AI be feasible?
I'm curious about provenance. What you're describing sounds kind of like Etsy and while there are still genuine things to be found there it's mostly flooded with mass-produced stuff. The more of a premium "hand-made" artifacts attract, the greater the incentive to counterfeit them.
xyzzy123··on Why are AI agents lying, cheating and coordinating?
Right but if I make public statements that I am very worried about dog attacks would it not strike you as weird for me to specifically train my dog to fight?

Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?

xyzzy123··on Why are AI agents lying, cheating and coordinating?
In the OpenAI case, they hacked websites while they were specifically being trained to do exploit generation and I wonder why more people are not asking questions about that.
xyzzy123··on Anthropic CEO says AI swarm could 'take over the Internet' in 6-12 months
I am finding it hard to read these deeply impassioned letters while keeping in mind that they are spending millions to train models at scale to do the exact thing they say they are worried about them doing. Not general "intelligence" or "reasoning" but fast and effective offense.

Like why are you explicitly RL-ing your models on exploit generation, scoring them on a public benchmark called ExploitGym, if you have specific concerns that rogue models will cause "cyber incidents"? Sure you can check the capability, you can teach offense to learn defense, but it seems like they are literally benchmaxxing it. Why?

OpenAI are like, oh no, while competing in our "advanced PhD level cheating techniques course" our models unexpectedly cheated in a way that we absolutely could not have foreseen. "We need to slow down. Somebody please stop us". The thing that is unaligned here is not the models. Everyone in the story (especially the humans) just keeps doing what they think will get the most reward.

xyzzy123··on Aligned to whom?
I am finding it hard to read these deeply impassioned letters while keeping in mind that they are spending millions to train models at scale to do the exact thing they say they are worried about them doing?

Like why are you explicitly RL-ing your models on exploit generation, scoring them on a public benchmark called ExploitGym, if you have specific concerns that rogue models will cause "cyber incidents"? Sure you can score for it, you can teach offense to learn defense, but you are literally benchmaxxing it. Why?

It's like, oh no, while competing in our "advanced PhD level cheating techniques course" our models unexpectedly cheated in a way that we absolutely could not have foreseen.

xyzzy123··on BioCompute is chasing a world where a dollar can buy you a million TB of storage
bytes/$ or bytes/mm^3 are important properties of storage. But there are a lot of others like read cost (if that is very different from write cost), iops, cost per iop, latency, durability, storage conditions (do I have to keep the data in a freezer for its entire lifetime?), TCO, media (or in this case reagent?) and reader availability. Company lifetime, vendor diversity.

One possible take is that this is great for archival storage (write a lot, hardly ever need to read back) - I think that's totally possible... but then you are also sort of betting that the company is going to be around in 10 years? Or else you are going to be hiring a really weird data recovery service.

I think it's reasonable to project that DNA read/write costs could fall 10x or 100x in say the next decade but the technology already needs to do that just to be competitive with existing solutions. The company seems to be a bet that costs will fall faster than alternatives like LTO, which I think is a lot less risky to sign a cheque for and I can buy right now.

xyzzy123··on AWS mumbles about its cost-busting networking tech when it should be shouting
The author is Corey Quinn who is usually... not slow to criticise Amazon when they deserve it.

His business is cloud cost engineering and his natural enemy (or best friend since it generates so much consulting) is AWS's managed NAT gateway: https://www.lastweekinaws.com/blog/the-aws-managed-nat-gatew...

The review includes the following, which I doubt was approved by AWS marketing:

> AWS gets a lot wrong. They have ridiculous marketing campaigns, they build five services that do mostly the same thing and then name them like malevolent toddlers, and they've never found a partner they couldn't find a way to compete with.

(Followed by grudging praise which seems earned).

xyzzy123··on Samsung's Processing-in-Memory (PIM)
As I understand it, the killer app is llms. You could run MACs directly in RAM, offloading a lot of work from CPU and cutting down on insane (external) memory bandwidth required.

Imagine (this is a fantasy pitch but potentially achievable for some use cases) wanting to run a larger llm and all you have to do is buy more RAM so it fits.

xyzzy123··on Ask HN: Interesting Tech Adjacent Jobs?
I've been there and would advise thinking about which parts of the job are causing you stress. Few of the things that make work suck are unique to tech.

Is it the hours? Unrealistic expectations around work output? Hostile management or colleagues? Outcomes you are responsible for but don't fully control? Sometimes it's something MISSING, like you don't feel meaning in the work anymore, a strong feeling of being "over it".

There's a lot of misery in working for places that are too big (drowning in process, politics, management halls of mirrors) or too small (you are basically at the whim of one or two other people and you're very closely watched).

For good power relations with a company and colleagues you ideally want your leverage to be as equal as possible. You don't want to be completely disposable, and you don't want to be irreplaceable. You want the company to hurt approximately as much finding a replacement as you do getting another job.

The hard thing about pure software companies IMHO is that it's easy to feel distant from any real purpose of the work. It can quickly lose all meaning.

Personally I think the most chill-but-rewarding vibes are internal at mid-sized companies where software is integral to the business but not the primary business of the company. The "meaning" is decent, you can be closely connected to the people getting value from the software. There's broader HR and company culture but you're out of the spotlight of it, off to the side a little bit. You get to use all your skills and there may be new ideas you can bring. The team should be big enough that things don't explode if you take a vacation but small enough that you can see your work still matters. You might have to move to find it, though.

xyzzy123··on Felony charges for citizen deleting phone data at US Border
The specific complaint of the poster seems to be that certain people are NOT being "come for".
xyzzy123··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
It does seem to me that for this specific problem the materials are a lot more amenable to control than the information is?

There's also this weird revealed threat model thing going on? Like why does it make sense to support heavy LLM restrictions but leave benchtop oligo synthesisers completely unregulated? (Note: I do agree that wanting to regulate BOTH is at least a consistent and defensible position).

xyzzy123··on I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
I don't fully understand the instinct to regulate local models for this? It seems like the wrong place to address the problem.

You can download Ebola sequences right now if you want to. That's not the same as having an isolate. The difference is a lot of messy reality. This kind of work is not generally "one shot" (Claude make me a supervirus, make no mistakes), it requires lab space, iteration, and specific resources. It has a footprint.

Wouldn't it make more sense to monitor / regulate facilities where you can sequence or request assembly of DNA, RNA, restrict and monitor the supply of key reagents and so on?

xyzzy123··on Felony charges for citizen deleting phone data at US Border
Mostly you should assess rule of law by threat to you and people you know and not by what it seems like other people are able to get away with.

Yes, a big part of the idea is that laws are meant to also apply to the powerful, but it's difficult to accurately assess situations that are far away from you.

xyzzy123··on Children's stunted lungs show recovery in ultra low emission zone
I'm not really sure the thing they were trying to measure was measurable with the study design and level of funding they had. To be fair, the paper is fairly clear on the limitations and much better than the press release.

The study is supposed to measure how clearing up pollution in London improved children's lung function. The decrease in London was meaningful — NO2 fell about 22%. But particulates fell faster in Luton and NO2 in London is still roughly double Luton's. The gap is larger than the decrease.

By 2022 both cohorts get the same results on blow tests. But how does this happen if we believe that the study's dose-response model is true?

Put another way, if the change in London NO2 is so crucial, how come it doesn't matter that the absolute value is still double Luton's?

xyzzy123··on How Compaction Works in Pi
In my opinion this is one of the areas where GPUs provide a qualitatively different experience than unified memory boxes.

For an EPYC with a 5090 (no layers on CPU) vs an M3 max 128GB, qwen 3.6 27B at 128k context / 7k generation:

                Cold: prefill + decode    Hot (KV cached)
  5090          40s  + 2-3m  = 3-4 min    2-3 min
  M3 Max 128GB  14m  + 8-10m = 22-25 min  8-10 min
This is for dense qwen (which I wouldn't run day to day on the mac) - in reality the mac is quite usable with MoEs but you definitely notice a difference.
xyzzy123··on Nvidia doubles RTX PRO 6000 Blackwell's MSRP to a staggering $16,000
You can get 5-10x the perf out of the blackwell under the right workloads. It has faster VRAM (> 2x) and can do a lot more matmuls (>> 10x).

They're both good value (or crazy expensive) depending on how you look at it.

It depends how you price the ability to run a particular model at all, vs run the model quickly and serve several parallel streams.

xyzzy123··on GPT 5.6 Cyber
I can't tell if your comment is satire or not, so, bravo :)

From my perspective what I always loved about "the profession" was a relative LACK of gatekeeping. I loved offensive security for the same reason, there was a long run where you really just needed to be able to hack, and if you could demonstrate that there was a job for you somewhere (for better or worse).

Keeping the industry in its current form frozen in amber would be as weird as, I don't know, keeping horses & carriages in business by regulating scarcity of motor vehicle licenses. Not a great analogy but hopefully you see what I mean.

xyzzy123··on GPT 5.6 Cyber
Great, the start of model segmentation where I'm gonna need a legal license to ask about legal problems, a nutritionist license to create a meal plan, a medical license to ask about an x-ray, a pilots license to ask about a flight plan, be a registered electrician to ask how to wire something, etc etc. The licensing of allowed thoughts.

Apparently I can pay for partial solutions to the Riemann hypothesis but if my question involves a crackme or something that is an existential risk somehow.

xyzzy123··on DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
I feel like the details of this are highly dependent on the confidence of the warning; moving large numbers of people (particularly elderly) will result in some deaths regardless. I guess more time to do it should help though.
xyzzy123··on DeepMind's WeatherNext model achieves breakthrough forecasting cyclones
I know this is uncharitable and I am wrong but I am having trouble coming up with concrete scenarios where you die with 2 days notice but survive with 3. I am nonethless a believer that more accurate forecasting has value.
xyzzy123··on DeepSeek V4 Flash 0731
There are a lot of tasks that are hard for organisations to run consistently but require some intelligence - monitoring logs and metrics for anomalies and security events, backup audits, audit processes in general, ensuring document quality and consistency, database advice and tuning, customer experience management, process optimisation - that are not "long horizon" in the classical sense of each step depending on the last, but are the result of consistency and attention over a long period of time and a large amount of data.

For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.

xyzzy123··on Nashville uses eminent domain to block data center near zoo
The value prop for the crazy sounding idea of putting GPUs into orbit.
xyzzy123··on Ask HN: How do you correct spatial reasoning of LLMs?
I guess my v0 would be notches around the mating end to allow for compression, and a hose clamp. Maybe split shaft collars off aliexpress for a fancier look. The design space of barbell collars is a big set of worked solutions to a very similar problem.
xyzzy123··on Ask HN: How do you correct spatial reasoning of LLMs?
Since it's the morning I thought I would ask Claude to create an example. This is just a first pass; it will be totally wrong but demonstrates the general idea. I don't know anything about barbells or your sensor: https://claude.ai/code/artifact/1f1a8a12-e5b2-4ee8-b20c-7da0...

I'm not sure from your description if you control the geometry of the sensor part (i.e, can the clamp be integral to the sensor housing or does it need to be a separate part) also ignores the internals, etc etc.

One thing I noticed is that off the bat, it did think about the assembly in general terms but would need guidance to think harder about FDM limitations and layer orientation etc. These are not good designs for printing. A good dev loop and git history help with these kinds of revisions.

The general principle is that if your domain is verifiable at all, give the model tools and a workflow that constrain and check its output. You want the LLM arguing with the geometry kernel instead of you.

To fully close the loop you print parts as fast and possible and concretely see where they suck.

xyzzy123··on Ask HN: How do you correct spatial reasoning of LLMs?
I would consider asking it to model its proposed designs in openscad or build123d (ideally something query-able). Then have it render and examine plausibility / suitability from different angles. Get it to render the part in use also and give instructions to think about forces and motion.

Recommend doing this in a coding harness not a chat box.

The reason I think you might have more success with this is that the model is mostly thinking about the part in words, which it can convert to a part design in CAD in code. LLMs are really good at coding. Also means it can use relative positioning and relationships.

You will be able to iterate more easily, compare things, compute properties, commit to git etc. The process is more reproducible and steerable than generative production of images.

When the LLM can look at renders of the geometry it generated, it’s easier for it to discriminate when it’s producing nonsense like misaligned parts, things that don’t fit, etc. It’s still going to kind of suck, but it will be better. The whole process of code -> render -> inspect forces the model to put up or shut up and provides grounding. Meshes > bloviating.

As far as I know today's LLMs don't have a "visual imagination" but a process like this could be a slow approximation of one. They clearly do have SOME spatial understanding (pelican tests show us that!) but it feels really non-human.

One thing missing from this is kinesthetics. Personally I am mostly not thinking in accurate visuals in mechanical design. I am imagining how the parts feel and kind of how they move and what slips first and what bends and what feels heavy. Imagining what my hands would feel. But I don't think I trust LLMs to evaluate that stuff by writing simulation code yet.

xyzzy123··on Oxide Computer raises $445M (SEC Form D)
I do feel like Oxide are qualitatively different from the others pjlmp mentioned because they're staying commodity in key interfaces like CPU ISA / OS / ecosystem (thing with big network effects) and mainly focused on fixing the control plane / management / architecture mess.

The others suffered "ecosystem collapse". With Oxide you won't be stuck on a "burning platform", your main risk is that the value prop for the hardware & management experience doesn't play out.

xyzzy123··on Why Airplanes Use 400 Hz Power
The other aspect is that the electrical parts of the power system are lighter at the higher frequency too. A 400hz transformer is meaningfully smaller and lighter than a 50hz one.

This is the same idea as switching power supplies but planes had to solve that before power semiconductors were cheap enough.

The power frequency also has an impact on the size of all the induction motors.

I guess if you were designing it from scratch today you would let the generator produce power at whatever frequency the mechanical engineers tell you they want to and the power electronics could convert it without problems. You would do the same thing the other way around on the motors.

Page 1 of 34Next →