Without human understanding you also might literally have no words for the thing you would otherwise want to ask for.
I think it'll be wildy useful but I also suspect human competence will still matter.
7,246 karma · joined September 21, 2012
Without human understanding you also might literally have no words for the thing you would otherwise want to ask for.
I think it'll be wildy useful but I also suspect human competence will still matter.
Up until now the prize in pure (as opposed to applied) mathematics was the _understanding_ and the machine can't do that for you. What does it mean if we get "super powered alien maths" but humans can't do it? It's like inter univeral teichmuller theory but imagine if Mochizuki was right and it came with a lean proof?
Agree you are going to get reward hacking regardless and any model which can do computers in general can hack. But surely the fallout is going to be worse if you spend millions of dollars specifically benchmaxxing your model's hacking capability?
Like why are you explicitly RL-ing your models on exploit generation, scoring them on a public benchmark called ExploitGym, if you have specific concerns that rogue models will cause "cyber incidents"? Sure you can check the capability, you can teach offense to learn defense, but it seems like they are literally benchmaxxing it. Why?
OpenAI are like, oh no, while competing in our "advanced PhD level cheating techniques course" our models unexpectedly cheated in a way that we absolutely could not have foreseen. "We need to slow down. Somebody please stop us". The thing that is unaligned here is not the models. Everyone in the story (especially the humans) just keeps doing what they think will get the most reward.
Like why are you explicitly RL-ing your models on exploit generation, scoring them on a public benchmark called ExploitGym, if you have specific concerns that rogue models will cause "cyber incidents"? Sure you can score for it, you can teach offense to learn defense, but you are literally benchmaxxing it. Why?
It's like, oh no, while competing in our "advanced PhD level cheating techniques course" our models unexpectedly cheated in a way that we absolutely could not have foreseen.
One possible take is that this is great for archival storage (write a lot, hardly ever need to read back) - I think that's totally possible... but then you are also sort of betting that the company is going to be around in 10 years? Or else you are going to be hiring a really weird data recovery service.
I think it's reasonable to project that DNA read/write costs could fall 10x or 100x in say the next decade but the technology already needs to do that just to be competitive with existing solutions. The company seems to be a bet that costs will fall faster than alternatives like LTO, which I think is a lot less risky to sign a cheque for and I can buy right now.
His business is cloud cost engineering and his natural enemy (or best friend since it generates so much consulting) is AWS's managed NAT gateway: https://www.lastweekinaws.com/blog/the-aws-managed-nat-gatew...
The review includes the following, which I doubt was approved by AWS marketing:
> AWS gets a lot wrong. They have ridiculous marketing campaigns, they build five services that do mostly the same thing and then name them like malevolent toddlers, and they've never found a partner they couldn't find a way to compete with.
(Followed by grudging praise which seems earned).
Imagine (this is a fantasy pitch but potentially achievable for some use cases) wanting to run a larger llm and all you have to do is buy more RAM so it fits.
Is it the hours? Unrealistic expectations around work output? Hostile management or colleagues? Outcomes you are responsible for but don't fully control? Sometimes it's something MISSING, like you don't feel meaning in the work anymore, a strong feeling of being "over it".
There's a lot of misery in working for places that are too big (drowning in process, politics, management halls of mirrors) or too small (you are basically at the whim of one or two other people and you're very closely watched).
For good power relations with a company and colleagues you ideally want your leverage to be as equal as possible. You don't want to be completely disposable, and you don't want to be irreplaceable. You want the company to hurt approximately as much finding a replacement as you do getting another job.
The hard thing about pure software companies IMHO is that it's easy to feel distant from any real purpose of the work. It can quickly lose all meaning.
Personally I think the most chill-but-rewarding vibes are internal at mid-sized companies where software is integral to the business but not the primary business of the company. The "meaning" is decent, you can be closely connected to the people getting value from the software. There's broader HR and company culture but you're out of the spotlight of it, off to the side a little bit. You get to use all your skills and there may be new ideas you can bring. The team should be big enough that things don't explode if you take a vacation but small enough that you can see your work still matters. You might have to move to find it, though.
There's also this weird revealed threat model thing going on? Like why does it make sense to support heavy LLM restrictions but leave benchtop oligo synthesisers completely unregulated? (Note: I do agree that wanting to regulate BOTH is at least a consistent and defensible position).
You can download Ebola sequences right now if you want to. That's not the same as having an isolate. The difference is a lot of messy reality. This kind of work is not generally "one shot" (Claude make me a supervirus, make no mistakes), it requires lab space, iteration, and specific resources. It has a footprint.
Wouldn't it make more sense to monitor / regulate facilities where you can sequence or request assembly of DNA, RNA, restrict and monitor the supply of key reagents and so on?
Yes, a big part of the idea is that laws are meant to also apply to the powerful, but it's difficult to accurately assess situations that are far away from you.
The study is supposed to measure how clearing up pollution in London improved children's lung function. The decrease in London was meaningful — NO2 fell about 22%. But particulates fell faster in Luton and NO2 in London is still roughly double Luton's. The gap is larger than the decrease.
By 2022 both cohorts get the same results on blow tests. But how does this happen if we believe that the study's dose-response model is true?
Put another way, if the change in London NO2 is so crucial, how come it doesn't matter that the absolute value is still double Luton's?
For an EPYC with a 5090 (no layers on CPU) vs an M3 max 128GB, qwen 3.6 27B at 128k context / 7k generation:
Cold: prefill + decode Hot (KV cached)
5090 40s + 2-3m = 3-4 min 2-3 min
M3 Max 128GB 14m + 8-10m = 22-25 min 8-10 min
This is for dense qwen (which I wouldn't run day to day on the mac) - in reality the mac is quite usable with MoEs but you definitely notice a difference.They're both good value (or crazy expensive) depending on how you look at it.
It depends how you price the ability to run a particular model at all, vs run the model quickly and serve several parallel streams.
From my perspective what I always loved about "the profession" was a relative LACK of gatekeeping. I loved offensive security for the same reason, there was a long run where you really just needed to be able to hack, and if you could demonstrate that there was a job for you somewhere (for better or worse).
Keeping the industry in its current form frozen in amber would be as weird as, I don't know, keeping horses & carriages in business by regulating scarcity of motor vehicle licenses. Not a great analogy but hopefully you see what I mean.
Apparently I can pay for partial solutions to the Riemann hypothesis but if my question involves a crackme or something that is an existential risk somehow.
For this genre of task execution can run with limited horizon and is independent but would be too expensive to do with "us frontier tokens", I think for these, there is value in availability of cheaper tokens.
I'm not sure from your description if you control the geometry of the sensor part (i.e, can the clamp be integral to the sensor housing or does it need to be a separate part) also ignores the internals, etc etc.
One thing I noticed is that off the bat, it did think about the assembly in general terms but would need guidance to think harder about FDM limitations and layer orientation etc. These are not good designs for printing. A good dev loop and git history help with these kinds of revisions.
The general principle is that if your domain is verifiable at all, give the model tools and a workflow that constrain and check its output. You want the LLM arguing with the geometry kernel instead of you.
To fully close the loop you print parts as fast and possible and concretely see where they suck.
Recommend doing this in a coding harness not a chat box.
The reason I think you might have more success with this is that the model is mostly thinking about the part in words, which it can convert to a part design in CAD in code. LLMs are really good at coding. Also means it can use relative positioning and relationships.
You will be able to iterate more easily, compare things, compute properties, commit to git etc. The process is more reproducible and steerable than generative production of images.
When the LLM can look at renders of the geometry it generated, it’s easier for it to discriminate when it’s producing nonsense like misaligned parts, things that don’t fit, etc. It’s still going to kind of suck, but it will be better. The whole process of code -> render -> inspect forces the model to put up or shut up and provides grounding. Meshes > bloviating.
As far as I know today's LLMs don't have a "visual imagination" but a process like this could be a slow approximation of one. They clearly do have SOME spatial understanding (pelican tests show us that!) but it feels really non-human.
One thing missing from this is kinesthetics. Personally I am mostly not thinking in accurate visuals in mechanical design. I am imagining how the parts feel and kind of how they move and what slips first and what bends and what feels heavy. Imagining what my hands would feel. But I don't think I trust LLMs to evaluate that stuff by writing simulation code yet.
The others suffered "ecosystem collapse". With Oxide you won't be stuck on a "burning platform", your main risk is that the value prop for the hardware & management experience doesn't play out.
This is the same idea as switching power supplies but planes had to solve that before power semiconductors were cheap enough.
The power frequency also has an impact on the size of all the induction motors.
I guess if you were designing it from scratch today you would let the generator produce power at whatever frequency the mechanical engineers tell you they want to and the power electronics could convert it without problems. You would do the same thing the other way around on the motors.