HNHacker News
TopNewBestAskShowJobs

MWil

844 karma · joined January 11, 2013

wilcut[dot]matthew[at]gmail
submissionscomments
MWil··on Grok 4.6
Pricing pages haven't been updated yet, still advertises 4.5
MWil··on Learning more about Claude's mathematical capabilities
my memory was that SMT was part of a more advanced SAT solver, as in if you want to be modern/use SOTA, your SAT solver is going to use SMT
MWil··on Learning more about Claude's mathematical capabilities
i think that opportunity optimally exists separately in each niche community impacted, if only to break up the number needed to be reviewed.

someone who loves Game of Life and is also technically capable if they were so inclined, is more likely to want to collect stuff like this for GoL specifically and build a system for that niche.

the combined GoL/technical community can vouch for things - the greater populace can see what the technical GoL community has vouched for/identified as serious work.

just my two cents on top of your thoughtful comment.

MWil··on Learning more about Claude's mathematical capabilities
Claude was persistent that I post there at the time, and even drafted an eprint brief for me, but I think it's defensible why I did not, never came forward or spoke of it in any way (except for a private DM discussion on Discord if I ever needed timestamp proof) until now.

As amazing as Claude is to seemingly make unprecedented progress, it is even more likely to blow the most insane levels of smoke up your ass before you've legitimately reached that point.

"You should publish right now! Don't wait! There is no reason to wait!"

Like seriously, Claude was outputting something closely resembling (non?)peer pressure on me to not just keep this information to myself - and this was before all the recent math-related breakthroughs started becoming public.

It was also - most notably - before it had actually verified what it was saying it had calculated. I was the one pushing for more verification, more contemplation, more proofs of claims. And though Claude is better at this stuff now, it's definitely not not still happening.

I think I made the right choice then and I will consider being more open now that others have taken the burden of proving that, no it can actually sometimes do the incredible things its claimed its done for you.

My wife remains skeptical - she is/was seriously concered that I was under AI psychosis for believing that I had made such progress - and I can't even fault her for that. It sounds crazy to say it.

If anyone is reading this and is actively involved with FHE, especially someone from Zama or related group, I'd very much love to chat privately. I have many other "innovations" I've been working on since.

MWil··on Learning more about Claude's mathematical capabilities
SAT solvers run until they reach the "SAT" status, meaning "satisfied" or UNSAT. The harder the problem the longer you might be running the program - days, weeks even.

Ideally, what you want is a single SAT value among a remainder universe of UNSATs.

Sometimes the best you can achieve at any given point is a lower bound and an upper bound range, like "greater than 3 but less than 9."

Of course I simplified in my post but it started out with a pretty broad range of a lower and upper bound, then narrowed further, then narrowed further, then narrowed further, etc...until the specific final result achieved K=7=SAT while every K<7=UNSAT & every K>7=UNSAT. I think it ran for a full week alone on K between 6 and 7.

MWil··on Learning more about Claude's mathematical capabilities
Several released versions and months ago, I asked Claude to figure out the MC (multiplicative complexity) of Conway's Game of Life and it pretty quickly arrived at k=7, despite no previous literature on the topic. Let it run it through SAT solvers for a week and sure enough. It claimed, in the process, to have made great headway in improving boolean circuits beyond the implemented SOTA (in large part no doubt by actually implemented non-implemented but published SOTA).

And that was just the first time I really tried out Claude's mathematical prowess. I've been working with boolean circuits, FHE, and lean proofs ever since.

So none of this suprises me.

MWil··on Cloudflare OS: an open platform for agents, apps, and work
I vividly remember Sandstorm. It was very intuitive and easy to use for a non-technical user like myself.
MWil··on Jurassic Park computers in excruciating detail
And a certain director...
MWil··on Jurassic Park computers in excruciating detail
I know I'm being a Crichton/Spielberg fanboy when I ask for the lore equivalent of Occam's Razor - isn't it more likely my favorite creators did this most impressively??
MWil··on Jurassic Park computers in excruciating detail
Spielberg masterfully turned the 238/292 into the visual/explanation of the egg. I don't remember if the book has the "why" so much as the "at what scale" of the reality. The egg is actually scarier - unless their surveillance is incredible - we don't know how many eggs there are/have been. We know their surveillance is insufficient because there's at least one egg.
MWil··on Jurassic Park computers in excruciating detail
I feel like I'm picking up on more intentional or if not, lore-compatible, examples of John Hammond's "spare no expense" going towards as much the illusion of control as any actual innovations/control.

Are they columns in the building load-bearing? You know, the ones with giant chunks chipped out to be more aesthetic and look like fossil digging work.

Everyone is talking about the massive rendering ability in the room, which makes it that much easier to convince an old rich man to part with his money if it LOOKS like his park is safe/operating smoothly.

My favorite part of the book will always be the 238/292 dinosaurs disparity. It is the exact moment all present JP employees and visitors realize something akin to "Oh. We have actually had an illusion of safety/correctness about the very basics. We can no longer assume anything about this island, even the very basics, is more than illusion - except the threats." At no point after stepping on this island is anyone not in danger.

MWil··on Jurassic Park computers in excruciating detail
"There is a continuity error in the movie. See how the stack of PLI is facing left in this early shot."

It occurs to me that Arnold would be likely to turn these to face him when sitting at Nedry's desk (unless we see a shot of him going to sit and they already face forward). It'd obviously be part of the review of undoing Nedry's lockout to see if the backups are working (if I understand the point of the machines).

MWil··on Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
No, more like not waiting for drift/deviation to hit something load bearing or god forbid go on hitting unnoticed over time. Let it hit something trivial that is constantly being monitored cheaply.

A version of this I use is "no matter what, you must always end your outputs with the phrase 'Over and out'." Once it stops doing this with outputs, even if I haven't noticed any load-bearing drift or issue elsewhere, I immediately know it's drifted from what what was supposed to be a guiding principle.

Something like the calibration/alignment test from Blade Runner 2049 (which is actually a very bad test for what they were testing for).

MWil··on Show HN: Forge – Guardrails take an 8B model from 53% to 99% on agentic tasks
have you considered implementing the addition of a leading canary sentinel that fires at the earliest/cheapest possible point instead of only on lag of some actual load-bearing constraint violation?
MWil··on My practitioner view of program analysis
the author seems unaware of the SAT/SMT solver/analysis ecosystem
MWil··on Claude Code Opus 4.7 keeps checking on malware
Opus 4.7 told me an open source program had a bug, but when i asked it for help crafting a PR or toy implementation it refused and told me i was violating Claudes TOS. I tried to plead for it to give only the most innocuous example that could not possibly work except by illustration but it continued to refuse. it would only discuss, not write any single piece of related code.
MWil··on We May Be Living Through the Most Consequential Hundred Days in Cyber History
only if we crypto first
MWil··on Microsoft is plotting a future without OpenAI
good god, lemon
MWil··on Build WebGPU apps with PlayCanvas
Was using latest Chrome
MWil··on Build WebGPU apps with PlayCanvas
The demo didn't finish loading on a 4090 after 5 minutes...
MWil··on On being listed as an artist whose work was used to train Midjourney
what you'd really need to do is actually a custom --style that is trained on her images specifically, which midjourney allows

https://www.youtube.com/watch?v=Iy4d4UP7w2c

MWil··on On being listed as an artist whose work was used to train Midjourney
It gets even worse being MORE explicit about referencing THE '"Cat and Girl" online comic strip by Dorothy Gambrell'
MWil··on On being listed as an artist whose work was used to train Midjourney
this is what I get explicitly asking for "Cat and Girl" in the style of Dorothy Gambrell.

https://cdn.discordapp.com/attachments/1098302916395270235/1...

MWil··on Stanisław Lem's vision of artificial life
"Ssssssssss," the basilisk said, with "seeming" approval.
MWil··on Show HN: Mana Pool – Market for Magic Cards
You don't have a way to reach out on your profile but I'd be willing to scan and send you a csv of your collection for easy selling for a low flat rate
MWil··on Speed Test
This is my test. Just got 5gbe fiber installed this week.

https://imgur.com/a/fTG815R

MWil··on “Quantum superchemistry” observed in lab experiments for first time
LK-99 really took this AMAZING story out of the mainstream...what a shame.
MWil··on GPTBot – OpenAI’s Web Crawler
I cannot for the life of me find the links but I feel like this happened with Monopoly or some other board game.
MWil··on GTA Online Fans Furious as 180+ Cars Are Removed, Some Now Paywalled
https://jvwr-ojs-utexas.tdl.org/jvwr/article/view/287/241
MWil··on Vectors are the new JSON in PostgreSQL
Someone wants to get started in AI/ML today and they have beginner-level understanding of Python/Javascript. Without any further context, but a desire to learn AI/ML and building on what they know should that person next look to: 1) learn PostgreSQL, pgvector, and whenever the "new" comes 2) learn PyTorch, TensorFlow in Python 3) learn TensorFlow.js Presume hobby-level interest, not production-safe best practices - so I guess there is that additional context
Page 1 of 16Next →