HNHacker News
TopNewBestAskShowJobs

Frost1x

6,166 karma · joined February 8, 2019

submissionscomments
Frost1x··on A beginning for mathematics
I can see how this trend is going to go.

Manager: “Why isn’t feature X available?”

Person B: “key pieces are delayed due to the developer not understanding all of the LLM doesn’t and implementation.”

Manager: “does it work? What are the risks?”

Person B: “well yes it works for now but we’re accumulating tech debt due to a lack of understanding and potential flaws that haven’t been thought out yet”

Manager: “they want feature X, ship it, we can deal with it later, I don’t care if it’s not coherent as long as it works.”

How many decades at this point has these been a push for functionality over everything at all costs? And you have a mechanical snow plow now. Most businesses don’t care about later risk or any future planning beyond the quarter horizon, they’re not concerned about how it will effect their performance in 3 quarters or lead to instability or issues, those are future problems for a future person and we’re here for money now.

Frost1x··on The Waymo effect: how AI is quietly making research less collaborative
I work at an intersection of tech, applied research, and science.

Something I’ve noticed in collaboration that does occur is an increased confidence in people outside their domains to say things with conviction. I have people who have limited experience with software pushing out layers and layers of abstracted code that’s fairly sophisticated but often misguided in intent who will say what they’re doing is correct, with conviction.

I also hear a lot more questioning people in their domains and challenging opinions, then hearing what I can only imagine are fragmented pieces of conversations they had with an LLM thinking through some argument. Then there’s silence when you discuss shortcomings, then they come back later with their memorized fragments of what you said, combined with memorized fragments of the LLM response to the argument.

It’s occurring, a lot more. People are treating their LLMs in collaboration as a source of truth and using then to focus on their specific path or goals they think or have bias towards going down, vs just opening discussing things, considering tradeoffs from experts multiple disciplines weigh in on and then taking an approach that everyone finds most agreeable.

It’s making me want to be a lot less collaborative with such individuals. I don’t want to sit around and refute Claude text outputs all day.

Frost1x··on Apple Watch Series 12
Don’t forget Japan! There’s some great Japanese mechanical watches. I’m a fan of the Seiko spring drive movement.
Frost1x··on iPhone 18 Pro and iPhone 18 Pro Max
How will this work against rehosted images on all the platforms people actually share photos from? It’s a good feature but many popular services apply metadata stripping and basic image adjustment before resharing. That seems like the place where people would actually want to verify authenticity.
Frost1x··on Political meddling at the Census Bureau damages the US statistical system
> How can you fuck up so monumentally, but then when Trump fires you somehow he’s the problem?

They can both be a problem, it doesn’t have to be mutually exclusive. I’m sure in all the turnover Trump removed some incompetent people. He also probably removed quite a few competent people. The ratio is what matters.

Frost1x··on OpenAI's GPT-6 Astra on ARC-AGI-3
It’s not that different than a lot of real world economies. Often paying for someone or something with better quality can reduce total costs. You have less failures, less mistakes, so on, so while the expertise or quality of the product is higher than cheaper solutions, they can be more reliable and over time ultimately cheaper.

The question I have is how far back that curve can go without relying on economies of scale to just drag all the points back to the left. And without overfitting a specific metric that I don’t need (like this test).

Frost1x··on OpenAI's GPT-6 Astra on ARC-AGI-3
So, you’re telling me I need to start a benchmark as a side gig to get a bunch of free compute.

Astra please create a benchmark that’s favorable to your reasoning skills with a human interface but don’t make the score too attainable add some small issues that keep you below 100% to look sensible and to keep my evaluation metric side gig going.

Alignment++

Frost1x··on Warp builds self-improving agents on Claude
That’s sort of, in my opinion, the power of agents that can assist in developing software. The parts that are deterministic are best baked into existing programming paradigms. In some cases it’s good to take the nondeterministic parts we tried to bake into programming languages (often using generic probabilistic means) to outsourcing back to agents. Sometimes even then if the nondeterministic part is well understood and probabilistic methods work (lots of modeling lands here) then leave that in programming paradigms as well.
Frost1x··on GLM-5.3 is now open-weight
> Part of it is knowing that whatever sort of enshittification the cloud providers do, my local programming environment won’t ever be less effective than it is today locally.

Is that true though? Many of the core LLMs need to be retrained as languages evolve to incorporate changes (language specifics, compilers, tooling, etc.). To some degree this can be handled via context injection in a variety do forms (agents looking up documentation and so on) but inevitably it’s not stationary in time, just as your OSS stack (probably) isn’t (depending on the languages, technologies, and use cases).

So your hardware is to some degree dependent on the good merit of groups like Z or Alibaba or whomever pushing out updated open weight models that dumped loads of capital into to train. You can keep using the existing models but at some point I suspect they’ll start to have more friction due to dated specs in language and so on. Again there are tuning and ways of layering this information on, and in theory you can even do some training on your own but I don’t think it’s as stationary as being portrayed here.

Those updated open weight models may not always be there (updated on new data). The usability of them is probably fairly long to be fair, but I suspect you’re going to see explosion in everything from libraries to languages etc due to LLMs so even the rate of change across your OSS stack may cause these models to be dated quite quickly, at least in the core model which will require layering fixes.

To be clear I’m on the fence thinking about much of the same issues and as close as I am to pulling the trigger, I keep thinking of very valid counter arguments as to why it’s me just wanting this thing I own. Which may be enough.

Frost1x··on 'AI refuser' quit her dream job, and hopes others follow
The underlying issue is really capital ownership, what classifies as capital, protections around it, and how much leverage capital provides on a society and democracy (that part being the most important, IMHO).

Who owns the looms and what they do with them isn’t inherently an issue if all loom owners can do is buy an extra Yacht. Instead they can enforce ungoverned law on society through a combination of disproportionate influence in government and through market forces where private policy (especially at large) become nearly undifferentiated from law (the policy that benefits them becomes so widespread and normalized that alternatives are for all intent and purposes impractical or unreasonable, therefor private policy within a capital ownership domain is law or they’re a monopoly so their policy is the policy).

But I think you’re right, as always we’re going to focus on the adjacent issues vs addressing the root of the problem. The issue is what wealth inequality affords one, when it is capable of infringing on rights and livelihoods of others, not that they necessarily have to share the other luxuries and rewards of their attained wealth. I care not how many luxuries in life Musk has, I may be envious from time to time but whatever. I care a lot more when things he does or says has unrealistic influence and affects me directly, just because he sits atop a mountain of capital and we pretend that mountain of capital somehow was bestowed upon him from divinity that he should have such influence. I’m picking on Musk because he’s the richest and has clear examples of this, he’s by no means alone… it’s that class of wealth at large.

Frost1x··on Feature Request: Support AGENTS.md
I was thinking ln -s AGENTS.md CLAUDE.md but to each their own.
Frost1x··on AI usage patterns in software teams
You don’t even have to go that far, recent Opus models are quite impressive.
Frost1x··on Working with AI Feels More Like Leadership Than Coding
It’s a management of management position I’d say. You’re not just managing AI, if you were already in a management or lead type role then now you’re managing people, the LLMs they manage, and any LLM you’re managing.
Frost1x··on Anthropic Risk August 2026 [pdf]
To play the devil’s advocate, some people and orgs that were highly inefficient that adopted it really could get massive productivity gains.

A performance improvement is relative to some baseline, and that baseline for some may be a lot lower than others, and if they adopt tech effectively it really could be a big boost. Across the board though I don’t think it’s sensible.

I work in an industry that very intentionally tries to be inefficient and I can tell you having certain tasks automated that before had a person barrier intentionally acting inefficiently that you can now sidestep by outsourcing their tasks to something like Claude gives me a massive performance increase because I’m not blocked as much anymore. I can literally just replace some external tasks that were intentionally slowing processes down for their own benefits with a few prompts and move along. I could have done the tasks before but then people would ask why I’m spending my time doing it, now I can just say “oh, I was blocked so I had Claude take care of that blocker” and move along.

Frost1x··on Qwen 3.8 27B
“Thinking” is just a guiding methodology to help iterations (between the initial prompt, results, and a mixture of harness back and forth to the LLM) converge on something sane in a massive parameter space.

I like to think of it much like (as a common example most people can relate to) the Newton-Rhapson method for finding roots of a (mathematic) function. Your initial prompt runs, then the ‘harness’ kicks in using whatever methodologies are behind them to iterate on that prompt (back and forth with the model, occasionally with the user to get better guidance) and refine the outputs to hopefully converge back to some sensible output or actions the user was initially looking for.

So you’re hoping for an LLM that sort of ‘zero shots’ or needs minimal iterations from a prompt to give usable results. I find from my anecdata it varies across models and what I’m trying to get it to converge on. I tend to prefer models to not zero shot attempt because they tend to not do great, I want them to get feedback often to let me push them down the route of convergence in spaces I already understand well, meanwhile I like them to explore and give me new paths in spaces I’m not too familiar with.

That’s really what all that “second guessing” is, it’s making sure you’re following a sane path in a massive parameter space of an ambiguously defined problem. Imagine if in Newton’s method you checked the slope and it didn’t decrease from the last iteration and you just say “screw it let’s keep trying that direction.” LLMs and their harnesses tend not to have that base assumption like iteration on decreasing slopes to guide them closer to convergence, it’s a lot messier.

Frost1x··on "Code was never the hard part" is an insult to all programmers
To some degree the defense just popped up because it’s part of the AI effect, at least in my opinion: https://en.wikipedia.org/wiki/AI_effect

As computing systems become increasingly capable and encroach in our territory that distinguishes us and lead to our success as a species, intelligence (whatever that is or isn’t), we redefine the problem and handwave away the new capabilities.

It’s getting increasingly more difficult to do that in knowledge domains with current frontier agentic systems. They’re not AGI, but they start to make it increasingly difficult to move the goal posts for many people’s comfort.

We really need a lot more philosophers, sociologists, and frankly economists working on this problem: in an era where physical needs were mechanized away and increasingly aspects of the knowledge economy are shifting away, what does it look like in modernity? How do we sustain or adapt our current economic models? What new models may be needed? Do we need to continue to enforce this whole work to survive in an environment where much work is disappearing or at the very least shifting around.

No, we’re not there yet. You still need experts to guide things around, but it’s becoming increasingly easier to do more in this space with less humans. That’s not a trivial change in the US where we put most our eggs in this whole knowledge economy basket.

Frost1x··on OpenJDK Interim Policy on Generative AI
There’s another issue where models and transparent wrappers around models that get exposed are shifting around often. Versioning is highly questionable, and not all closed models will be supported indefinitely… so determinism becomes highly questionable at a purely prompt level.
Frost1x··on US citizen charged after GrapheneOS phone wipes during airport search
Well I don’t have any legal need to hire a lawyer or anything I would need a lawyer for. It’s a rather fast way to surface legal information and precedent. I don’t see how it’s any more depressing than Google diving on a topic you’re interested in for an hour..
Frost1x··on US citizen charged after GrapheneOS phone wipes during airport search
I’ve been arguing against some LLMs about this point for a good hour and there’s a whole lot of linking intent to action where you can be liable if a court can prove it. Not that an LLM is legal gold but it’s the best thing I have to pass ideas around with.

The entire situation is sort of nonsensical and boils down to lots of minutia in law that no normal person would know about.

For example having normal widely known security features like wiping the device after N failed PIN attempts is fine. Even having long standing security practices that can’t be related are fine, like having a timed touch point where if you don’t enter the PIN every… 15 days or whatever the device wipes, perfectly fine if it can’t be connected towards the crime and you’re not compelled to tell officers you have such a security mechanism.

Even if you were to set a trap where you use the same PIN for your bank, your laptop, and some other security devices in repetition then decide to set your duress PIN to that by assuming it would be discovered as a probable option they’d use, you’d be ok but it could be questionable if that was by design…

It’s so obscure really as to how and how you’re not allowed to protect your data, even if you’re not the one performing the action to clear destroy the potential evidence yourself. The entire thing seems pretty absurd a frankly arbitrary to me, and I don’t know how people could know which cases are and aren’t legal. I know not to destroy evidence myself but I wouldn’t know to tell someone to not use the duress pin or that even giving them my duress pin could somehow be my liability. It’s madness if you ask me.

Frost1x··on I wanted a clock that never needed setting. Things escalated
I’ve had a Phillips “atomic” alarm clock beside my bed for 23 years that just sets based on radio signal and auto restores if the power is lost, to which it has a backup 9v battery.

There’s a time zone setting offset slider and a DST slider. I basically touch it twice a year, maybe another if I move time zones. I’ve only had to touch it for DST (never switched time zones). Takes me conservatively 10 seconds to find and flip, so it’s taken me 460 seconds or a little under 8 minutes in the past 23 years to do time adjustments.

While these efforts are definitely fun hobby projects, there are cheap reliable solutions out there with minimal intervention that consume the NIST radio signal for time.

Just for anyone interested who wasn’t aware there’s some “old school” time broadcast solutions out there too besides NTP: https://www.nist.gov/pml/time-and-frequency-division/time-di...

Frost1x··on Nvidia, Microsoft, Meta warn against overregulating open-weight models
To some degree even unintentionally this is all inevitable. I know this isn’t what you’re taking about, but it’s a similar line of thought. These models are consuming public free information, and they’re all also producing free public information, so they’re all pissing and drinking into same pool.

If we go down the line of dead internet theory which I’m becoming more convinced of these days, the volume of information that’s not necessarily original or extracted from reality and just interpolated and extrapolated from existing information in different ways by LLMs will greatly outnumber human information coming up.

In which case these models should.. start to converge on the same data I imagine, with slightly different behaviors within that. One big generative orgy feedback loop.

Frost1x··on Terence Tao's ChatGPT conversation about the Jacobian Conjecture counterexample
I had a similar reaction with the language but also response times and language. It reminded me of a quote about Jon von Neumann from Edward Teller:

"Von Neumann would carry on a conversation with my 3-year-old son, and the two of them would talk as equals, and I sometimes wondered if he used the same principle when he talked to the rest of us."

Frost1x··on GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]
I’m all for it since it’s value directly returned to humanity.
Frost1x··on Potential session/cache leakage between workspace instances or consumer accounts
I noticed you were linking a file vs creating a correct CLAUDE.md implementation. Would you like me to fix that for you?
Frost1x··on Please stop the AI confidence theater
Marketing is just a proxy for the underlying goal: growing profit. So any evil done by marketing is driven by the pursuit of wealth. Greed underlies anything marketing does as a purer form of evil.

There’s plenty of marketing out there that just tries to make information about a product and service available without focusing on driving home higher revenue at any cost. That’s usually advertising, not marketing though, but it does exist.

Frost1x··on Claude-real-video - any LLM can watch a video
So I did this yesterday for a video analysis sample with ChatGPT and it took the video, pulled out frames, did difference tests across the frames to look for significant frames to focus on, did image recognition on each frame, and interpolated motion and action between.

So I’m not sure why this says ChatGPT doesn’t “see” video and reads transcripts. Obviously if the video is already labeled that’s the shortcut. But it did an impressive job describing a video I have no inclination it would have in its training data. One could argue it wasn’t “native” and had an agent orchestrator to rely on external tools to accomplish the goal… but it worked.

Frost1x··on Virginia bans sale of geolocation data
Good question, I’m curious too. 911 services and cell providers come to mind, as well as subpoenaed data from law enforcement? Perhaps?

Third party commercial entities like cell providers are collecting and sharing it out of necessity but I’m guessing not selling it?

But that opens an interesting loop hole it seems where you could open a share agreement and then through other mechanisms recover the fee you’d otherwise charge for.

Provider A wants to sell data to provider B and provider B wants to buy from provider A but they legally can’t. So instead provider A just tucks the cost in some other unrelated contract with provider B with a wink wink, handshake, nod, their “relationship” then just makes them want to share the data at “no charge.” Both know the fees are tucked in other agreements, although only provider A knows the itemized cost, provider B just wonders if the cost of the other package + their friendship handshake sharing of geolocation data is worth that total cost.

To be fair, until money comes into play people tend to be less nefarious about their uses of information and intentions. Not always, but on average.

Frost1x··on AI can't be listed as inventor on patent applications, Japan's top court rules
AI use is slowly creeping into pure mathematics and proving theorems or providing legging to mathematical breakthroughs. Just go watch some Terrance Tao videos to see some recent work. In addition, theorem provers and the likes have been around for awhile. Some of these systems create novel ideas or bridge novel ideas in ways that are arguably not “obvious” in any sense of the term.

While as a species our key strength has been our intelligence and it’s been core to our identity, and computing has slowly over decades infringed on this forcing us to rewrite what it is to be human, I understand the defensive view.

I also see LLMs and other AI systems spit out complete nonsense that’s truly obvious to most people. But that doesn’t make any of these systems, in my opinion, incapable of creating or bridging novel new ideas that I would call far from obvious had we substituted a human in place of it. I didn’t look at the patents in question, plenty of obvious patents make it through anymore, so that could be the case here, but I believe AI isn’t far away if not already there of creating truly patentable inventions if someone were to push it.

Frost1x··on Why software engineers are grieving
Then we need to redesign our entire economic system because none of it hinges on you enjoying your productivity to survive or reaping the rewards. Not saying I entirely disagree with you, just saying our economic system isn’t configured for the theoretical ideals you suggest. I’m not sure we can do that or the people with enough wealth and power to shift things would ever want to change these things.
Frost1x··on DeepSeek V4 Pro beats GPT-5.5 Pro on precision
I’m not sure all models will converge on your acceptance criteria. I’ve done quite a bit of varied agent based modeling and scientific modeling in that domain and just because you have some grounding to check against and some ideas on how you might go about getting to a convergence point doesn’t mean you’ll actually converge, you can absolutely get stuck in the information space iterating away, never finding your desired solutions.

It helps but you often have to step in the failure cases and guide them or forcibly fix certain paths to get a solution.

Page 1 of 34Next →