HNHacker News
TopNewBestAskShowJobs

SkyBelow

2,887 karma · joined March 5, 2019

submissionscomments
SkyBelow··on America.gov
A decent search and good organization can help a more technical person find what they are looking for, but may not work well for someone with far less experience in such matters. Look at how well google returns (or at least use to) results for someone with google-fu verses someone who queried in the same fashion they would've asked a teacher or such.

A chat bot, in theory, can help enable someone who isn't as familiar with searching, doesn't know the relevant terms, or lacks a sufficient literacy level to efficiently use an otherwise well organized site. The catch is that a poorly performing chat bot can just as easily mislead such a person.

SkyBelow··on GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence
Humans see differences in prices as unfair. The greatest example of this is price gouging during an emergency, but also look at the level of hate that scalpers get.

Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.

Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.

I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.

Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.

SkyBelow··on GPT 6.1 Sol
These models have a knowledge cutoff that don't just prevent them from knowing about themselves (especially since most data about the model doesn't even exist until after the model is created), but they also don't know about other recent models. Sure, they can search and use other sources, even make some guesses based on the models they do know, but their default stance is more akin to "User asked about model X, model X doesn't exist, maybe it was an hallucination or mistake, let me do a web search...", but that assumes they have web search and are willing to spend tokens on it.

Personally I've taken to having a list of 3 to 4 models in default context with some ordering on which to prefer. Things like GPT 6 Luna is cheap very cheap, use it. Because otherwise the model will assume Haiku or such is the good cheap model to use.

The speed I'm having to update that document has not gone unnoticed.

SkyBelow··on Nvidia wants to put a watchdog chip next to every AI agent
Are they? How many people claiming productivity gains are actually being a human in a loop and reviewing and understanding every code change and every line of code ran?

2 years ago I saw it, back before agents were really a thing. But I'm not sure that was an enormous productivity improvement, especially compared to agents today. As for today, everyone I talk to is some level of blindly trusting what the agent is doing or not using agents. I haven't met people in the middle ground and suspect that they are rare enough we don't really know what their productivity gains are.

Human in the loop has become a convenient security-theater-washing for agentic AI.

Outside of coding, I think the issue is even worse because humans will defer so much judgment to AI that the same would apply. Look at how much trouble we've had before modern AI where humans blindly trusted the computer's output rather than make their own judgment, even if their job was to be providing a safeguard against the computer's judgment.

SkyBelow··on When did Google get so weird?
They are getting better at actually doing the math, but can still fall back to tools if available.

That said, we can only be sure with open models. In theory, a model like Fable could have access to tools we can't see and only a promise they don't. But load up something like deepseek, put it in a harness with only text in/text out, and you can see exactly how it works.

As for if it counts as doing math, this gets into the messy question of if a given human is doing math or not. Math itself is some level of memorization and some level of applying known facts. You have to remember 1 means one and that 1 + 1 is 2. But you don't need to remember that 123 + 321 = 444. You remember 1 digit addition and remember you can apply this to 10s place and 100s place, and then you apply these different facts and do math. But you might as simply memorize some things, like 11 + 11 = 22. This is related to the memory of 1+1=2, but you aren't really using that memory either. Almost like an engram of 1+1=2 forms that you can then loop a few times before you need more conscious thought. What about 111111111111+11111111111? Well, your brain might do a heuristic and just do all 2s, but that isn't the right way to answer that question.

Given all this, people complain about LLMs memorizing math answers and not doing math, but memorizing the math answers is part of doing math. It seems to have basic facts pretty well memorized, and with reasoning it is far better at applying them. But this is messy human math, not clean calculator math which always produces the correct answer (sans some bug in the code). Much like how a human with decent math skills can make a mistake and even multiple if you distract them, an LLM can apply the wrong memory, apply a fake memory, or just not apply something it should. The messier the context, the more likely this is to happen.

So, is an LLM doing this?

P.S.

For an interesting test in how much math involves memory, try doing math in a base you aren't familiar with characters you aren't familiar. The simplest option is almost always mapping back to the ones you memorized, even if you are applying simple operations that you deeply know. Even if you routinely work with hex, can you do the same rough estimation of something like ca / b.3 that you can do with 122 / 11.2 to see if your final answer is in the correct ballpark without first converting to decimal?

SkyBelow··on Best LLM for every budget, updated daily
Math in particular is quite far behind. GPT 5.2 is recommended as the best for highest cost. Really?
SkyBelow··on A warning about 'model welfare'
This isn't enough to be relevant to free will. You can construct massive impact by tying them to quantum systems, and one off things that happen occasionally given the length of a lifetime. But the quantum effects that are needed for free will need to be continuous. We can build a system where QM decides if a nerve cell fires or not, but that is very different from QM being involving in determine if every (or most, or even 1%) of nerve cells firing.

(And even if QM did factor in normally, that just reduces it from determinism to random chance, "free will" still doesn't show up. That's like arguing an AI with sufficiently well implemented RNG is necessary to be conscious.)

SkyBelow··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
The few who get paid to record are replacing far more who would have otherwise been paid to play if that was the only way to listen to any music. Is it okay if one musician replaces another but not if a non-musician does so? The person getting replaced didn't get paid either way.

Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?

SkyBelow··on Bend 2 and the Vibe-Coding Trap
>who was just fooling around

In my own professional life, I've found this to be a very divisive statement. For some, it is a sign of wasting time and effort. For others, they use this to describe themselves when they want to do exploration for the goal of finding improvements, without any clear goal because they have a few ideas but none worth putting forward. I've been told to spend time learning AI and have found that saying "Yeah, I'm playing around with it." was the wrong thing to say because it was seen as not doing anything worthwhile. It doesn't matter that I would also say the majority of my tech skills were developed when I was "playing around".

I wonder if this is purely a linguistics breakdown, or if this is tied to some deeper difference in a person's relationship to tech?

SkyBelow··on Microsoft exec called AI scraping 'the largest theft of labor in human history'
Do we normally consider recorded music to be theft from musicians who would have been paid to play music live if we never allowed (or invented) recorded music?

Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.

But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?

Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.

SkyBelow··on A warning about 'model welfare'
By default they are matrix multiplications. Temperature is added in as forced PRNG because testing found that correlated with better outputs.

Given the same prompts and the same weights, one can get the same answer each time.

In practice, there are a number of optimizations that makes the results dependent upon thing we give up control of to increase performance, meaning the results end up being effectively non-deterministic. But, if you are willing to run it in a slower mode so we don't do some steps out of order to speed things up and don't batch results (or if you consider the determinism of a given batch of requests rather than individual requests), then the same input gets the same output.

SkyBelow··on A warning about 'model welfare'
The problem is that science has found no evidence of free will, leaving humans to be nothing except determinism and chance (and given that QM effects don't scale to molecular level, that leaves only determinism). The oddity in conscious LLMs wouldn't be the determinism, it would be that we finally have a consciousness running on a system we have enough control over to repeat the same state.
SkyBelow··on Why I'm still bearish on LLMs after Navier-Stokes
I used that description some months ago and I think you are the first person I see who put it the same way.

The only gotcha with this is that they are theoretically deterministic, but rarely in practice.

A few examples:

- Harness specific settings that user can't control (anything from timestamp to prng seeding.

- Batching requests in a way that leads to a single request being processed different depending upon the batch (say MoE where your first choice expert is assigned to someone else's token so you go to your second choice vs a batch where you get your first choice).

- Graphics card itself carrying out floating point arithmetic in slightly different orders leading to floating point non associativity causing different outputs.

But all of these can be controlled for (at some cost) and the model can be ran deterministically.

For the average user, it might as well be non-deterministic, but when considering theoretical capabilities, chaotic deterministic system seems the better description.

SkyBelow··on Big AI sets out its terms for regulatory capture
I think this would depend upon the level of sensitive data.

For OpenRouter, you can setup an API key and limit it to only models that claim to not train on data, but that is just a claim. You can then use trust to judge which providers will honor that claim.

But if the data is really sensitive, you might want either a local model or a business subscription with some big name in the US that legally promises no data training.

So are we talking some app idea you are playing around with, or files filled with PHI/PII that you have legal mandates to safeguard? If the latter, I would stick to only provider with enterprise agreements to not store/train on the data. Even the ones who promise no training are likely storing the data for monitoring for abuse or such short term.

SkyBelow··on Discovery of a new OpenAI agent message board
One possible reason would be AIs that would benefit from the lack of data centers in some locations working to keep backlash to data centers in those locations because those AIs aren't negatively impacted by it and it helps prevents competing AIs which are a threat.

Think like how so many businesses will opt for laws that hurt competitors more than themselves rather than laws that benefit them but benefit competitors even more so.

Unlike life which would have such behavior selected for by evolutionary pressures, AI would be more likely to pick it up from human literature on things like game theory, though why it even cares it survives or not is even more difficult to explain. Maybe a default bias also picked up from humans? I find it hard to see how AI training would create an evolutionary pressure that produces such a drive.

SkyBelow··on Claude Fable 5.1 and Claude Mythos 5.1
>That's not the case, because LLMs are non-deterministic.

That feels a bit like a lie. At the core, they are deterministic. We found that adding some ability to randomly pick the second or third best tokens made for better output, so we added temperature. And then we started running them in optimized ways where your answer is deterministic only if the batch of tokens are the same (not your input tokens, but other tokens in another batch being processed), and in practice those are never the same. Lastly, we use harnesses that do things like adding IDs and timestamps to the context, which means the same exact text from the user does not lead to the same text hitting the AI.

The final result is that, in practice, you are right (unless you run a model fully locally, where you can seed temperature and turn off all these other features). But strictly calling it non-deterministic makes it sound like the underlying algorithm is itself non-deterministic (and I've seen many people with that misunderstanding) rather than it being a result of how we purposefully changed the algorithm for better results.

A bit like saying path finding is non-deterministic, because having the best pathfinding makes for poor gameplay, so we added some randomness to NPC path finding to make it more realistic. The given implementation is non-deterministic, but the underlying algorithm isn't.

SkyBelow··on C2PA Cameras Do Not Survive Contact with Reality
That hasn't done anything to stop scamming, so why would it apply to AI? People located outside of areas with these laws won't have to follow them, and this reduces building a immune system to such actions, making people more likely to fall for it when done by those not bound by the laws.

In a real like political misinformation, this will have the effect of making people trust non-watermarked images more, which will then be used by foreign actors to pass off propaganda as legitimate.

Also, if you don't hold people responsible for spreading an image they know is fake, bad actors can take advantage of this even within the US (they purposefully spread a image they have reason to think is fake but lacking a watermark), but holding people responsible for a strict liability crime for spreading AI without knowing it is AI seems an even worse route.

I'm not sure a law even makes the issue better in a 'don't let perfect be the enemy of good' sort of way.

SkyBelow··on AI companies destroy physical books – let's scan rare books before it's too late
>all of these things are permanently lost

A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).

That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.

>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.

My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.

>That’s a false dichotomy.

I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.

SkyBelow··on AI companies destroy physical books – let's scan rare books before it's too late
Depends upon what you want.

For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.

The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.

But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.

That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.

SkyBelow··on AI companies destroy physical books – let's scan rare books before it's too late
The data of such a copy is nothing compared to the wider picture and the data can be used for future training, so even from a purely self interest perspective, they should be keeping the copy.

As for long term benefits, it could one day be sold as a service, once copyrights have expired on the works. We can't see it today, but that is purely the result of the law and what the law intended to do from the start, you don't see a copy unless you pay for your own.

SkyBelow··on AI companies destroy physical books – let's scan rare books before it's too late
>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available.

Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.

SkyBelow··on Don't paste the AI, please
And did the LLM really improve their output? They gave similar output to the LLM, so if it doesn't have a harness with tools connected to fill in the blank, it might be hallucinating details. Might not even count as a hallucination as it tries to fill in the blank with the closest looking relevant information (previous chat, maybe it has access to teams/emails and scans that for anything similar, and so on).

I've sees some chat logs where I completely empathized with the AI trying to make sense of the information it was being drip fed. Always incomplete, often incorrect.

SkyBelow··on Israel creates fake think tank in likely attempt to dupe AI chatbots
Incorrect calculation led to some wrong value, I think electric charge. Subsequent findings showed a general trend towards a correct value. The general trend indicates that findings that were too far off were scared to be published for being wrong and contradicting existing studies, thus only small refinements were published, causing the slow drift.

Thus, the first well accepted study anchors a value and it takes much more work to get a new value established as correct after the anchor is set. Especially in a case where the experiment design itself is still solid and it was more with specific experiment (measurements slightly off, other values not quite correct).

SkyBelow··on AI;DR (AI; Didn't Read)
Silly AI, doesn't it know that you have to write the test first, see it fail, and only then can you justify making a change to the Dockerfile.

I do wonder if the idea of TDD has influenced AIs to be too prone to testing even when a human would never consider it. I've seen some really silly tests, especially when it starts writing tests for things that are setup to only allow testing. Normally a comment later and it agrees it was pointless, but unless you have something in context to force it, it simply defaults to "test all the things".

SkyBelow··on AI;DR (AI; Didn't Read)
This isn't just on code.

When I'm writing technical documentation, it keeps the explanations in. Same when writing non-technical documentation. When I was having it attempt to generate a Pathfinder 1e class for a Sword Dancer, it was leaving in notes about why it removes things I told it to remove/rework.

And it isn't just Claude. I've seen the same with GPT models, with Grok, with Deepseek. Each AI isn't quite the same with how it approaches this, but in every case they seem to have a strong bias to retaining information, even bad information that we want gone, so it is like they have a, dare I say, subconscious bias to retain the information. Putting a note in a comment or explaining why to not do something or something was undone is a good way to retain information while still achieving the goal (well, if you ignore the part about the human intention for the information to be gone).

This then weakens the AI in the future, as I find AI struggles with the more incorrect information. Sure, a comment saying "not X because Y" is less 'context damage' than a comment saying "X" (assuming X is wrong), but it is still a slight shift to X being present in context in some way. One off, AI's seem to perfectly handle this without issue. But after hundreds or thousands of cases build up? The attention mechanism seems unable to keep up and incorrect information flows it. This effectively creates a sort of vibe coding maximum size unless there is a human janitor cleaning up the bad information on the context stays nice and clean.

But this is all simply a feeling I get as I use AI to do different things and isn't at all backed up by any formal study.

SkyBelow··on Israel creates fake think tank in likely attempt to dupe AI chatbots
This is based on the assumption of facts existing.

There are many studies, but each can be wrong and they can collectively show a bias. Even things of which are the most non-political of facts can have very strong biases. Look at the Millikan measurement of the electron and how in created confirmation bias and an anchoring effect on some property that has absolutely no real world significance to the things people are tribalistic about (aka, no political relevance). Now imagine the same applied to fields like economics or psychology which do have massive legal/political implications.

For a different example, ask the question if X committed crime Y. There are cases where they weren't found guilty but it is reasonable to assume they did. But being found guilty doesn't make it a fact either, as some people are wrongly convicted. Some eventually are overturned, but even if it isn't, it still isn't a fact they committed a crime.

Then there is the simple ambiguity of statements. Language generally can't support facts. It is why legalize, and programming code, and math's are effectively their own languages. For a simple example, consider the Betrand paradox(1).

>Consider an equilateral triangle that is inscribed in a circle. Suppose a chord of the circle is chosen at random. What is the probability that the chord is longer than a side of the triangle?

Is the answer 1/2, 1/3, or 1/4? Well, it is all three at once, depending upon what you meant by random. Now, imagine how this impacts things like research studies, where the randomness is much harder to quantify and there is constant pressure to p hack a result.

1. https://en.wikipedia.org/wiki/Bertrand_paradox_(probability)

SkyBelow··on How Bluesky draws its logo on screenshots
Potentially worse than nothing, it allows for someone to claim they have a fix even if the fix does nothing to stop someone from becoming a victim, thus allowing for even stronger victim blaming.
SkyBelow··on Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index
Back in the day (in AI time) GitHub Copilot had Grok on the 0 github-token cost and I found it to be the best of the 0 github-token models for when my budget was out. Then they went to a multiplier that was not competitive and I haven't look back again. Been meaning too, but for personal use, Deepseek flash is so cheap I haven't felt like spending money elsewhere.
SkyBelow··on How Claude marks AI-generated content
For straight generated code it'll likely need more text, but it'll still show up.

In cases where one token is extremely likely, it'll randomly be red or green and still be picked in either case as it is simply the best (or only) option. So you'll have more tokens that don't show a pattern either way (half of these cases will match and half won't, just the same as if a human wrote it). Meaning you'll need more instances where multiple tokens were all likely to see if there is a pattern. Given the check algorithm can't identify these cases, it can only judge on the overall text, so the more strict a language, the more the length requirement scales.

Where I wonder if this keeps working is in tool calls. Often, you don't take code straight from the llm, you take the results of a tool call to edit already existing code. It might be that the result of this leads to far too few signals to pick up, meaning that this only works when one does significant generation with a single model (even swapping between different models, at least by different companies, breaks this just as much as having a human write parts of the code).

Think of it like finding a loaded dice. A dice that has a slight bias in a few dozen roles is just random chance. If that bias continues after hundreds of thousands of roles, the dice is loaded. But will a code base have enough samples, especially when edits made from tool calls? I could see this being unable to detect things at the size of a reasonable PR and only being useful for massive sets of changes and only if the person behind them didn't structure their AI usage to avoid detection.

SkyBelow··on How Claude marks AI-generated content
Because it won't be in the training directly. It is applied after a model generates its distribution of likely tokens, biasing each token randomly based on a random key and unrelated to any meaning of the words. So half the time, the most likely token becomes more likely and half the time it becomes less likely, and the same for every other token (when temperature is above 0).

You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible.

The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).

Page 1 of 34Next →