HNHacker News
TopNewBestAskShowJobs

Borealid

1,063 karma · joined November 3, 2014

submissionscomments
Borealid··on Various Projects Find Hidden SDR Capabilities in ESP32 Microcontrollers
Unfortunately the use of the Wi-Fi 5Ghz spectrum is not at all in sync worldwide. See https://en.wikipedia.org/wiki/List_of_WLAN_channels#5_GHz_(8... and how meaningfully different the rules are for how you can set up your own home wireless network in different parts of the world...
Borealid··on Various Projects Find Hidden SDR Capabilities in ESP32 Microcontrollers
I believe the default state, in the absence of any law whatsoever, is that you can do whatever you want. Nothing is forbidden by default. You could assault another person, and it would be legal absent a law forbidding it.

Then there are laws that prohibit things like broadcasting radio signals without a license.

Then within those very same laws there are exceptions allowing certain transmissions WITHOUT a license, in specific circumstances.

The "default" is that things are legal unless they are made illegal, and I think it's pointless wordplay to try to argue the default state is somehow the law prohibiting unlicensed operation, and not the part of those same laws that allow unlicensed operation in particular bands with particular transmission power limits.

I understand the point you're trying to make that some parts of U.S. law say that all not-explicitly-mentioned radio transmissions are unlawful, but that's not meaningfully different from drug policies that say that all "harmful substances" are unlawful and so forth. They all just boil down to what the law defines as being within its own scope.

Borealid··on Git 3.0's upcoming SHA-256 default will be a costly mistake
Correct. The content of the git `tree` is partially controlled by the attacker because filenames in the repository are part of it. So if it's feasible to manufacture a generic hash collision it MAY (not MUST) be feasible to generate two colliding `tree` objects.
Borealid··on An AI sovereign wealth fund isn't progressive – it's techno-imperialism
The specific problem isn't "it's not X, it's Y", it's using that pattern where X and Y are not correctly related categorical items.

Yes, I know how I structured that sentence.

"It's not a frying pan, it's a revolution" sounds like AI. "It's not an octopus, it's some type of crustacean" does not. The thing that makes it an AI tell is specifically FALSE analogy.

Borealid··on OpenDLSS: A Vulkan Reimplementation of Nvidia's DLSS 5 Neural Rendering Network
Am I the only one who feels a sense of disinterest in a project where the main README is LLM-generated? Does the author not have time to write what they did and how it's used?
Borealid··on The last time my family was replaced by technology
That doesn't make intuitive sense to me.

First, you make a robot that mines rare earth metals. Then construction robots. Then an automated factory for assembling robots. The whole premise is that the robots are faster at these tasks than humans would be.

If the robotics are feasible their rise should be exponential, not linear, since each robot built makes building the next faster and helps to remove whatever bottleneck you had.

Borealid··on 500k facial scans at UK stations yield no arrests, 1 false positive
We have no reason to believe a given face is linked to a person's name by this system.
Borealid··on 500k facial scans at UK stations yield no arrests, 1 false positive
The data necessary to determine a face had previously been seen are biometric data.
Borealid··on 500k facial scans at UK stations yield no arrests, 1 false positive
If we're going to be precise, the word "data" is plural and its singular is "datum". But treating it as a substance noun does not change the argument I made in any way - "I gave away the water" does not mean "I gave away some of the water (and kept some for myself)".
Borealid··on 500k facial scans at UK stations yield no arrests, 1 false positive
The sentence says "the [...] data [...] are automatically and immediately deleted". It does not say "some of the data are automatically deleted". It would be a staight lie to delete some, but not all, of the data and then issue the statement that "the data are deleted". This is not an ambiguous case or weasel-wording: it is an unambiguous falsehood.

If you have two coins and you spend one, you cannot truthfully say "I spent the coins". The definite article "the" means all of something you have in this usage, not a single member of a class of something you have. It would be true to say "I spent a coin" or "I spent some of the coins".

Borealid··on California farmers are struggling to sell grapes as demand for wine drops
Question: what is the ratio of US wine imported into the EU vs US wine imported into Canada, by dollar value?

Question: what is the percentage increase in tariff amounts on wine between the pre- and post-USMCA state of play between the US and Canada?

Sometimes seeing why something has happened requires considering wider context, which is why it's usually good for a nation to ensure primary education includes solid economics, history, media literacy, etc curricula.

Borealid··on California farmers are struggling to sell grapes as demand for wine drops
Countries other than the US have, for reasons that you may choose to find mysterious, imposed tariffs on products exported from the US.

Tariffs raise the consumer-visible price of those goods. This then, due to demand elasticity, causes consumers in those non-US countries to buy fewer of those products.

Some impact may be badwill as well but quite a bit is first-order impact of second-order tariffs.

Borealid··on Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest
I meant vfio-pci.
Borealid··on Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest
The percentage-overhead comparison is pretty choice nonsense. It has only percentages to try and "explain" that overheads don't matter if the system is slow anyhow.

A fair comparison would be this project vs virtio.

Borealid··on Let's make quality the norm again
There absolutely are still rich-people brands that sell absurd quality at absurd price.

See, for example, Marine Layer in San Francisco, or boutiques like Red Ants Pants.

Borealid··on Why is Google still serving dodgy ads?
I can believe that, I'm just here to point out that some of the laws being not-broken are the laws of thermodnamics. You can cool a room cheaply with dry ice, for instance. There are tricks.
Borealid··on Why is Google still serving dodgy ads?
There do exist "evaporative coolers", sometimes called "swamp coolers". The trick is that they require water and raise the humidity of their surroundings. But they tick all the boxes you listed - plenty of delta-t, extremely low power, no external tubes or parts.

It's not magic because it relies on the "fuel" of cool water to evaporate.

Borealid··on Large language models develop novel social biases through adaptive exploration
It's a bias even if the true population distribution isn't linear.

For example, if you have a training corpus where 50% of the text follows "black bobblehead" with "arrested" and 20% of the text follows "white bobblehead" with "arrested", and your LLM is trained such that it produces "arrest" 50% of time after "<color> bobblehead" regardless of color, that's a bias - the output frequency distribution fails to match the "population" (training) frequency distribution. This has nothing to do with races, ethnicities, whatever - it's just statistics and text. To be unbiased, it would need to be less likely to produce the text "arrested" after "white bobblehead" than after "black bobblehead".

A die is supposed to land on each face evenly - a linear probability distribution. So anything other than a linear distribution is biased. But bias can exist for any desired probability distribution. And for an LLM the desired probability distribution of the model output is one that exactly matches the infinitely-many distributions of the various facets of the training data.

Your point about how in the absense of information a token shouldn't influence the distribution is spot-on. But unfortunately almost any token does condition the output, which means you get biased output all the time.

Borealid··on Large language models develop novel social biases through adaptive exploration
No, a "bias" is a statistical term meaning a probability distribution that has an expected value differing from the population's expected value.

A human's discriminatory bias against an ethnicity is just one type of bias. The LLM isn't a racist, it merely produces text where that text does not perfectly reflect the training data's frequencies.

Borealid··on I resigned from Anthropic today
The reason nuclear reactors are dangerous is because if you turn off the power cooling them down, they react (and radiate) more.

If you turn off the power cooling a data center, the servers within rapidly stop doing any computing.

Positive feedback loops are dangerous. Negative ones self-regulate.

Borealid··on Large language models develop novel social biases through adaptive exploration
My comment is literally explaining the result of the paper, in which it is shown that LLMs can and do develop biases based on text appearing in their training data set even where such text is not in any training example connected with a systematically more positive or systematically more negative outcome.

In other words, if the text "X is wet" and the text "Y is wet" and the text "X is dry" and the text "Y is dry" each appeared exactly one time in the corpus, it's still possible for a model to end up being produced that is more likely to write wet-like words when it sees X in the context window than when it sees Y.

On a side note, it's very unrewarding to try to explain this type of statistical observation when it feels like (anecdotally, hypocritcally...) the entire world wants to use words like "think" and "understand" and "pick up on" to describe inference and training processes. I'm not making a stochastic-parrot argument here, just pointing out that understanding an LLM's behavior is best done by understanding its conditioning.

Borealid··on Large language models develop novel social biases through adaptive exploration
I think you're missing the point of TFA.

The LLMs take in text which conditions their output. That means even nonsense text - such as a "tribal affiliation" to a tribe that may not have ever existed - ALSO condition the output, because the tribe name is a token in the context window and there's no such thing as a perfectly neutral token.

Taking away the race/ethnicity layer for a moment, it might be that an LLM develops a predisposition to emit positive terms (like "accept") when the prompt contains "banananow", and negative terms when it contains "pearian". That's the very definition of bias, and hacking those biases could give individuals serious socioeconomic benefits!

Borealid··on We Must Return to the Office to Use AI in Person
That's not how training works.

Training tries to produce something that scores highly in training evaluations. With one data point, the evaluation is solely how closely the model output resembles the single input text.

Let's say you do that, and the training text is 58,100 tokens long. Let's say you ask the model to produce 58,101 tokens. Will it "reproduce [the] text verbatim"? No, it can't, because of the dissimilar requested length. Something "new" will come out.

It's also entirely possible that no matter how long you train, the model never converges on generating exactly the same output as its training data - you could end up with an average loss value of 0.001 instead of 0.0. It's not a perfectly deterministic process.

You're correct in principle, but in reality even with limited training data real-world models produce something that isn't exactly their training, especially when sampled stochastically. They're biased toward their training data, not forced to it.

Borealid··on DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
KYC stands for Know Your Customer, the regulations that require institutions moving money between two parties to positively identify each of those two parties.

I think the intellectual position "it should be illegal for institutions transmitting money between two parties to identify either of those parties" might require some kind of logical argument behind it. Are you saying all financial transactions should be anonmyous by law? How would banks function if they were required to be blind to their customers? How would the government prosecute money laundering if all cash-trails went cold after the first time they passed a bank?

I understand people often like to express extreme positions on the Internet, but I think it's pretty easy to see an ideal society has rules somewhere in between "you're not allowed to know your customers" and "you can't accept a penny unless the giver does a blood draw in front of you and is confirmed to be in a central register of DNA".

Borealid··on Jellyfin 12.0
I don't think this is what the GP meant, but he's accidentally correct in that there is indeed no good reason for a password to reach the server.

The client could do a challenge-response PAKE with the server, proving the client held the password while not transmitting it. The session cookie could be sent from the server as part of that exchange, encrypted so only the holder of the user's password could decrypt it.

That's not how the web works, but it's strictly superior to the accepted standard of the client sending the raw password to the server, and it would be secure even over an active MITM unencrypted link.

Borealid··on "Next-token predictor" is the wrong mental model for LLMs
> and by the same token

I don't think you intended this, but the word choice here gave me a chortle.

Borealid··on “Next-token predictor” is the wrong mental model for LLMs
I think the most useful word in both cases is "extrapolating".

An LLM extrapolates from its context window to the immediate next token. This word applies whether you view what's happening as "reasoning", "prediction", or as a math function.

Borealid··on How Fairphone built the Fairphone Gen 6+
I believe the GrapheneOS team demands, in particular, a security feature called Memory Tagging Extensions (MTE). This feature has a phone's processor attach a "tag" to each memory location saying what's stored there and requires accesses to that memory to bear a matching tag. Buffer overflow type attacks are effectively prevented because a read into the adjoining region having a different tag gets rejected.

This feature is usually NOT present in desktop computers, and has only really been implemented by the Google Pixel hardware so far. It's unclear why it's viewed as being so critically necessary for Graphene: MTE is certainly useful but I personally wouldn't say its absence indicates an "insecure" device, and it imposes both a clear performance (and battery life!) penalty and a significant hardware burden on the SoC manufacturer. It's a pretty significant trade-off.

Anyway, that's the biggest single reason why GrapheneOS devs say "this hardware is not adequately secure": 99% of phone chips are excluded because they do not support MTE.

Borealid··on GrapheneOS says Pixel 11 has MTE support after all
I do not feel you addressed my point at all.

I want to install an app now, and protect against its developer doing a rugpull on my local files later. I do not want to need to review each new app version in advance as I install it.

This is clearly a legitimate case where what the Android security model says I should be permitted to do makes me less secure against an attack by the app developer.

Borealid··on GrapheneOS says Pixel 11 has MTE support after all
Let's say I want to secure my system against an app developer deciding to delete my data stored in their app.

To do that, I wish to store a copy of all the files the app has written to my filesystem, and put that copy outside the app's control. This is explicitly to contain data the app's developer does not WANT me to be able to keep.

A. Is my being unable to do this "more secure"? If so, why is the specific threat I described to my data integrity - an app developer deleting my data - invalid?

B. Does GrapheneOS support this protection, ensuring the device owner is secure against the app developer, or do they instead secure the app developer against the user?

Page 1 of 12Next →