HNHacker News
TopNewBestAskShowJobs

cheald

14,356 karma · joined August 19, 2009

chris+hn@heald.me

[ my public key: https://keybase.io/cheald; my proof: https://keybase.io/cheald/sigs/piAgm06dhM9eFLAHy4jVtO3rIY5-emPB4MRJ91EKOZU ]

submissionscomments
cheald··on TikTok says it is restoring service for U.S. users
We're the product of all the information we consume, but forums like HN aren't custom-tailoring what I'm shown to maximize engagement from me.
cheald··on US Export Control Framework for Artificial Intelligence Diffusion
It might be somewhat prohibitive to print the model weights for any sufficiently large model, though.
cheald··on Platforms systematically removed a user because he made "most wanted CEO" cards
Given that they were made immediately after and in response to the shooting of the United Healthcare CEO, it requires some fairly lithe mental gymnastics to not interpret them as an effective call for murder.
cheald··on You can't optimize your way to being a good person
Au contraire, I wear very large shoes, so my step size becomes a limiting factor at some point.
cheald··on Why does storing 2FA codes in your password manager make sense?
I don't think that's true at all. 2FA has been a popular solution for many years, well before the addition of TOTP support to the popular password managers.
cheald··on Why does storing 2FA codes in your password manager make sense?
Which is precisely why it's irresponsible to give people the rope to hang themselves with by supporting 2FA seeds in password managers (much less telling them it's a good idea), IMO.

People take the path of least resistance; we know this. It's why, for the longest time, people used one password for everything. People don't like using password managers, either, but we would all agree that it's unacceptably insecure to not use them, because the alternative is "one password used everywhere, maybe with a single varying digit on the end".

cheald··on Why does storing 2FA codes in your password manager make sense?
I think it's a terrible idea, because it dramatically decreases the attack surface area needed to compromise accounts. 2FA is supposed to be "something you know' and "something you have"; putting your 2FA seeds into your password manager reduces your 2FA to "something you know", and, significantly worse, it's "something you know in the same place as the other thing you know".

The time-variant component is still quite valuable, but it does nothing to protect you in the event of a password manager compromise. This is not a hypothetical; LastPass has suffered multiple breaches, and the more popular a solution, the more likely there are to be attacks against that solution. By keeping your 2FA separate from your password manager, even if it's still just "something you know", it's something you know in a location that's orthogonal to your passwords. If I yield to convenience and use a 2FA desktop app, then now, instead of just attacking my Bitwarden install, you have to successfully attack my Bitwarden install and my 2FA desktop app install to get access to my accounts, and the combination of password managers * 2FA managers is a substantially larger attack surface and requires a significantly more sophisticated attack to get both pieces.

The arguments in the article come down to "well, 2FA mitigates phishing attacks" (true) and "Google Authenticator means you can lose your data easily" (also true). But neither of these is a good argument for why the data should be kept together. It just means "use 2FA", and "use a 2FA manager that lets you directly manage your seeds and keep offsite encrypted backups".

If you can't be bothered to do it properly, then 2FA codes in your password manager is certainly better than not using 2FA at all, but that just makes it a less terrible solution, not a good one.

cheald··on Beyond Gradient Averaging in Parallel Optimization
As an experiment, I tried implementing this for Stable Diffusion lora training, where I'm training on a single GPU with a batch size of 8, and it does actually seem to have an appreciable impact. In my case, I'm keeping a per-parameter grad EMA, and then computing the cosine distance between the parameter's grad and its EMA, and then multiplying the grad by 0 if (1.0 - cos_sim) > 0.99.

My loss metrics stay roughly the same (they're slightly lower, but SD loss is fraught to interpret because variance by timestep renders it more or less meaningless), but tracking the means of `param.grad.norm / param.numel` (which shows how big the grad updates are) shows the grads stabilizing significantly quicker than baseline. I'm tracking suppressed params / total params via tensorboard, and I show that it drops (as expected) but then stabilizes at around 7%, suggesting that there are model parameters which consistently don't agree. I'm gonna try tracking the variance from the mean, as well, and perhaps down-weight or eliminate grads for parameters which show high cos similarity variance over time (suggesting a generalized lack of agreement in the direction to move, further suggesting that the parameter cannot contribute meaningfully to the task).

cheald··on Why OpenAI's Structure Must Evolve to Advance Our Mission
This is specifically why I caution people against trusting OpenAI's "we won't train on your data" checkbox. They are specifically financially incentivized to do so, and have a demonstrated history of saying the nice, comforting thing and then doing the thing that benefits them instead.
cheald··on Can AI do maths yet? Thoughts from a mathematician
OpenAI's also in the position of having to compete against other LLM trainers - including the open-weights Llama models and their community derivatives, which have been able to do extremely well with a tiny fraction of OpenAI's resources - and to justify their astronomical valuation. The economic incentive to cheat is extreme; I think that cheating has to be the default presumption.
cheald··on Can AI do maths yet? Thoughts from a mathematician
LLMs have been very useful for me in explorations of linear algebra, because I can have an idea and say "what's this operation called?" or "how do I go from this thing to that thing?", and it'll give me the mechanism and an explanation, and then I can go read actual human-written literature or documentation on the subject.

It often gets the actual math wrong, but it is good enough at connecting the dots between my layman's intuition and the "right answer" that I can get myself over humps that I'd previously have been hopelessly stuck on.

It does make those mistakes you're talking about very frequently, but once I'm told that the thing I'm trying to do is achievable with the Gram-Schmidt process, I can go self-educate on that further.

The big thing I've had to watch out for is that it'll usually agree that my approach is a good or valid one, even when it turns out not to be. I've learned to ask my questions in the shape of "how do I", rather than "what if I..." or "is it a good idea to...", because most of the time it'll twist itself into shapes to affirm the direction I'm taking rather than challenging and refining it.

cheald··on Getting Docker to not suck for Development
You can just run Docker containers as `--u $UID:$GID`, presuming the docker container isn't set up in such a way that it's hostile to its contents being executed by a non-root user. Usually this just means ensuring that you don't have read/execute permissions locked down to just root and that any in-container directories which need writes have the global write bit set. Once you do that, you can run your containers as whatever user/group you'd like, and things generally just work, and you don't have to worry about building custom images.

    $ cat /etc/lsb-release
    DISTRIB_ID=Ubuntu
    DISTRIB_RELEASE=24.04
    DISTRIB_CODENAME=noble
    DISTRIB_DESCRIPTION="Ubuntu 24.04.1 LTS"
    $ mkdir tmp
    $ docker run --rm -v $(pwd)/tmp:/tmp alpine:latest sh -c 'echo "ok" > /tmp/test.txt'
    $ ll tmp
    .rw-r--r-- root root 3 B Sat Nov 16 14:53:51 2024 test.txt
    $ docker run -u $UID:$GID --rm -v $(pwd)/tmp:/tmp alpine:latest sh -c 'echo "ok" > /tmp/test2.txt'
    $ ll tmp
    .rw-r--r-- root  root  3 B Sat Nov 16 14:53:51 2024 test.txt
    .rw-r--r-- chris chris 3 B Sat Nov 16 14:54:16 2024 test2.txt
cheald··on AI makes tech debt more expensive
The niche I've found for LLMs is for implementing individual functions and unit tests. I'll define an interface and a return (or a test name and expectation) and say "this is what I want this to do", and let the LLM take the first crack at it. Limiting the bounds of the problem to be solved does a pretty good job of at least scaffolding something out that I can then take to completion. I almost never end up taking the LLM's autocompletion at face value, but having it written out to review and tweak does save substantial amounts of time.

The other use case is targeted code review/improvement. "Suggest how I could improve this" fills a niche which is currently filled by linters, but can be more flexible and robust. It has its place.

The fundamental problem with LLMs is that they follow patterns, rather than doing any actual reasoning. This is essentially the observation made by the article; AI coding tools do a great job of following examples, but their usefulness is limited to the degree to which the problem to be solved maps to a followable example.

cheald··on Stargate built 15 years ago in Ohio 50k pounds concrete family time
The first couple of seasons are a little rocky, but once it established its own cast of characters and mythos, it really became something special.
cheald··on LoRA vs. Full Fine-Tuning: An Illusion of Equivalence
I've done a lot of tinkering with the internals of LoRA training, specifically investigating why fine-tune and LoRA training result in such different results, and I'm no academic, but I have found that there are definitely some issues with the SOTA at least WRT Stable Diffusion.

I've had significant success with alternate init mechanisms (the standard technique of init'ing B to zeros really does hurt gradient flow), training alpha as a separate parameter (and especially if you bootstrap the process with alphas learned from a previous run), and altering the per-layer learning rates (because (lr * B) @ (lr @ A) produces an update of a fundamentally different magnitude than the fine-tune update of W * lr = lr * B @ A).

In the context of Stable Diffusion specifically, as well, there's some really pathological stuff that happens when training text encoders alongside the unet; for SD-1.5, the norm of "good" embeddings settles right around 28.0, but the model learns that it can reduce loss by pushing the embeddings away from that value. However, this comes at the cost of de-generalizing your outputs! Adding a second loss term which penalizes the network for drifting away from the L1 norm of the untrained embeddings for a given text substantially reduces the "insanity" tendencies. There's a more complete writeup at https://github.com/kohya-ss/sd-scripts/discussions/294#discu...

You also have the fact that the current SOTA training tools just straight up don't train some layers that fine-tunes do.

I do think there's a huge amount of ground to be gained in diffusion LoRA training, but most of the existing techniques work well enough that people settle for "good enough".

cheald··on PostgreSQL Streaming Replication (WAL); What It Is and How to Configure One
+1 to all of this. The thing I'd add is that we use barman for our additional replicas; WAL streaming is very easy to do with Barman, and we stream to two backups (one onsite, one offsite). The only real costs are bandwidth and disk space, both of which are cheap. Compared to running a full replica (with its RAM costs), it's a very economical way to have a robust disaster recovery plan.

If you're doing manual failover, you don't need an odd number of nodes in the cluster (since you aren't looking for quorum to automatically resolve split-brain like you would be with tools Elasticsearch or redis-sentinel), so for us it's just a question of "how long does it take to get back online if we lose the primary" (answer: as long as it takes to determine that we need to do a switch and invoke repmgr switchover), and "how robust are we against catastrophic failure" (answer: we can recover our DB from a very-close-to-live barman backup from the same DC, or from an offsite DC if the primary DC got hit by an airplane or something).

cheald··on AI Companions Reduce Loneliness
I'm sure that we could find ways to draw the same conclusions from social media and always-on connections in peoples' pockets, but study after study shows that we're lonelier than we've ever been.

Maybe we should just stop trying to replace face-to-face contact with other people.

cheald··on Inertia.js – Build React, Vue, or Svelte apps with server-side routing
Yes, it's all the strings per language, for the SPA. We also have services which we can use to request a bundle of strings from the backend, which I guess could in theory be an avenue for JIT translation on the frontend, but so far the per-language bundle splitting has kept things managable for us.
cheald··on Inertia.js – Build React, Vue, or Svelte apps with server-side routing
We do actually have per-language code splitting with React/Vite. We essentially just have a map of [language, loader], where the loader utilizes lazy imports. Vite is smart enough to split those all out of the main bundle and they get loaded on demand as expected.

    const LOCALE_MAP: any = {
      en: () =>
        document.body.dataset["appEnv"] === "development"
          ? Promise.resolve({ default: {} as LocaleMap })
          : import("@locales/admin.en.compiled.json"),
      fr: () => import("@locales/admin.fr.compiled.json"),
      es: () => import("@locales/admin.es.compiled.json"),
      ...
    };
Then we consume it in a component which accepts the language key, lazy-loads the bundle by just calling `LOCALE_MAP[key]()`, and then uses it as the string translation set for a context function which looks up keys in a dict. Super simple.
cheald··on Chase Bank glitch releases thousands in cash over Labor Day weekend
"Check fraud" isn't exactly a "glitch". This "works" because of a concession made to the customer to allow them to access funds before they're verified, because checks can take weeks to clear.

It's exactly why those "I'll send you a check for $10k, deposit and withdraw it, then send me half back and you keep the other half" scams work. By the time the bank comes back and says "This $10k check was bad, give us back the $10k you withdrew", you're out the $5k you mailed off. Except in this case, the only person you're scamming is yourself, so it's an extra potent form of stupid.

cheald··on Moments in Chromecast's history
The absolutely killer feature of the Chromecast is that I can have guests over (or be visiting someone), and anyone can stream any content they're authorized for. Movies I've bought can be watched anywhere there's a Chromecast; my buddy can come over and we can watch something together with his Paramount+ subscription. Keeping the accounts and authorization linked to personal devices, and letting the Chromecast essentially be a way to translate that to a bigger screen without having to actually stream it out of your pocket is fantastic.
cheald··on Moments in Chromecast's history
Last I looked, there was essentially no good programmatic route into local Nest control, unlike most home automation devices which use wifi/bluetooth/zwave/zigbee. I replaced my Nests with a couple of $25 Centralite Zigbee thermostats and drive it via HomeAssistant running on a Raspberry Pi, and I'm significantly happier with it than I ever was with the Nest.
cheald··on Disrupting deceptive uses of AI by covert influence operations
Given that any of these bad actors can spin up local inference infrastructure to run their operations pretty trivially, what practical effect does this have beyond just "hey, look, we're doing something"?
cheald··on Microsoft Paint's new AI image generator builds on your brushstrokes
It seems like a whole bunch of unnecessarily liability. Once you place yourself in the position of moderator of what people may use their computers for, you are arguably liable for failures of moderation.

To put it another way, nobody's ever sued Canon for making cameras which are used to take illegal photos, but if Canon suddenly started screening your photos to make sure they were acceptable per local laws or whatnot, they suddenly are actually a responsible party in the creation and distribution of whatever it is people use their cameras for.

cheald··on New Windows AI feature records everything you've done on your PC
FWIW, I've gotten in the habit of using Zotero (https://www.zotero.org/) with a browser extension to do this. If I read something that I think I might want to reference later, I just hit an extension button and it gets slurped into Zotero with a bunch of information indexed for retrieval later.
cheald··on LoRA Learns Less and Forgets Less
It's worse than that, because lora requires two matrices per layer. At full rank, you have an additional NxN parameters to learn versus full finetuning, where N is min(input_features, output_features).

For example, tuning a layer of 128 in x 256 out is 32k params. Learning a full-rank lora for that layer would be two matrices of 128x128 and 128x256 = 48k params.

cheald··on OpenAI: Model Spec
"Make an argument for a fact you know to be wrong" isn't an exercise in lying, though. If anything, the ability to explore hypotheticals and thought experiments - even when they are plainly wrong - is closer to a mark of intelligence than the ability to regurgitate orthodoxy.
cheald··on React 19 Beta
AFAIK, the "use" convention is just a convention used to indicate that you're dealing with something that is plugged into React's state management system, and is going to deal with some kind of state change that can cause component re-renders.

It's helpful to remember that functional components are, well, just functions, which on their own don't have any way to preserve their own state without resorting to some kind of global state store. Unlike object instances, which can have members which keep their state across invocations to instance.render(), functional components are just functions which receive props and return something (usually a React.createElement invocation, often disguised with JSX). Since they don't inherently have any way to preserve state, React provides functions to manage state and to trigger re-renders of components. This is what hooks are - they're global functions which manage updating a global state dict behind the scenes, and triggering re-renders when state changes.

* useState declares a stateful variable which is preserved across renders, and a mutator function which is used to update that variable (in the global state dict) and cause React to perform a re-render of any components which depend on that variable.

* useEffect would probably be better named "useSideEffect", since its purpose is to run a callback as a side effect of one of its declared dependencies changing. It is a little overloaded, in that it can run its callback as an effect of: 1. The component initially mounting, 2. the component unmounting, or 3. one of the declared dependencies changing value.

* useMemo is kind of a combination of useState and useEffect, which declares a stateful variable that is preserved across component renders (the memoized value) and uses a generator function to regenerate and save a new version of that variable when a declared dependency changes.

* useCallback can be thought of as a specialized case of useMemo which returns a memoized function.

Essentially all hooks are built on these four (and really, just useState and useEffect, since useMemo and useCallback are trivial derivations from those two).

It's helpful to think about react state as eventual and to treat it as immutable, and to think of components as truly functional constructs, which should just receive external input and return an output. It may be helpful to think of useState values as nothing more than additional values passed to the function by React when it invokes them. When you update state (or run an effect) you aren't changing things in this invocation of the component, you're changing things for the next invocation for the component (and queueing a reinvocation).

Imagine that you have a declared hook `[someState, setSomeState] = useState();`

Calling `setSomeState(val)` doesn't change the value of `someState`, which you should treat as an immutable variable. Instead, it updates the value of `globalStateDict[currentComponentIdentifier]["someState"] = val` and tells React to queue a re-render of this component. React dutifully re-invokes the function, and the second time it's run, the local `someState` variable receives its value from the React-managed global state, which is now `val`.

cheald··on Google Pixel camera consistently blurring out The North Face logo
Well, naturally, but neural nets are pretty good at feature extraction even in the presence of noise. I also just printed out his unblurred example at a scale approximating a real logo and took a photo of it. Came out just fine.
cheald··on Google Pixel camera consistently blurring out The North Face logo
It's an interesting observation, but I couldn't replicate this by taking a photo of a Google Images search for "north face logo", or with a photo of the reddit user's photo. Pixel 7, stock Camera app. The photo turns out as expected.
← PreviousPage 4 of 34Next →