14,356 karma · joined August 19, 2009
[ my public key: https://keybase.io/cheald; my proof: https://keybase.io/cheald/sigs/piAgm06dhM9eFLAHy4jVtO3rIY5-emPB4MRJ91EKOZU ]
People take the path of least resistance; we know this. It's why, for the longest time, people used one password for everything. People don't like using password managers, either, but we would all agree that it's unacceptably insecure to not use them, because the alternative is "one password used everywhere, maybe with a single varying digit on the end".
The time-variant component is still quite valuable, but it does nothing to protect you in the event of a password manager compromise. This is not a hypothetical; LastPass has suffered multiple breaches, and the more popular a solution, the more likely there are to be attacks against that solution. By keeping your 2FA separate from your password manager, even if it's still just "something you know", it's something you know in a location that's orthogonal to your passwords. If I yield to convenience and use a 2FA desktop app, then now, instead of just attacking my Bitwarden install, you have to successfully attack my Bitwarden install and my 2FA desktop app install to get access to my accounts, and the combination of password managers * 2FA managers is a substantially larger attack surface and requires a significantly more sophisticated attack to get both pieces.
The arguments in the article come down to "well, 2FA mitigates phishing attacks" (true) and "Google Authenticator means you can lose your data easily" (also true). But neither of these is a good argument for why the data should be kept together. It just means "use 2FA", and "use a 2FA manager that lets you directly manage your seeds and keep offsite encrypted backups".
If you can't be bothered to do it properly, then 2FA codes in your password manager is certainly better than not using 2FA at all, but that just makes it a less terrible solution, not a good one.
My loss metrics stay roughly the same (they're slightly lower, but SD loss is fraught to interpret because variance by timestep renders it more or less meaningless), but tracking the means of `param.grad.norm / param.numel` (which shows how big the grad updates are) shows the grads stabilizing significantly quicker than baseline. I'm tracking suppressed params / total params via tensorboard, and I show that it drops (as expected) but then stabilizes at around 7%, suggesting that there are model parameters which consistently don't agree. I'm gonna try tracking the variance from the mean, as well, and perhaps down-weight or eliminate grads for parameters which show high cos similarity variance over time (suggesting a generalized lack of agreement in the direction to move, further suggesting that the parameter cannot contribute meaningfully to the task).
It often gets the actual math wrong, but it is good enough at connecting the dots between my layman's intuition and the "right answer" that I can get myself over humps that I'd previously have been hopelessly stuck on.
It does make those mistakes you're talking about very frequently, but once I'm told that the thing I'm trying to do is achievable with the Gram-Schmidt process, I can go self-educate on that further.
The big thing I've had to watch out for is that it'll usually agree that my approach is a good or valid one, even when it turns out not to be. I've learned to ask my questions in the shape of "how do I", rather than "what if I..." or "is it a good idea to...", because most of the time it'll twist itself into shapes to affirm the direction I'm taking rather than challenging and refining it.
$ cat /etc/lsb-release
DISTRIB_ID=Ubuntu
DISTRIB_RELEASE=24.04
DISTRIB_CODENAME=noble
DISTRIB_DESCRIPTION="Ubuntu 24.04.1 LTS"
$ mkdir tmp
$ docker run --rm -v $(pwd)/tmp:/tmp alpine:latest sh -c 'echo "ok" > /tmp/test.txt'
$ ll tmp
.rw-r--r-- root root 3 B Sat Nov 16 14:53:51 2024 test.txt
$ docker run -u $UID:$GID --rm -v $(pwd)/tmp:/tmp alpine:latest sh -c 'echo "ok" > /tmp/test2.txt'
$ ll tmp
.rw-r--r-- root root 3 B Sat Nov 16 14:53:51 2024 test.txt
.rw-r--r-- chris chris 3 B Sat Nov 16 14:54:16 2024 test2.txtThe other use case is targeted code review/improvement. "Suggest how I could improve this" fills a niche which is currently filled by linters, but can be more flexible and robust. It has its place.
The fundamental problem with LLMs is that they follow patterns, rather than doing any actual reasoning. This is essentially the observation made by the article; AI coding tools do a great job of following examples, but their usefulness is limited to the degree to which the problem to be solved maps to a followable example.
I've had significant success with alternate init mechanisms (the standard technique of init'ing B to zeros really does hurt gradient flow), training alpha as a separate parameter (and especially if you bootstrap the process with alphas learned from a previous run), and altering the per-layer learning rates (because (lr * B) @ (lr @ A) produces an update of a fundamentally different magnitude than the fine-tune update of W * lr = lr * B @ A).
In the context of Stable Diffusion specifically, as well, there's some really pathological stuff that happens when training text encoders alongside the unet; for SD-1.5, the norm of "good" embeddings settles right around 28.0, but the model learns that it can reduce loss by pushing the embeddings away from that value. However, this comes at the cost of de-generalizing your outputs! Adding a second loss term which penalizes the network for drifting away from the L1 norm of the untrained embeddings for a given text substantially reduces the "insanity" tendencies. There's a more complete writeup at https://github.com/kohya-ss/sd-scripts/discussions/294#discu...
You also have the fact that the current SOTA training tools just straight up don't train some layers that fine-tunes do.
I do think there's a huge amount of ground to be gained in diffusion LoRA training, but most of the existing techniques work well enough that people settle for "good enough".
If you're doing manual failover, you don't need an odd number of nodes in the cluster (since you aren't looking for quorum to automatically resolve split-brain like you would be with tools Elasticsearch or redis-sentinel), so for us it's just a question of "how long does it take to get back online if we lose the primary" (answer: as long as it takes to determine that we need to do a switch and invoke repmgr switchover), and "how robust are we against catastrophic failure" (answer: we can recover our DB from a very-close-to-live barman backup from the same DC, or from an offsite DC if the primary DC got hit by an airplane or something).
Maybe we should just stop trying to replace face-to-face contact with other people.
const LOCALE_MAP: any = {
en: () =>
document.body.dataset["appEnv"] === "development"
? Promise.resolve({ default: {} as LocaleMap })
: import("@locales/admin.en.compiled.json"),
fr: () => import("@locales/admin.fr.compiled.json"),
es: () => import("@locales/admin.es.compiled.json"),
...
};
Then we consume it in a component which accepts the language key, lazy-loads the bundle by just calling `LOCALE_MAP[key]()`, and then uses it as the string translation set for a context function which looks up keys in a dict. Super simple.It's exactly why those "I'll send you a check for $10k, deposit and withdraw it, then send me half back and you keep the other half" scams work. By the time the bank comes back and says "This $10k check was bad, give us back the $10k you withdrew", you're out the $5k you mailed off. Except in this case, the only person you're scamming is yourself, so it's an extra potent form of stupid.
To put it another way, nobody's ever sued Canon for making cameras which are used to take illegal photos, but if Canon suddenly started screening your photos to make sure they were acceptable per local laws or whatnot, they suddenly are actually a responsible party in the creation and distribution of whatever it is people use their cameras for.
For example, tuning a layer of 128 in x 256 out is 32k params. Learning a full-rank lora for that layer would be two matrices of 128x128 and 128x256 = 48k params.
It's helpful to remember that functional components are, well, just functions, which on their own don't have any way to preserve their own state without resorting to some kind of global state store. Unlike object instances, which can have members which keep their state across invocations to instance.render(), functional components are just functions which receive props and return something (usually a React.createElement invocation, often disguised with JSX). Since they don't inherently have any way to preserve state, React provides functions to manage state and to trigger re-renders of components. This is what hooks are - they're global functions which manage updating a global state dict behind the scenes, and triggering re-renders when state changes.
* useState declares a stateful variable which is preserved across renders, and a mutator function which is used to update that variable (in the global state dict) and cause React to perform a re-render of any components which depend on that variable.
* useEffect would probably be better named "useSideEffect", since its purpose is to run a callback as a side effect of one of its declared dependencies changing. It is a little overloaded, in that it can run its callback as an effect of: 1. The component initially mounting, 2. the component unmounting, or 3. one of the declared dependencies changing value.
* useMemo is kind of a combination of useState and useEffect, which declares a stateful variable that is preserved across component renders (the memoized value) and uses a generator function to regenerate and save a new version of that variable when a declared dependency changes.
* useCallback can be thought of as a specialized case of useMemo which returns a memoized function.
Essentially all hooks are built on these four (and really, just useState and useEffect, since useMemo and useCallback are trivial derivations from those two).
It's helpful to think about react state as eventual and to treat it as immutable, and to think of components as truly functional constructs, which should just receive external input and return an output. It may be helpful to think of useState values as nothing more than additional values passed to the function by React when it invokes them. When you update state (or run an effect) you aren't changing things in this invocation of the component, you're changing things for the next invocation for the component (and queueing a reinvocation).
Imagine that you have a declared hook `[someState, setSomeState] = useState();`
Calling `setSomeState(val)` doesn't change the value of `someState`, which you should treat as an immutable variable. Instead, it updates the value of `globalStateDict[currentComponentIdentifier]["someState"] = val` and tells React to queue a re-render of this component. React dutifully re-invokes the function, and the second time it's run, the local `someState` variable receives its value from the React-managed global state, which is now `val`.