405 karma · joined June 4, 2012
One big differentiator between various Brother models that's not immediately obvious is the tape sizes they accept. The cheaper models don't accept the fatter 24mm/1" labels, gotta step up to the $99ish models for that.
the A connectors with the extra contacts in
between the normal USB 2 ones still aren't enough.
Oh jesus, I hadn't even realized that was a thing!There are only a couple of places where I need the fancier cables that can carry video and have super high xfer speeds. Those don't move around because they're connected to monitors and external hard drives and usually come labeled USB4/Thunderbolt from the factory anyway.
The rest of my USB cables are just rando charging cables... attached to various chargers sticking out of various outlets around my house. No data functionality required and I think iPhone and Android both warn you now if you're plugged into a slow charger.... so I don't really need to differentiate these...
> Re: stiffness: I mean I guess it depends, how big is your bag. I'm thinking like a freezer bag not a grocery bag. I could bend them to fit but then I'd probably mess them up.
A few years back I bought a crapton of 2mil and 4mil clear ziploc style bags in various sizes from amazon. The 6x6" ones are plenty big enough for nearly any modern cable, except for long HDMI/Ethernet. Then those get grouped into big 2x2ft ish 4mil bags.I guess maybe it is a bit of work in an absolute sense. But, it requires basically zero brainpower. I just sorted and bagged them all up when I was watching a game on TV or something. So it didn't really feel like work lmao
Visually, if it's something I'm going to be staring at on a daily basis, I like this much better than having tags hanging off of the cable. Tip: consider white-on-black labels.
The labels have been plenty durable in my experience, but if you need extra durability, you could probably just brush some lacquer over it with a Q-tip.
The sounds hard
Doesn't feel hard to me. I have a bag for USB A, a bag for USB C, and a bag for USB A<->C.My USB-C charging cables are all 3' or longer, so they're visually distinct from the crappy little cables that come for free inside the package for various budget-priced gizmos.
My Thunderbolt/USB4 cables are labeled as such, and I am very confused that you are talking about Thunderbolt/USB4 cables so stiff that they cannot bend to fit into a container or bag.... I mean, mine are stiff but not that stiff. Plus these are expensive cables. Not sure I have any spares. Those are bought for purpose.
I very much recommend them. As a bonus, you can label them if desired. (Print a label out, or slap some tape on there and write on that, or just stick a sheet of paper inside the bag and write on that)
I imagine decompressing from something like Burning Man would be even more intense... even without psychedelics.
(I know psychedelics are obviously a thing at Burning Man, but I don't know if it's 25% or 50% or 75% or 99.99% of people using them...)
As a side note, I know somebody who permanently changed their gender identity after a mushroom trip revelation. I've always wondered how long they waited after the trip to "finalize" their decision.
(FWIW, that was years ago and it seems like it was indeed the right choice for them)
IM/FWIW: I personally have found 4mil ziploc-style bags to be pretty close to indestructible. They're like thick evidence locker bags
Grouping them is key to deduplicating.
It's easy to look at a single legacy USB A-to-B "printer" cable in isolation and think "I might need this someday!" Because you really might need it someday. However, if you group them you might see that you have ten of them. And then you can get rid of... maybe 8 of them.
I also (mostly) put individual cables into baggies. You can get clear 2mil generic ziploc style baggies for super cheap on Amazon or elsewhere. $15 for 200 or something. Prvents tangles and way less effort than wrapping or tying them.
I have not spun up my own homelab, but, I have researched it and the electricity costs can be mitigated to a large extent.
One... the GPUs can be massively clocked down during idle state, to the point where the fans can be shut off as well. The machines themselves can be shut down and wait for a magic wake-on-LAN packet if needed.
Two... even under load, the GPU cores can be significantly underclocked with very little performance loss. The GPU and VRAM/HBM clocks are independent, and the GPU is largely bottlenecked on VRAM/HBM so you can just drop the GPU speed. A commonly reported figure I saw was, basically, 350W nominal cards being underclocked to consume "only" 200W under load with ~5-10% perf loss.
This all assumes you're using discrete GPUs and not AIO systems like a Mac Studio which is going to be pretty efficient just by design; they idle at 35W or so and in practice max out at a few hundred watts. I believe DGX Spark and Strix Halo are similar.
I must stress that this is second-hand anecdata here, admittedly, but my understanding is that it can be pretty manageable.
A $20/month Gemini subscription is truly all you need, then yeah, sure.... obviously a homelab setup is a ridiculous alternative on a pure cost basis. For most people doing "real" work with LLMs 40+ hours per week, a more apt comparison would be one or multiple $200/month subscriptions. At which point the break-even point of a homelab is much sooner.
However, most people running homelabs are doing it for other reasons. Independence, learning, and/or privacy issues.
Wild oversimplification, and benchmarks vary widely, but I've read a lot of benchmarks suggesting that Qwen3.8-27B (xhigh effort) competes with near-frontier models at a lot of coding tasks. To the best of my understanding it's not going to run very feasibly in 16GB of VRAM at usable quants however.
r/LocalLLM and r/LocalLlama are noisy, but valuable sources of anecdata if you have the time (or the tokens, hah) to comb through them. You are going to see a lot of modest setups there, and also guys with $20K+ of hardware.
The two things (besides my bank account) that keep me from investing heavily in local are (1) we are not guaranteed to get a steady release of open models in the future (2) a lot of the "fun" stuff LLM stuff that interests me involves orchestrating lots of parallel agents, which of course multiples the hardware you need to achieve it.
For example, I've been having good results having both Sol and Opus review the same PR, and then I have them cross-review each others' PRs. A next step I'd like to consider is maybe having a swarm of Luna agents review the same PR and have them fight it out... maybe with Sol doing final arbitration? I suspect 5-10 Lunas might outperform a single Opus. Or maybe not. But at any rate, that would be impractical in a homelab without a pretty big hardware (or time) budget.
This is of course anecdata. I know plenty of outliers, too. I know a principal engineer who uses many multiples of the number I quoted above. I am sure we also know many people making do with much much smaller budgets as well, via all kinds of well-discussed methods.
But, "$500-$1500 per month per full-time developer" is just kind of the personal mental baseline I use when making my decisions with regards to thinking about whether any of this makes any economic sense.
The Sparks admittedly are kind of anemic: 273GB/sec is the same bandwidth as a midrange 4060, although (depending on how you configure things) you can effectively have much greater bandwidth by connecting them.
Compared to 1-2 years worth of LLM tokens for a full-time software engineer making $100K+/year, a one-time spend of $12K for 4 Sparks for on-prem private LLM inference starts looking reasonable, particularly if privacy is an important consideration. It starts looking even more reasonable if running something like a private cloud to service multiple developers because then you likely need less hardware per developer.
(Also, it is going to be a long time until RAM+GPU prices return to what we used to call "normal." If ever. I am not endorsing the current state of affairs and I am not saying you wrong to find it insane, but it is definitely the new reality)
I'm usually running about ~4 instances of it locally in separate Git worktrees, which (in our setup) means I'm also running 4 copies of the whole stack: web, frontend, Redis, Postgres, etc. These are rebuilt and torn down multiple times per day. Lot of frontend assets to compile.
Additionally, like many people in corporate environments, we do of course also pay a performance tax thanks to the security software we need to run.
So perf does matter for us.
We could, of course, alternatively... just run cheap shitboxes locally and have all our dev environments in the cloud. Certainly appealing in many ways from a cost perspective and you don't even need to give up DX really.
However, in the next ~5 years, as "actually useful" local LLM inference starts to become genuinely useful (for software engineering, this seems to roughly mean "the ability to run 28B+ models at usable speeds") then having local compute horsepower might start to look more attractive again. Even if it's not an absolute winner on cost/utility alone, the data privacy and cost certainty have value to many.
In general, the reason why modern CPUs are so complex is because the gap in performance between CPU and memory has grown massively over time.
In the old days, something like a 6502 was running nearly synchronously with RAM.
That gap has grown massively over time; a modern CPU is orders of magnitude faster than RAM. So they have to jump through a lot of loops to avoid simply idling 99.99% of the time while they wait for some new data or instructions from RAM. On-CPU cache memory is one answer. Branch prediction and speculative execution are others. As you may imagine, speculatively executing code based on branch prediction is very complex because you must roll back any side effects from that execution if your prediction turns out wrong.
Example:
# assume `i` is a value stored in main memory
if i == 42
j += 1
k -= 1
l = 666
else
z = 123
q = 5879873
Waiting for `i` to arrive from main memory might take thousands of CPU cycles. So instead we will execute one, and possibly both of those branches. But we'll need to undo those side effects if turns out we executed something with an invalid prediction. It's complex, and messy, but still better than sitting around doing nothing for thousands of cycles.That's why we can't just take an R2000 and scale it up to 3ghz. I mean, we could, but it wouldn't work very well unless we also had low-latency 3ghz main memory to pair with it.
But when humanoid enemies behave in plainly stupid ways it's a real immersion-breaker for me. I've been gaming a looooong time, so I'm quite adept at the necessary mental gymnastics to enjoy stuff anyway... but... still... games could be better here. (And by "better" I mean "more fun")
I tried the qwen3.6-27b Q6_k GUFF in llama.cpp
and LM Studio on my M2 MacBook Pro 32GB machine
last week, and I barely get a token a second with either.
The fact that it was this slow makes me suspect it's a matter of insufficient free RAM. The entire model needs to fit into RAM (and stay there the entire time) for acceptable performance.(not sure of exact diagnosis/fix, but definitely look in that direction if you're still having this issue when you give it another shot)
Also, there are two stages - prompt processing, and token generation. Prompt processing is notoriously slow on Apple Silicon unfortunately. If you have large context (which includes system prompts, lots of tools loaded by a harness like Claude Code, OpenCode, etc) it can take minutes for prompt processing before you see the first output token. On the bright side, the tokens are cached between turns, so subsequent turns won't be so bad.
It's highly improbable that the US government has a secret team inside Anthropic and OpenAI manipulating their training regimen.
Two thoughts.One: it would be relatively technically trivial for $GOVERNMENT_AGENCY to just monitor all the prompts + context we send over the wire to OpenAI/Anthropic/etc. That's a goldmine of sensitive personal and corporate data, no secret team needed (although, the LLM providers obviously would need to cooperate)
Two: Rather than secret infiltration teams influencing model training I think what's more likely on the training side of things is simply self-censoring by the LLM providers, so that they don't risk angering the government.
I highly doubt that China has government interlopers, secret or otherwise, inside Qwen's training team. Nonetheless, "sensitive" issues like Tiananmen Square are censored. I would imagine that much/most such censorship in China is self-censorship that doesn't leave a legal/paper trail. That's what we're in danger of seeing (more of) in America IMO.
Open-source model inference providers (who do not have to bear the cost of training) seem able to do it at much lower prices.
https://www.together.ai/pricing
https://fireworks.ai/pricing#serverless-pricing (scroll down to headline models)
Of course, it's possible that they are burning through investor cash as well, and apples-to-apples comparisons are not possible because AFAIK Google does not mention the size/paramcount for 3.5 Flash.
But if the prevailing wisdom is true, I think it's actually encouraging. It suggests that OpenAI and Anthropic could perhaps, if they need to, achieve profitability if they slow down model development and focus on tooling etc. instead. If true that's probably good news for everybody w.r.t. preventing a bursting of this economic bubble.
...my opinions here are of course, conjecture built on top of conjecture....
You can get part of the way there by Sublime Text to run queries and browse results. In Sublime, it's easy - you can create a custom "build tool" with (essentially) just one line of config... you're basically just piping the current file's contents to command-line PSQL: https://www.youtube.com/watch?v=tPd4m3PLVqU
Then I use the SublimeAllAutocomplete package. It works just like "regular" Sublime autocomplete except it gives you autocomplete suggestions from all open files, not just the current one - so if you have your DB schema dump in another window it will use that: https://github.com/alienhard/SublimeAllAutocomplete
Obviously, that doesn't really give you the smartest autocomplete ever but it's pretty productive for me. Tons of room for improvement though.
I actually enjoy early-morning coding (before the rest of the world wakes up) more than late-night coding now!
What I'm saying, though, is that the people around them at the bar are far more likely to be able to assist than a conference organizer who's not physically there.
"The point is that conferences should have a policy for how to deal with these kinds of situations. Those policies are probably going to depend on a certain amount of judgement from the conference organizer."
I totally agree that there need to be policies. But policies are just words, and they won't get the job done alone.
What will get the job done is looking out for each other. In social situations, particularly bars, we should all be looking out for each other, especially women, since they're disproportionally the ones on the receiving end of predatory/harassing behavior.
If she went to the bar with a group, every single one of them ought to have been looking out for each other. And yes, I've been in scores of similar situations where I and others have looked out for the welfare of others. You don't need to be an obnoxious "white knight" about it; you don't even need to be overt. Creeps tend to look for girls whose friends aren't paying attention.
Example: Your female friend is caught in a conversation with a potential El Creepo. Play dumb and introduce yourself to the guy in a purely friendly way. Heck, maybe buy a round of drinks. Just knowing that somebody noticed him will often nip things in the bud.
On one hand, running a conference/convention is difficult enough without being asked to referee personal conflicts. It almost inevitably devolves into a he-said/she-said kind of argument.
But on the other hand, if we take that kind of an attitude, that's essentially a signal for predatory men to go ahead and harass women (or worse). That's pretty much what predatory men have been doing since the beginning of history - acting with impunity since claims of rape or harassment are almost impossible to prove if there are no additional witnesses or physical evidence.
What we've done is to stress proper conduct before the event, and if there are repeat complaints about somebody they're removed from the community permanently.
It's not a perfect process and we've given some "second chances" to people that we've later regretted.