5,621 karma · joined January 25, 2014
So google's tap to pay doesn't work, but others do, like garmin pay. Random bank apps are hit and miss.
Or the Mesh Node T114 Development Board V2 (23mm wide).
Sad that most of ham radio, which has similar APRS for tracking, doesn't do async or multi-hop.
Beware that they take much longer to start since they don't know the GPS coords of the nearest cell tower, nor do they download the GPS almanac over the cell connection. They will however remember the almanac so you can leave it on for 30 minutes and not have to redownload it the next day.
Public hearing set for nearly 650 diesel generators tied to Amazon data center.
2 days ago, Inside the Secretive Deal for a $10 Billion Data Center in ... 21 buildings and roughly 600 diesel generators"
6 days ago "Public hearing set for nearly 650 diesel generators tied to Amazon data center"
4 hours ago, Trump’s Vision for A.I. Dominance Comes With Major Air Pollution ... The first 366 megawatts would come from a power plant made up of more than 800 small gas-burning generators bolted to concrete pads.
Trick is money can influence local governments, even without a clear benefit to the local residents. Then result is people end up with more expensive water, less water, more expensive power, and crazy amounts of noise. Similarly DCs often seem to get special deals from the power company, which raises costs to all users.
Similarly they seem to get a pass on permits: Elon Musk’s artificial intelligence company xAI has installed 59 natural gas turbines for its Colossus 2 data center project in Tennessee without securing federal clean air permits, according to communications between regulators and xAI representatives.
I don't think most people would care if DCs could be hidden by a line of trees, sadly that's not the case.
Not to mention consuming quite a bit of water to provide cheaper cooling, instead of using closed loop systems.
Or just get an AMD Ryzen.
Go has less flexibility, no pointer arithmetic, a healthy package system, and a smaller domain. Mostly consuming or providing network services. My favorite feature is channels. For me they make levering the performance of multi-core CPUs straight forward, and dramatically nicer than the C approaches I've tried like pthreads and mutexes.
I wouldn't rate go as secure as rust, but has a pleasingly developer friendly approach. Seems way more secure than C.
Making a pipeline where each stage is 1 to N threads is pleasingly easy, reliable, and performant.
However the motherboard vendors get annoyingly hide that from you by claiming DDR4 is dual channel (2 x 64 bit which means two outstanding cache misses, one per channel) and just glossing over the difference by saying DDR5 dual channel (4 x 32 bit which means 4 outstanding cache misses).
> Yes, because AMD64 has 64-bit words.
It's a bit more complicate than that. First you have 3 levels of cache, the last of which triggers a cache line load, which is 64 bytes (not 64 bits). That goes to one of the 4 channels, there's a long latency for the first 64 bits. Then there's the complications of opening the row, which makes the columns available, which can speed up things if you need more than one row. But the general idea is that you get at the maximum one cache line per channel after waiting for the memory latency.
So DDR4 on a 128 bit system can have 2 cache lines in flight. So 128 bytes * memory latency. On a DDR5 system you can have 4 cache lines in flight per memory latency. Sure you need the bandwidth and 32 bit channels have half the bandwidth per clock, but the trick is the memory bus spends most of it's time waiting on memory to start a transfer. So waiting 50ns then getting 32bit @ 8000 MT/sec isn't that different than waiting 50ns and getting 64 bit @ 8000MT/sec.
Each 32 bit subchannel can handle a unique address, which is turned into a row/column, and a separate transfer when done. So a normal DDR5 system can look up 4 addresses in parallel, wait for the memory latency and return a cache line of 64 bytes.
Even better when you have something like strix halo that actually has a 256 bit wide memory system (twice any normal tablet, laptop, or desktop), but also has 16 channels x 16 bit, so it can handle 16 cache misses in flight. I suspect this is mostly to get it's aggressive iGPU fed.
So with 2 dimms or 4 you still get 128 bit wide memory. With DDR4 that means 2 channels x 64 bit each. With DDR5 that means 4 channels x 32 bit each.
Keep in mind that memory controller is in the CPU, which is where the DDR4/5 memory controller is. The motherboards job is to connect the right pins on the DIMMs to the right pins on the CPU socket. The days of a off chip memory controller/north bridge are long gone.
So if you look at an AM5 CPU it clearly states:
* Memory Type: DDR5-only (no DDR4 compatibility).
* Channels: 2 Channel (Dual-Channel).
* Memory Width: 2x32-bit sub-channels (128-bit total for 2 sticks).I don't think you can wire up DDR5 dimms willy nilly as if they were 2 separate 32 bit dimms.
For some workloads the extra channels help, despite having the same bandwidth. This is one of the reasons that it's possible for a DDR5 system to be slightly faster than a DDR4 system, even if the memory runs at the same speed.
Yes, but tablets, laptops, and normal (non-HEDT) desktops have 4 channels, 4x32 bit = 128 bit wide. Modern memory with DDR5 allows two 32 bit channels on a 64 bit dimm. The previous gen DDR4 would allow 1 64 bit channel on a 64 bit dimm.
So strix halo (on laptops, tablets, and desktops) allows for a 256 bit wide memory system, providing twice the memory bandwidth of any ryzen or intel i3/i5/i7/i9. The Apple pro (256 bit), max (512 bit), and ultra (1024 bit) lines of apple silicon have greater than 128 bit wide memory systems. On the AMD size it's just the Threadripper (256 bit) and Threadripper pro (512 bit), but those are typically in expensive workstations that are physically large, expensive, and need substantial cooling.
So the HALO is pretty unique (outside of Apple) for providing twice the memory bandwidth of anything else that fits in the tablet, laptop, or small desktop category.
Sure modules just play with env variables. But it's easy to inspect (module show), easy to document "use modules load ...", allows admins to change the default when things improve/bug fixed, but also allows users to pin the version. It's very transparent, very discover-able, and very "stale". Research needs dictate that you can reproduce research from years past. It's much easier to look at your output file and see the exact version of compiler, MPI stack, libraries, and application than trying to dig into a container build file or similar. Not to mention it's crazy more efficient to look at a few lines of output than to keep the container around.
As for slurm, I find it quite useful. Your main complaint is no default systemd service files? Not like it's hard to setup systemd and dependencies. Slurms job is scheduling, which involves matching job requests for resources, deciding who to run, and where to run it. It does that well and runs jobs efficiently. Cgroup v2, pinning tasks to the CPU it needs, placing jobs on CPU closest to the GPU it's using, etc. When combined with PMIX2 it allows impressive launch speeds across large clusters. I guess if your biggest complaint is the systemd service files that's actually high praise. You did mention logging, I find it pretty good, you can increase the verbosity and focus on server (slurmctld) or client side (slurmd) and enable turning on just what you are interested, like say +backfill. I've gotten pretty deep into the weeds and basically everything slurm does can be logged, if you ask for it.
Sounds like you've used some poorly run clusters, I don't doubt it, but I wouldn't assume that's HPC in general. I've built HPC clusters and did not use the university's AD, specifically because it wasn't reliable enough. IMO a cluster should continue to schedule and run jobs, even if the uplink is down. Running a past EoL OS on an HPC cluster is definitely a sign that it's not run well and seems common when a heroic student ends up managing a cluster and then graduates leaving the cluster unmanaged. Sadly it's pretty common for IT to run a HPC cluster poorly, it's really a different set of contraints, thus the need for a HPC group.
Plenty of HPC clusters out there a happy to support the tools that helps their users get the most research done.
You can get tablets, laptops, and desktops. I think windows is more limited and might require static allocation of video memory, not because it's a separate pool, just because windows isn't as flexible.
With linux you can just select the lowest number in bios (usually 256 or 512MB) then let linux balance the needs of the CPU/GPU. So you could easily run a model that requires 96GB or more.
Or strix halo.
Seems rather over simplified.
The different levels of quants, for Qwen3.6 it's 10GB to 38.5GB.
Qwen supports a context length of 262,144 natively, but can be extended to 1,010,000 and of course the context length can always be shortened.
Just use one of the calculators and you'll get much more useful number.
However I still find it crazy that when you slow down the laser and one photon at a time goes through either slit you still get the bands. Which begs the question, what exactly is it constructively or destructively interfering with?
Still seems like there's much to be learned about the quantum world, gravity, and things like dark energy vs MOND.
Or use NVIS, which at least makes triangulation harder.
Don't forget to setup sound, preferably in stereo.
She used the hell out of it for years. One time on a call she was fascinated by fish. I printed out one of her drawings remotely while we were on the phone and she loved it.
Middle click to open new tabs is compatible with middle click to paste.