HNHacker News
TopNewBestAskShowJobs

brucethemoose2

7,874 karma · joined February 13, 2023

submissionscomments
brucethemoose2··on The Reasonable Effectiveness of Using Old Phones as Servers
Some phones have limiters to keep the battery at 60%-80%.

I believe most can do this with the right software.

It's not a fix, but it should extend the life considerably.

brucethemoose2··on The Reasonable Effectiveness of Using Old Phones as Servers
Another benefit: modern smartphones have large GPUs, large media blocks, and fast RAM.

With the right software, they can be a surprisingly powerful AI host or transcoding server.

brucethemoose2··on Towards 1-bit Machine Learning Models
Real world GPU performance is hugely influenced by hand optimization of the CUDA kernels.
brucethemoose2··on Mazda’s rotary engine in the age of the electric car
Power/Weight is extremely high. A tiny wankel will do the job, and weight is everything on cars.

It does prefer a narrow RPM band, which is fine.

Reliability is the biggest concern TBH, but maybe that's not a huge bummer if its more of a backup/assistant engine.

brucethemoose2··on DBRX: A new open LLM
Yeah, its an unspoken but rampant thing in the llm community. Basically no one respects licenses for training data.

I'd say the majority of instruct tunes, for instance, use OpenAI output (which is against their TOS).

But its all just research! So who cares! Or at least, that seems to be the mood.

brucethemoose2··on DBRX: A new open LLM
Yeah I know, hence its odd I found it kind of dumb for personal use. Moreso with the smaller models, which lost an objective benchmark I have to some Mistral finetunes.

And I don't think I was using it wrong. I know, for instance, the Chinese language models are funny about sampling since I run Yi all the time.

brucethemoose2··on DBRX: A new open LLM
I would note the actual leading models right now (IMO) are:

- Miqu 70B (General Chat)

- Deepseed 33B (Coding)

- Yi 34B (for chat over 32K context)

And of course, there are finetunes of all these.

And there are some others in the 34B-70B range I have not tried (and some I have tried, like Qwen, which I was not impressed with).

Point being that Llama 70B, Mixtral and Grok as seen in the charts are not what I would call SOTA (though mixtral is excellent for the batch size 1 speed)

brucethemoose2··on JPEG XL exceeds 20% of served images
The conspiracy theorist in me says thats a low priority due to perverse incentives (namely selling more storage at a huge markup).

Another rationale is that the Apple ecosystems tends to not use JPEG by default anyway, right? It uses HEIC or something.

brucethemoose2··on Show HN: Memories – FOSS Google Photos alternative built for high performance
Being a "hero" open source dev for a project like that can require a lot of neuroticism.

Sometimes it works, but sometimes the project is just too big, I think.

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
It's not either or, you can use different vendors for different tasks.

tinygrad isn't in the realm of production ready though, AFAIK.

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
The MI300 is the best accelerator you can buy, for many current workloads.

It's technically way more advanced. Not as outrageously priced as an H100 either.

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
I think you are preaching to the choir, and AMD is not listening.

AMD would be selling 48GB 7900s or AI-only W7900s if they really wanted a consumer card ramp.

They don't. Not because they can't (they literally prevent OEMs from doing so, who would double up VRAM in a heartbeat without AMD lifting a finger), but because AMD doesn't want that.

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
I never followed Hotz, so perhaps I missed something cool. But I never understood the hype myself.
brucethemoose2··on Key Stable Diffusion Researchers Leave Stability AI as Company Flounders
Well, personally, SDXL just blows 1.5 out of the water for me. I haven't had a reason to even touch 1.5 in months.

But note that SDXL is really awful in automatic1111 or vanilla HF diffusers for me. You have to use something with proper augmentations (like ComfyUI or Fooocus(which runs on ComfyUI)).

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
Used 3090 prices are absolutely outrageous.

And the 4090 MSRP was outrageous to begin with.

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
I was talking about renting!

There are some boutique hosts like Hot Aisle serving MI300s (who I really should reach out to), but for the immediate future our little startup is stuck with the big cloud providers. No MI300s for us mere mortals, not even to rent.

brucethemoose2··on The AMD tinybox is on hold until we can build and run the firmware on our GPUs
But is this going to blow over in a few days? Again?

I can certainly appreciate frustration with the AMD stack, but be blunt, I was not impressed with Hotz's YouTube rant from before.[1] It didn't give the impression of a stable framework, and this doesn't either.

Also (at least from the end user llm inference side of things) ROCm is not nearly as unusable as it used to be. We would certainly be renting MI300s over A100s (or even H100s) if we could get any, and we use a number of different inference backends.

1: https://news.ycombinator.com/item?id=36193625

brucethemoose2··on Key Stable Diffusion Researchers Leave Stability AI as Company Flounders
SDXL is amazing.

The community is entrechend in 1.5 because that's what everyone is now familiar with, IMO

brucethemoose2··on The rise and fall of a Halifax man's illegal TV streaming empire
Its more like the store being a literal hedge maze, with an entrance fee, and once you get to the actual products, they are outrageously priced junk.

And the store is price fixing with nearby stores.

I think theres a difference between opportunistic crime, and reasonably upstanding people being utterly frustrated with a market status quo.

If course there is a ton of piracy that is just straight up theft from perfectly convenient platforms, but the unacceptable, anticompetitive commercial platforms are the core that keeps the community going, I think.

brucethemoose2··on Show HN: Not sure you're talking to a human? Create a human check
My "oh no" moment was a vision model reading this perfectly:

https://abadguide.files.wordpress.com/2012/01/jh66.jpg?w=640

Not an OCR program or anything specialized, just some generic (vision) llm that can run on my desktop with a dumb prompt...

brucethemoose2··on Grok
Tests are not out yet, but:

- It's very large, yes.

- It's a base model, so its not really practical to use without further finetuning.

- Based on Grok-1 API performance (which itself is probably a finetune) its... not great at all.

brucethemoose2··on I Put 4M Suns in a Black Hole over New York [video]
Going to leave this gem here:

https://www.vttoth.com/CMS/physics-notes/311-hawking-radiati...

Black holes are weird because they are essentially macroscopic particles with only one variable, mass (ignoring angular momentum, charge and other details for the moment). But they scale very strangely.

For instance, its interesting to see what a black hole is like that emits precisely 100W of radiation. Or that is as big as a apple, or weighs as much as a apple. Or one that lives exactly 100 years, or that is as dense as water (as black holes have the odd property of getting less 'dense' as their mass grows). Its very unintuitive, and just punching values into the calculator illustrates it perfectly.

I bring this up because the illustration (a town sized black hole sitting over New York) is a very specific configuration, and it doesn't really illustrate how oddly these objects scale.

brucethemoose2··on AI chip startup Groq grabs the spotlight
Groq's inference strategy appears to be "SRAM only." There is no external memory, like GGDR or HBM. Instead, large models are split between networked cards, and the inputs/outputs and pipelined.

This is a great idea... In theory. But it seems like the implementation (IMO) missed the mark.

They are using reticle size dies, running at high TDPs, at 1 die per card, with long wires running the interconnect.

A recent microsoft paper proposed a similar strategy, but with much more economical engineering. Instead, much smaller, cheaper SRAM heavy chips would be tiled across a motherboard, with no need for a power hungry long-range interconnect, no expensive dies on expensive PCIe cards. The interconnect is physically so much shorter and lower power by virtue of being on a motherboard.

In other words, I feel that Groq took an interesting inference strategy and ignored a big part of what makes it cool, packaging them like PCIe GPUs instead of tiled accelerators. Combined with the node disadvantage and compatibility disadvantage, I'm not sure how they can avoid falling into obscurity like Graphcore, which took a very similar SRAM heavy approach.

brucethemoose2··on U.S. is investigating Meta for role in drug sales
> Proper moderation

I would point to oldschool forums (and HN!), where the communities were just large enough to moderate themselves and stop nasty off topic junk like that from appearing.

One problem is modern social media has (mostly) yanked this self-moderation power from users. And, as you said, they are too big and too cheap to hire enough mods on payroll to replace them.

Another, that I have, is that they make money from the drug dealer posts! If Facebook has to leave it up, OK, but they sure as heck shouldn't display ads or collect metrics. That should be a far more zealous check.

brucethemoose2··on Ollama now supports AMD graphics cards
Interestingly, Ollama is not popular at all in the "localllama" community (which also extends to related discords and repos).

And I think thats because of capabilities... Ollama is somewhat restrictive compared to other frontends. I have a littany of reasons I personally wouldn't run it over exui or koboldcpp, both for performance and output quality.

This is a necessity of being stable and one-click though.

brucethemoose2··on IBM CEO pay jumps 23% in 2023, average employee gets 7%
I mean, they have so much potential...

As an example, they have expertise designing huge silicon chips, big iron servers, and fast interconnects. They even made a ternery chip in the past, and plenty of fabrication research. They could absolutely be a horse in the AI accelerator race.

Thats just one example of many.

...But they dont take advantage of any of that, like they are stuck in a corporate quagmire or something.

brucethemoose2··on 4T transistors, one giant chip (Cerebras WSE-3) [video]
As I understand it, the WSE-2's interconnect is actually quite good, and models are split across chips kinda like GPUs.

And keep in mind that these nodes are hilariously "fat" compared to a GPU node (or even an 8x GPU node), meaning less congestion and overhead from the topology.

brucethemoose2··on 4T transistors, one giant chip (Cerebras WSE-3) [video]
And even mundane details are so difficult... thermal expansion, the sheer number of pins that have to line up. This thing is a marvel.
brucethemoose2··on 4T transistors, one giant chip (Cerebras WSE-3) [video]
Reposting the CS-2 teardown in case anyone missed it. The thermal and electrical engineering is absolutely nuts:

https://vimeo.com/853557623

https://web.archive.org/web/20230812020202/https://www.youtu...

(Vimeo/Archive because the original video was taken down from YouTube)

brucethemoose2··on Building Meta's GenAI infrastructure
Facebook very specifically bought and customized Intel SKUs tailored for AI workloads for some time.
Page 1 of 34Next →