HNHacker News
TopNewBestAskShowJobs

bick_nyers

829 karma · joined August 31, 2018

submissionscomments
bick_nyers··on Why DeepSeek is cheap at scale but expensive to run locally
Or merge the bottom 1/8 (or whatever) experts together and (optionally) do some minimal training with all other weights frozen. Would need to modify the MoE routers slightly to map old -> new expert indices so you don't need to retrain the routers.
bick_nyers··on Why DeepSeek is cheap at scale but expensive to run locally
The general rule of thumb when assessing MoE <-> Dense model intelligence is SQRT(Total_Params*Active_Params). For Deepseek, you end up with ~158B params. The economics of batch inferencing a ~158B model at scale are different when compared to something like Deepseek (it is ~4x more FLOPS per inference after all), particularly if users care about latency.
bick_nyers··on Why DeepSeek is cheap at scale but expensive to run locally
There's still a lot of opportunity for software optimizations here. Trouble is that really only two classes of systems get optimizations for Deepseek, namely 1 small GPU + a lot of RAM (ktransformers) and the system that has all the VRAM in the world.

A system with say 192GB VRAM and rest standard memory (DGX station, 2xRTX Pro 6000, 4xB60 Dual, etc.) could still in theory run Deepseek @4bit quite quickly because of the power law type usage of the experts.

If you aren't prompting Deepseek in Chinese, a lot of the experts don't activate.

This would be an easier job for pruning, but still I think enthusiast systems are going to trend in a way the next couple years that makes these types of software optimizations useful on a much larger scale.

There's a user on Reddit with a 16x 3090 system (PCIE 3.0 x4 interconnect which doesn't seem to be using full bandwidth during tensor parallelism) that gets 7 token/s in llama.cpp. A single 3090 has enough VRAM bandwidth to scan over its 24GB of memory 39 times per second, so there's something else going on limiting performance.

bick_nyers··on Zed: High-performance AI Code Editor
I've been using PyCharm for the debugger (and everything else) and VSCode + RooCode + Local LLM lately.

I've heard decent things about the Windsurf extension in PyCharm, but not being able to use a local LLM is an absolute non-starter for me.

bick_nyers··on Nvidia's latest AI PC boxes sound great – for data scientists with $3k to spare
MoE inference wouldn't be terrible. That being said, there's not a good MoE model in the 70-160B range as far as I'm aware.
bick_nyers··on Apple M3 Ultra
If you want to split tensorwise yes. Layerwise splits could go over Ethernet.

I would be interested to see how feasible hybrid approaches would be, e.g. connect each pair up directly via ConnectX and then connect the sets together via Ethernet.

bick_nyers··on Apple M3 Ultra
About $12k when Project Digits comes out.
bick_nyers··on Apple M3 Ultra
Just to add onto this point, you expect different experts to be activated for every token, so not having all of the weights in fast memory can still be quite slow as you need to load/unload memory every token.
bick_nyers··on Video encoding requires using your eyes
It's not really possible to say what's "best" because the criteria is super subjective.

I personally like the Spline family, and I default to Spline36 for both upscaling and downscaling in ffmpeg. Most people can't tell the difference between Spline36 and Lanczos3. If you want more sharpness, go for Spline64, for less sharpness, try Spline16.

Edit: As far as I'm aware though OpenCV doesn't have Spline as an option for resizing.

bick_nyers··on How to Run DeepSeek R1 671B Locally on a $2000 EPYC Server
It would not be that slow as it is an MoE model with 37b activated parameters.

Still, 8x3090 gives you ~2.25 bits per weight, which is not a healthy quantization. Doing bifurcation to get up to 16x3090 would be necessary for lightning fast inference with 4bit quants.

At that point though it becomes very hard to build a system due to PCIE lanes, signal integrity, the volume of space you require, the heat generated, and the power requirements.

This is the advantage of moving up to Quadro cards, half the power for 2-4x the VRAM (top end Blackwell Quadro expected to be 96GB).

bick_nyers··on How to Run DeepSeek R1 671B Locally on a $2000 EPYC Server
It will be slower for a 70b model since Deepseek is an MoE that only activates 37b at a time. That's what makes CPU inference remotely feasible here.
bick_nyers··on Promising results from DeepSeek R1 for code
An actual hardcore technical AI "psychology" program would actually be really cool. Could be a good onboarding for prompt engineering (if it still exists in 5 years).
bick_nyers··on DeepSeek: X2 Speed for WASM with SIMD
I definitely agree with you in the interim regarding junior developers. However, I do think we will eventually have the AI coding equivalent of CICD built into perhaps our IDE. Basically, when an AI generated some code to implement something, you chain out more AI queries to test it, modify it, check it for security vulnerabilities etc.

Now, the first response some folks may have is, how can you trust that the AI is good at security? Well, in this example, it only needs to be better than the junior developers at security to provide them with benefits/learning opportunities. We need to remember that the junior developers of today can also just as easily write insecure code.

bick_nyers··on Nvidia announces next-gen RTX 5090 and RTX 5080 GPUs
Check out their project digits announcement, 128GB unified memory with infiniband capabilities for $3k.

For more of the fast VRAM you would be in Quadro territory.

bick_nyers··on Does current AI represent a dead end?
I suspect the big AI companies try to adversarially train that out as it could be used to "jailbreak" their AI.

I wonder though, what would be considered a meaningful punishment/reward to an AI agent? More/less training compute? Web search rate limits? That assumes that what the AI "wants" is to increase its own intelligence.

bick_nyers··on Training LLMs to Reason in a Continuous Latent Space
I wonder if you would want to use an earlier layer as opposed to the penultimate layer, I would imagine that the LLM uses that layer to "prepare" for the final dimensionality reduction to clean the signal such that it scores well on the loss function.
bick_nyers··on The Need to Grind Concrete Examples Before Jumping Up a Level of Abstraction
I kinda wish you could just take a course on a specific distribution. Like, here's the Poisson class where you learn all of its interesting properties and apply it to e.g. queuing problems.
bick_nyers··on Ask HN: How much would you pay for local LLMs?
If you are comfortable with purchasing used hardware, used 3090 are great value, they can be had for roughly a third of the price of a new 4090.

How many GPUs you need is completely dependent on the size of your team, their frequency of usage, and the size of the models you are comfortable with.

I generally recommend you rent instances on something like runpod to build out a good estimate of your actual usage before commiting a bunch of money to hardware.

bick_nyers··on OpenCoder: Open Cookbook for Top-Tier Code Large Language Models
I would probably refer to category 1 as "Open Architecture". I wouldn't want to give anyone the false impression that category 1 is comparable in the slightest to Open Weights, which is vastly more useful.
bick_nyers··on Tencent Hunyuan-Large
You could always split one of the experts up across multiple GPUs. I tend to agree with your sentiment, I think researchers in this space tend to not optimize that well for inference deployment scenarios. To be fair, there is a lot of different ways to deploy something, and a lot of quantization techniques and parameters.
bick_nyers··on Tencent Hunyuan-Large
Generally speaking this works well, pending your definition of node and the interconnect between them. If by node you mean GPU, and you have multiple of them on the same system (interconnect is PCIE, doesn't need to be full speed however for inference), you're good. If you mean multiple computers connected by 1 Gigabit Ethernet? More challenging.

When splitting models layer by layer, users in r/LocalLLaMA have reported good results with as low as PCIE 3.0 x4 as the interconnect (4GB/s). For tensor parallelism, the interconnect requirements are higher but the upside can be faster speeds in accordance to number of GPUs split across (whereas layer by layer operated like a pipeline, so isn't necessarily faster than what a single GPU can provide, even if splitting across 8 GPUs).

bick_nyers··on Tencent Hunyuan-Large
You would need to fit the 389B parameters in VRAM to have a speed that is usable. Different experts are activated on a per token basis, so you would need to load/unload a large chunk of the 52B active parameters every token if you were trying to offload parameters to system RAM or SSD. PCIE 4.0 x16 speed is 64GB/s, so you can load those active parameters maybe 1 or 2 times per second, yielding an output speed of 1-2 tokens per second, which most would consider "unusable".
bick_nyers··on Quantized Llama models with increased speed and a reduced memory footprint
Perhaps I'm being charitable but I read OP's comment in the light of what you described with context length. Batching, context length, and attention implementation vary these numbers wildly. I can fit a 6bit quant Mistral Small (22b) on a 3090 with ~10-12k context, but Qwen2VL (7b, well 8.3b if you include vision encoder) also maxes out my 3090 VRAM with an 8bit quant and ~16k context.

I do think it would be good to include some info. on "what we expect to be common deployment scenarios, and here's some sample VRAM values".

Tangentially, whenever these models get released with fine-tuning scripts (FFT and Lora) I've yet to find a model that provides accurate information on the actual amount of VRAM required to train the model. Often times it's always 8x80GB for FFT, even for a 7B model, but you can tweak the batch sizes and DeepSpeed config. to drop that down to 4x80GB, then with some tricks (8bit Adam, Activation Checkpointing), drop it down to 2x80GB.

bick_nyers··on Efficient high-resolution image synthesis with linear diffusion transformer
CogVideoX seems to be the best offline model so far
bick_nyers··on Llama 3.2: Revolutionizing edge AI and vision with open, customizable models
Does anyone know of a CoT dataset somewhere for finetuning? I would think exposing it to that type of modality during a finetune/lora would help.
bick_nyers··on Coffee Stats – Maximize Caffeine Intake and Get to Bed at Night
Nowadays whole genome sequencing can be done for $300-400 (for 30x WGS, with 100x around $1000?), it's been on my TODO list for a few years.
bick_nyers··on Coffee Stats – Maximize Caffeine Intake and Get to Bed at Night
Hmm, now I'm interested in seeing if an LLM could be finetuned to give an effective quality score on such a study.
bick_nyers··on DOJ sues realpage for algorithmic pricing scheme that harms renters
I can't read the article because of the paywall, but Kalamazoo is unique due to the Kalamazoo Promise, where college tuition (in the state of MI) is paid for if you attend Kalamazoo Public Schools. For me that was $60k of tuition I didn't have to take out loans for. Housing/grocery prices have gone up there but not nearly as dramatically as other places.
bick_nyers··on DOJ sues realpage for algorithmic pricing scheme that harms renters
I agree that the underlying issue is a lack of supply. However, if a landlord is commanding a 30% profit margin on a non-luxury apartment then I think they are contributing to the problem (to be clear, I'm not insinuating anyone in this thread is doing this). I think the only objective way to tell if it's priced too high is by the profit margin, but of course even that can be inflated if e.g. a developer took a huge margin on it before selling it to a new owner.

As to what an "ethical and not terrible for society" profit margin would be is above my pay grade, but I would estimate 15%. It also probably depends on how easily you can make money on the stock market as well.

bick_nyers··on DOJ sues realpage for algorithmic pricing scheme that harms renters
Housing is pretty inelastic. I think people are just willing to suffer financially to avoid the fate of homelessness. Just because there aren't vacancies anywhere doesn't mean that the price is fully justified, because housing will take priority over groceries for a lot of people.
Page 1 of 19Next →