HNHacker News
TopNewBestAskShowJobs

vladf

424 karma · joined May 16, 2017

vladfeinberg.com
submissionscomments
vladf··on NanoGPT Slowrun: Language Modeling with Limited Data, Infinite Compute
That still looks like a “converge faster” paper.

https://arxiv.org/abs/2006.10732

The above provides a nuanced theoretical view. GD inductive bias is probably better unless your model is misspecified

vladf··on Show HN: AutoThink – Boosts local LLM performance with adaptive reasoning
This is available for Flash
vladf··on Stuff I Learned at Carta
why
vladf··on 4o Image Generation
That's pretty disappointing, it has been out for a while, and we still get top comments like (https://news.ycombinator.com/item?id=43475043) where people clearly think native image generation capability is new. Where do you usually get your updates from for this kind of thing?
vladf··on New York claims a small victory in 'forever war on rats'
Yes perhaps one day cities like Tokyo will catch up
vladf··on Git Bisect-Find
Ah! Finally a real world use for egg drop (with eggs and floors both equal to num commits since init, but maybe fewer eggs for those less patient).
vladf··on Towards 1-bit Machine Learning Models
> You still have a group-size of 64 in 4-bit fyi.

Results may vary :)

> Again, and I keep repeating this but it seems to be ignored every time: this is experimental work and it's still in progress. This story of small group-sizes on large models should not be an issue.

Apologies if something I said (or I guess did not say...) offended you! It's a hypothetical, and one IME is not so easy to achieve, but maybe you have different results. So I didn't want to comment on this, maybe it's possible (but LLMs don't scale up as easily in terms of quantization than other networks like image classifiers, in my experience).

> The extreme quant buys you potentially 70x more efficient matmul via binary/ternary operations.

To be clear, such hardware does not yet exist, and it's unclear if you really can have more efficient binary/ternary matmul if you need high-precision accumulators and more frequent broadcasting shiftss. It's again a complicated hardware question to answer if the sum total latency of doing many high-precision accumulations and many scales/shifts will be smaller (or, chip-area-wise, even feasible to implement), compared to a 4-bit baseline.

vladf··on Towards 1-bit Machine Learning Models
If you’re willing to pay for the latency cost of per layer cpu fetching/offloading, I don’t see what extreme quant buys you.

You could just do a layer-by-layer fetching scheme with 4 bit weights.

For training too, just fetch each layer twice per step as needed for fwd/bwd.

And all for hbm cost equal to one layer’s worth

vladf··on Towards 1-bit Machine Learning Models
I see, so we’re still fetching the metadata to gpu, and rescaling on gpu, just on-demand and discarding metadata when we’re done with that layer?

Why not do the same optimization for layer weights themselves?

vladf··on Towards 1-bit Machine Learning Models
Thanks for the reply. I’m quite familiar with subchannel quant, but still feel like my questions did not get addressed.

1 Could you post the full memory use of the methods? E.g. you include quip metadata in its GB but not hqq metadata in its GB.

2 If you have to go to cpu to shift and scale, how did you get latency lower than pure on device? Was this bsz1? No speculative decoding?

3 how can lora absorb shifts with only increasing rank by 1 if you have a shift per group?

vladf··on Towards 1-bit Machine Learning Models
Err, you are just restating what I’m saying, without addressing the concerns.

1 - is it fair to use ram in two places and report only one of them without any asterisk? (If you think this is fair-oh boy wait till you hear about my 0GB hbm use inference algorithm)

2 - i know how subchannel quantization works. Are they hitting those reported latency numbers with per layer cpu pingpong to rescale?

vladf··on Towards 1-bit Machine Learning Models
Really strong binary results. So strong it was fishy. I hope someone can explain my confusion below.

> We compared the performance of the Llama2-7B model in three configurations: FP16 (full precision), HQQ (without fine-tuning), and HQQ+ (with adapter layers) using a group-size of 8.

Interesting, what is "group-size of 8"?

From their HQQ post (https://mobiusml.github.io/hqq_blog/), it's the block size at which they add scales (presumably 16-bit) and shifts (in that post, it's 8-bit).

So for every 8 binary weights we have a 16-bit scale and 8-bit shift?

> Fine-tuning with Low-Rank Adapters

They say they inline the shift into the LoRA but how can you do this, block-wise, without increasing your LoRA rank by num-blocks (they claim to only use 1 additional rank)?

Then, the reported 7B sizes, in GB:

> 13.5 (fp16) 1.76 (HQQ 1-bit) 1.85 (HQQ+ 1-bit) 2.72 (quip# 2-bit)

those numbers would make sense if it was _actually_ 1 bit. But if you include the overhead of 16-bit scales (and why is the shift inlineable into lora? still unexplained) it'd be more like 3-bit.

From their HF page:

> This version offloads the meta-data to the CPU, so only the binary weights and the low-rank adapters are stored in the GPU memory.

Interesting, so we have to go back to CPU to rescale? Is this how they counted GB? This should have been clearly caveated in the table. I also am amazed they got latency lower than quip if they pingpong to CPU.

vladf··on Jim Keller criticizes Nvidia's CUDA, x86
I’m not sure if you’re being facetious, but this is literary available for early access, but not ga yet. https://simonwillison.net/2024/Feb/21/gemini-pro-video/
vladf··on 'Baby Bust': Why Fewer Young People Expect to Become Parents (2013)
Isn't this literally the case in the US? You list dependents on your tax form.
vladf··on Ask HN: Do companies hire principal and staff level engineers from job postings?
What is your current level? Is it shown on your LinkedIn?
vladf··on Push ifs up and fors down
And yet, a Rust Option (or really any option) can just be viewed as a list of one or zero elements. https://rust-unofficial.github.io/patterns/idioms/option-ite...

In fact, in Haskell, operating on an option conditionally has the exact same functor as a list: `map`.

So what am I to do, with an iterator? It's conflicting advice! An if is a for for an option.

vladf··on What scientists must know about hardware to write fast code (2020)
I ended up needing this so often for graph processing, and for values which might be inexact if using floating point, that I saved the formula in a blog post. https://vladfeinberg.com/2020/03/07/subset-isomorphism.html

The formula can be "oblivious" to the final size of the matrix too, which is helpful if you're doing some sparse ML training on edges (e.g., GNNs).

vladf··on Hutter Prize for compressing human knowledge
What
vladf··on Public restrooms are hard to find in America
What are some non-sanctuary cities which also have year round temperate, but not too hot or humid, weather?
vladf··on Chunking 2M files a day for code search using syntax trees
Which papers?
vladf··on Invisible asymptotes (2018)
I don't know if your questions are rhetorical, but have you ever been in the car with an exec? They very much do work in the car and take meetings from the car.
vladf··on FTC investigating ChatGPT over potential consumer harm
Intelectual parity to an average human by 2030? Would you be willing to bet on this?
vladf··on Getting to know the right people (2022)
You forgot the part where the only reason they shared the note with you was as a show of power.
vladf··on A Swedish startup’s bid to build a green rival to AWS
Sure, but the reasons offsets are dodgy is because they actually aren't fungible in the way you're hoping for.

Say a paper maker is about to cut down a forest in the US. It buys up the rights to do so for $1M. Then, instead, it sells $1M worth of carbon credits to Amazon (maybe for a premium) for its data center greenwashing and doesn't end up cutting the US forest.

But then, secretly, or through some shell company machinery, it buys a plot of land in the Amazon and cuts trees down anyway to match its demand. Even though US law makers evaluate the counterfactual of the paper maker cutting trees down in the US, the net amount of trees in the world goes down.

vladf··on Chain-of-Thought Hub: Measuring LLMs' Reasoning Performance
> Also be careful that GPT-4/ 3.5's performance on GSM8K is not true few-shot -- in GPT-4 report they said that they mixed a portion of GSM8K training set to train the model

It'd be really valuable to have "fuzzed" versions of these benchmarks, where you replace quantities in the questions with randomly-sampled values, so that this wasn't a concern. Of course, then the score would itself be a random variable, but you could just return an interval.

vladf··on Google Calendar and Assistant Reminders Will Migrate to Google Tasks Soon
Have you seen this? https://xkcd.com/1172/
vladf··on NewsNotFound: An open-source, unbiased news company
A bit of an obscure reference for the anglocentric crowd, but Night Guard’s news in the Twilight World is exactly what your lede is imagining.
vladf··on Spotify abandons Heardle less than a year after buying it
Same; I switched to YT Music b/c of better offline playback than Spotify Premium (and ofc YT Premium simply includes music!)

however i've found that the mixes introduce me to recommendations that I like at roughly the same rate as Discover Weekly.

What are you listening to for such good recommendations? Is the the automatic "Mix" playlists it creates?

vladf··on DyLoRA: Parameter Efficient Tuning of Pre-Trained Models
Err, I suppose trivially, the higher rank terms include the lower-rank subnets, so they dominate in terms of quality.

But if you have some capacity constraint (e.g., memory, I guess?) then you can imagine dynamic rank allocation helping in the case where the maximum rank across all layers isn't within budget.

It's a bit of a stretch though, I agree

vladf··on DyLoRA: Parameter Efficient Tuning of Pre-Trained Models
The optimal rank could differ across layers
Page 1 of 5Next →