HNHacker News
TopNewBestAskShowJobs

alexhutcheson

3,780 karma · joined February 17, 2013

Nothing I post reflects the views or policies of my employers, past or present.
submissionscomments
alexhutcheson··on RLHF Book
DeepSeek-R1 had an RLHF step in their post-training pipeline (section 2.3.4 of their technical report[1]).

In addition, the "reasoning-oriented reinforcement learning" step (section 2.3.2) used an approach that is almost identical to RLHF in theory and implementation. The main difference is that they used a rule-based reward system, rather than a model trained on human preference data.

If you want to train a model like DeepSeek-R1, you'll need to know the fundamentals of reinforcement learning on language models, including RLHF.

[1] https://arxiv.org/pdf/2501.12948

alexhutcheson··on RLHF Book
Glad to see the author making a serious effort to fill the gap in public documentation of RLHF theory and practice. The current state of the art seems to be primarily documented in arXiv papers, but each paper is more like a "diff" than a "snapshot" - you need to patch together the knowledge from many previous papers to understand the current state. It's extremely valuable to "snapshot" the current state of the art in a way that is easy to reference.

My friendly feedback on this work-in-progress: I believe it could benefit from more introductory material to establish motivations and set expectations for what is achievable with RLHF. In particular, I think it would be useful to situate RLHF in comparison with supervised fine-tuning (SFT), which readers are likely familiar with.

Stuff I'd cover (from the background of an RLHF user but non-specialist):

Advantages of RLHF over SFT:

- Tunes on the full generation (which is what you ultimately care about), not just token-by-token.

- Can tune on problems where there are many acceptable answers (or ways to word the answer), and you don't want to push the model into one specific series of tokens.

- Can incorporate negative feedback (e.g. don't generate this).

Disadvantages of RLHF over SFT:

- Regularization (KL or otherwise) puts an upper bound on how much impact RLHF can have on the model. Because of this, RLHF is almost never enough to get you "all the way there" by itself.

- Very sensitive to reward model quality, which can be hard to evaluate.

- Much more resource and time intensive.

Non-obvious practical considerations:

- How to evaluate quality? If you have a good measurement of quality, it's tempting to just incorporate it in your reward model. But you want to make sure you're able to measure "is this actually good for my final use-case", not just "does this score well on my reward model?".

- How prompt engineering interacts with fine-tuning (both SFT and RLHF). Often some iteration on the system prompt will make fine-tuning converge faster, and with higher quality. Conversely, attempting to tune on examples that don't include a task-specific prompt (surprisingly common) will often yield subpar results. This is a "boring" implementation detail that I don't normally see included in papers.

Excited to see where this goes, and thanks to the author for willingness to share a work in progress!

alexhutcheson··on Seer: A GUI front end to GDB for Linux
I prefer the GDB Graphical Interface in Emacs[1] (M-x gdb), rather than the more basic integration via GUD[2] (M-x gud-gdb). I’ve had to switch to GUD to run lldb recently, and I miss having dedicated windows that show breakpoints, threads, the current stack, etc.

The one nice thing about GUD is that the interface is consistent across debuggers, so I don’t need to refresh myself on the keyboard shortcuts when switching between debugging Python with pdb and C++ with lldb.

[1] https://www.gnu.org/software/emacs/manual/html_node/emacs/GD...

[2] https://www.gnu.org/software/emacs/manual/html_node/emacs/St...

alexhutcheson··on Seer: A GUI front end to GDB for Linux
GDB also has a built-in text user interface (TUI) that is surprisingly easy to use[1]. It even supports mouse interaction.

[1] https://sourceware.org/gdb/current/onlinedocs/gdb.html/TUI.h...

alexhutcheson··on Ask HN: Alternative to Emacs with undo-tree functionality?
Try disabling VC over Tramp connections[1]:

  (setq vc-ignore-dir-regexp
        (format "\\(%s\\)\\|\\(%s\\)"
                vc-ignore-dir-regexp
                tramp-file-name-regexp))
VC is quite chatty and assumes that filesystem operations have a negligible cost. Before I disabled it, VC was adding >1 second to every find-file operation over Tramp.

I also recommend using the direct-async-process connection property[2], which significantly decreases the latency of async process creation.

[1] https://www.gnu.org/software/emacs/manual/html_node/tramp/Fr...

[2] https://www.gnu.org/software/emacs/manual/html_node/tramp/Re...

alexhutcheson··on Probably pay attention to tokenizers
It depends if they are using a “vanilla” instruction-tuned model or are applying additional task-specific fine-tuning. Fine-tuning with data that doesn’t have misspellings can make the model “forget” how to handle them.

In general, fine-tuned models often fail to generalize well on inputs that aren’t very close to examples in the fine-tuning data set.

alexhutcheson··on Visual Studio Code is designed to fracture (2022)
If you want to be cautious, I have somewhat higher confidence in the versions of Emacs packages published on the Debian repositories[1] than the ones on ELPA/MELPA.

The downside is that not every package is packaged for Debian, and the versions are a bit stale.

https://packages.debian.org/search?keywords=ELPA+&searchon=n...

alexhutcheson··on Pivotal Tracker will shut down
Are there any open source self-hostable tracking/project management tools that still have a committed team and forward momentum?

I used to self-host a Phabricator instance, which I liked a lot, but the upstream maintainer made the reasonable decision to step away.

My guess is there is not much of a niche for self-hosted solutions anymore. The GitHub Issues free tier covers most of the low-complexity use-cases, while higher-complexity use-cases are addressed by enterprise SaaS.

alexhutcheson··on Calendar Queues: A Fast O(1) Priority Queue Implementation (1988)
Boost.Heap has this functionality. Or if you want to stick with the standard library it’s fairly easy to use the *_heap functions from <algorithm> and just hand-code your own fix_heap(first, last, changed) function. Agree it would be more convenient to have it built-in, though.
alexhutcheson··on Visit Bletchley Park
On a related note, if you’re around Washington, DC in the US, then the National Cryptologic Museum[1] is well worth a visit. They have a ton of equipment on exhibit, including a variety of mechanical encryption machines as well as early supercomputers.

It doesn’t get that many visitors, because it’s within Fort Meade (home of the NSA), and most people probably don’t even realize it’s open to the public.

[1] https://www.nsa.gov/museum/

alexhutcheson··on SIMD Matters: Graph Coloring
If you are accessing elements in memory with a constant stride[1], then hardware prefetchers[2] do a surprisingly good job at “reading ahead” and avoiding cache misses.

A typical example would be: you have an array of objects of constant size, and you’re reading a double field from a constant offset within each object. The hardware prefetcher will “recognize” this access pattern and prefetch that offset every sizeof(obj) bytes.

The major downsides (vs. a struct-of-arrays design with full spatial locality) are:

1. Every prefetch pulls a full cache line, but the cache line will include data you don’t need. In this example, every cache line might have 64 bytes of data, but you only needed the one double field (8 bytes) - the rest is not useful. If you were iterating over an array of doubles you could have pulled 8 double fields in a single cache line.

2. Specific performance is hardware-dependent, so it’s hard to guarantee performance on e.g. low-end cores, short loops, or unusually long strides.

[1] https://en.wikipedia.org/wiki/Stride_of_an_array

[2] https://en.wikipedia.org/wiki/Cache_prefetching

alexhutcheson··on Grace Hopper, Nvidia's Halfway APU
Somewhat tangential, but did Nvidia ever confirm if they cancelled their project to develop custom cores implementing the ARM instruction set (Project Denver, and later Carmel)?

It’s interesting to me that they’ve settled on using standard Neoverse cores, when almost everything else is custom designed and tuned for the expected workloads.

alexhutcheson··on The lie of music discovery algorithms
Google has a great free online course on Recommendation Systems that goes through the various common approaches, with working code in Colab notebooks: https://developers.google.com/machine-learning/recommendatio...

[Disclosure: Work at Google, but not on that. Just thought that course was particularly well-designed.]

alexhutcheson··on It's not just you, Next.js is getting harder to use
What makes it the closest to Spring and ASP.NET?
alexhutcheson··on Google Sheets ported its calculation worker from JavaScript to WasmGC
Do you use Flutter Web, or some other framework for web UI?
alexhutcheson··on The Migrants Spending $72,000 to Illegally Immigrate to the US
I’ve seen this book recommended, but have not personally read it: https://www.unshackled.club/book
alexhutcheson··on Ask HN: Compiler speed-up or Build Caching tool. Hard to find?
Specifically: https://bazel.build/remote/caching
alexhutcheson··on Intel's Lion Cove Architecture Preview
Modern hash table implementations use vector instructions for lookups:

- Folly: https://github.com/facebook/folly/blob/main/folly/container/...

- Abseil: https://abseil.io/about/design/swisstables

alexhutcheson··on High performers job hop when they can't find a high performance culture
Useful blog post about this sort of situation: https://lethain.com/hard-to-work-with/
alexhutcheson··on Aboriginal Linux
The successor project uses musl: https://github.com/landley/toybox/tree/master/mkroot
alexhutcheson··on Angle-grinder: Slice and dice logs on the command line
Angle grinders are useful when working with pipes, though.
alexhutcheson··on Angle-grinder: Slice and dice logs on the command line
For those not familiar, Google had a tradition of choosing names of wood-processing tools for “logs” analysis:

- Sawzall: https://research.google/pubs/interpreting-the-data-parallel-...

- Dremel: https://research.google/pubs/dremel-interactive-analysis-of-...

- PowerDrill: https://research.google/pubs/processing-a-trillion-cells-per...

alexhutcheson··on Gnuplotlib: A gnuplot-based plotting backend for NumPy
Vega-Altair is pretty great as well. It uses a grammar of graphics that’s slightly different from ggplot, but has most of the same advantages.

https://altair-viz.github.io/

alexhutcheson··on You (probably) don't need to learn C
Honestly I find Python's value vs. reference semantics fairly confusing, and it's not always obvious whether an operation is going to make a copy or modify an existing instance of an object. Good API design and docstrings mitigate this, but it's one more source of complexity to think about. I run into this much less frequently in C++ code - it's normally pretty clear from parameter and return types whether I'm dealing with a copy or a reference.

This might just be because C++ broke my brain into assuming:

   Object foo = other_object;
is a copy operation, and taking a reference or pointer would require extra characters (e.g. copy by default). Most other languages are the opposite: assignment creates a reference by default, and making a copy would require extra characters (e.g. .copy() in Python or .clone() in Java). That's my biggest mental adjustment when moving back-and-forth between C++ and Python.
alexhutcheson··on You (probably) don't need to learn C
This is a huge pain point. If you're writing small programs that fit in a couple files in a single directory, then you can get by with manual invocations of clang or gcc at the command line. If you want to build anything of moderate complexity, then you need to get familiar with all the complexity of a separate build system like Make/CMake/Bazel/etc. Java has a similar problem - once your program grows beyond ~1 directory, you end up needing to learn Maven/Gradle/etc.

Build and testing tooling is one thing that younger languages have generally done much better. Rust, Go, Dart and others all have standard build tools integrated with the language that scale to large projects, with standard testing frameworks integrated with those tools. Scripting languages like Python have import features within the language itself. Much lower cognitive overhead.

alexhutcheson··on Airbus Shatters Record for Jet Orders as Demand Soars
Airlines with Airbus fleets are dealing with lengthy periods of unavailability due to issues with the Pratt & Whitney engines that power ~60% of A320neos and all A220s[1]. My understanding is that the airlines have agreements where P&W will eventually compensate them for this, but Airbus fleets haven't been without headaches.

[1] https://crankyflier.com/2023/09/26/the-problem-with-pratt-wh...

alexhutcheson··on Netflix never used its $1M algorithm (2012)
This course is a great intro to some of the “workhorse” models for recommendation systems: https://developers.google.com/machine-learning/recommendatio...

The same general methods are also relevant to other matching problems, like search.

Disclosure: Work at Google, but not on anything related to this course.

alexhutcheson··on Show HN: I made a GPU VRAM calculator for transformer-based models
Sincere question: Is installing and running a mini split actually cheaper than racking them in a colo, or paying for time on one of the GPU cloud providers?

Regardless, I can understand the hobby value of running that kind of rig at home.

alexhutcheson··on Show HN: I made a GPU VRAM calculator for transformer-based models
You can only fit 1-2 graphics cards in a “normal” ATX case (each card takes 2-3 “slots”). If you want 4 cards on one machine, you need a bigger/more expensive motherboard, case, PSU, etc. I haven’t personally seen anyone put 6 cards in a workstation.
alexhutcheson··on Show HN: I made a GPU VRAM calculator for transformer-based models
Nvidia’s workstation cards are available with more RAM than the consumer cards, at a lower price than the datacenter cards. RTX 6000 Ada has 48 GB VRAM and retails for $6800, and RTX 5000 Ada has 32 GB VRAM and retails for $4000[1].

Very large models have to be distributed across multiple GPUs though, even if you’re using datacenter chips like H100s.

[1] https://store.nvidia.com/en-us/nvidia-rtx/store/

Page 1 of 29Next →