HNHacker News
TopNewBestAskShowJobs

gthompson512

30 karma · joined March 19, 2025

submissionscomments
gthompson512··on I Wrote a Compiler
@dang I feel like this is getting close to the lowest level of discourse and occurs somewhat often these days where people overly reference the output of some AI and then are challenging the results without googling. I feel a proper hacker ethos would spur someone to find a real answer. So in honor of the name of the site maybe disallow refutations or other dismissals of a submission based on the output of an AI or overly long discussion that is offtopic about an AI?
gthompson512··on “Secret Mall Apartment,” a Protest for Place
Sorry I read about this way long ago, what is meant by "worked for me" "project", "behind my back" and such. I didn't get the connotations for most of that from reading articles about this. What is the real story?
gthompson512··on AI in my plasma physics research didn’t go the way I expected
There is a bunch of new phrases people have been using around AI topics, so it can be hard to tell what exactly is being talked about, like "Roko's Basilisk", the "lottery card hypothesis?"(not quite remembering the phrasing on this one), "the bitter lesson", "paperclip maximizing", "stochastic parrot", etc.. Thank you for clarifying. I was kind of hoping there was a fun blog or story with a "banana zone".
gthompson512··on A simple search engine from scratch
I have been thinking a bit lately about how much sense that makes compared to just using word vectors, since traditional queries are super short and often keyword based(like searching for "ground beef" when wanting "ground beef recipes I can cook easily tonight") and so lack most of the context that BERT or similar gives you. I know there are methods like using seperate embeddings for queries and such, but maybe a basic word based search could be more useful, especially with something like fastText for out of vocabulary terms.
gthompson512··on AI in my plasma physics research didn’t go the way I expected
> "the start of the banana zone"

What does this mean? Is it some slang for exponential growth, or is it a reference to something like the "paperclip maximizer"?

gthompson512··on Poll: How often do you bathe?
I usually exercise ~5 days a week and on those days will shower or bathe twice. On days when I am doing yard work or grilling it might bathe 3 times depending on how poorly I planned out my day.
gthompson512··on Show HN: I modeled the Voynich Manuscript with SBERT to test for structure
Sorry if I missed it, but what about keeping the suffixes and trying to do some finetuning on the source then clustering sentences or at least pages which given the media should be consistent-ish
gthompson512··on Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust
Sorry, looking more, it doesn't seem like you are doing what you are saying. This is just poorly breaking text into bad chunks with no regard for semantics and is like ~200 lines of actual code. What is this for? Most models can handle fairly large contexts.

Edit: That wasn't intended to be mean, although it may come off that way, but what is this supposed to be for? Myself I have text >8k tokens that need to be embedded and test things regularly.

gthompson512··on Show HN: Model2vec-Rs – Fast Static Text Embeddings in Rust
How does it handle documents longer than the context length of the model? Sorry there are a ton of these regularly and they don't usually think about this.

Edit: it seems like it just splits in to sentences which is a weird thing to do given in English only 95%ish percent agreement is even possible on what a sentence is. ``` // Process in batches for batch in sentences.chunks(batch_size) { // Truncate each sentence to max_length * median_token_length chars let truncated: Vec<&str> = batch .iter() .map(|text| { if let Some(max_tok) = max_length { Self::truncate_str(text, max_tok, self.median_token_length) } else { text.as_str() } }) .collect(); ```

gthompson512··on COBOL front-end added to GCC
This has been planned for a while, is this news?
gthompson512··on Phi-4 Reasoning Models
Sorry if this comment is outdated or ill-informed, but it is hard to follow the current news. Do the Phi models still have issues with training on the test set, or have they fixed that?
gthompson512··on Migrating away from Rust
> So in D, is it now natural to mix borrow checking and garbage collection?

I think "natural" is a bit loaded, there is native support in the frontend for doing both. You have to go out of your way to annotate functions with @live and it is still experimental(https://dlang.org/spec/ob.html). The garbage collection is natural and happens if you do nothing, but you can turn it off with proper annotations like @nogc(https://dlang.org/spec/function.html#nogc-functions) or using betterC(https://dlang.org/spec/betterc.html). There is also @safe, @system and @trusted(https://dlang.org/spec/memory-safe-d.html).

So natural is a stretch at the moment, but you can use all kinds of different techniques, what is needed is more community and library standardization around some solutions.

gthompson512··on Are polynomial features the root of all evil? (2024)
There is a simple corollary to Stone-Weierstrass that extends to infinite intervals, but requires the use of rational functions.
gthompson512··on Show HN: I rewrote few of the common core string.h functions
Depending on what exactly you are trying to learn from this, I would recommend looking at the source for musl like https://git.musl-libc.org/cgit/musl/tree/src/string and trying to understand why it looks so much more complicated than the simple implementation.
gthompson512··on Hacktical C: practical hacker's guide to the C programming language
Minor correction, macros CANT have newlines, you need to splice them during preprocessing using \ followed by a new line, the actual code has these:

from https://github.com/codr7/hacktical-c/blob/main/macro/macro.h

#define hc_align(base, size) ({ \ __auto_type _base = base; \ __auto_type _size = hc_min((size), _Alignof(max_align_t)); \ (_base) + _size - ((ptrdiff_t)(_base)) % _size; \ }) \

After preprocessing it is a single line.

gthompson512··on Stop using e for compound interest
> - "My name is Hardy, G.H. Hardy.": A unique function satisfies exp(x+y) = exp(x)exp(y).

This has nothing to do with e and is satified by 2^x or any a^x, so this wouldn't work for introducing e in particular.

- "The Classic": There exists a unique function equal to its own derivative up to a constant.

Same for this, but if you fix the constant to be 1, then e^x is the only one that works.

I will give the series works too.

gthompson512··on PEP 750 – Template Strings
This led to the OpenD language fork (https://opendlang.org/index.html) which is led by some contributors who had other more general gripes with D. The fork is trying to merge in useful stuff from main D, while advancing the language. They have a Discord which unfortunately is the main source of info.
gthompson512··on A love letter to the CSV format
This forces each field to be quoted, and it assumes that each row has the same fields in the same order. A library can handle the quoting issues and fields more reliably. Not sure why you went with a generator for this either.

Most people expect something like `12,,213,3` instead of `"12","213","3"` which yours might give.

https://en.wikipedia.org/wiki/Comma-separated_values#Basic_r...

gthompson512··on A love letter to the CSV format
> It's just that people tend to use specialized tools for encoding and decoding it instead of like ",".join(row) and row.split(",")

You really super can't just split on commas for csv. You need to handle the string encodings since records can have commas occur in a string, and you need to handle quoting since you need to know when a string ends and that string may have internal quote characters. For either format unless you know your data super well you need to use a library.