HNHacker News
TopNewBestAskShowJobs

chewxy

5,014 karma · joined November 3, 2010

You can contact me here: chewxy [at] gmail dot com

My personal blog is http://blog.chewxy.com . Be warned. Lots of nonsense in there.

submissionscomments
chewxy··on Opus 4.7 knows the real Kelsey
So I have been practicing writing fiction the past year or so. It identifies a fiction piece I wrote as Greg Egan[0]. Another paragraph from another piece was identified as China Mieville[1]. The accompanying blog posts explaining the making of the fiction pieces were identified as me.

Both pieces have never been published. Neither have the blog posts.

[0] in https://blog.chewxy.com/2026/04/01/how-i-write/ this is the story titled "there is no constant non-zero derivative in nature". It does not read like Egan at all.

[1] in https://blog.chewxy.com/2026/04/01/how-i-write/ this is the story titled "The Case of the Liquidated Corps". I use a lot of biological metaphors. Once again, nothing like Mieville.

If only I could write like them! These pieces were all rejected by the major scifi mags

chewxy··on Waiting for SQL:202y: Group by All
BigQuery has that and I've been loving using it since they introduced it
chewxy··on Don't Force Your LLM to Write Terse [Q/Kdb] Code: An Information Theory Argument
> Limited training data, as compared to kanguages like Python or javascript.

I use my own APL to build neural networks. This is probably the correct answer, and inline with my experience as well.

I changed the semantics and definition of a bunch of functions and none of the coding LLMs out there can even approach writing semidecent APL.

chewxy··on Do Large Language Models know who did what to whom?
Maybe read the paper first?

> This study asked whether Large Language Models (LLMs) understand sentences in the minimal sense of representing “who did what to whom”. In Experiment 1, we found that the overall geometry of LLM distributed activity patterns failed to capture this information: similaritiesbetween sentences reflected whether they shared syntax more than whether they shared thematic role assignments. Human judgments, in contrast, were strongly driven by this aspect of meaning.

> In Experiment 2, we found limited evidence that thematic role information was available even in a subset of hidden units. Whereas activity patterns in subsets of hidden units often allowed for significant classification of whether sentence pairs had shared vs. opposite thematic role assignments, the effect sizes were small; even the best-performing case appeared to lag behind humans, and its representation of thematic roles did not seem robust across syntactic structures.

> However, thematic role information was reliably available in a large number of attention heads, demonstrating LLMs have the capacity to extract thematic role information. In some cases, information present in attention heads descriptively exceeded human performance.

chewxy··on Are Plants Farming Us?
I wrote a short story ten years ago making fun of this concept: https://blog.chewxy.com/2014/05/20/the-long-term-plan/
chewxy··on Tree Calculus
Barry Jay's got an upcoming paper at PEPM regarding typed tree calculus. Good read too.
chewxy··on Ask HN: What are you working on (September 2024)?
Thanks :)
chewxy··on Ask HN: What are you working on (September 2024)?
At this point I'm writing mostly for myself. GTM strategies for novels... that's an interesting way to think about things. I've not thought about it just yet. Happy to hear if you have any ideas tho.
chewxy··on Ask HN: What are you working on (September 2024)?
I'm working on my scifi novel. I had started writing it when LLMs started taking off - I had been doing AI for two decades and I was well-placed to be in a good position to profit with the rise of LLMs, but I ended up gaining nothing much and I was depressed about it - so I started writing instead. Been picking at it for about a year before befriending an editor who encouraged me to keep writing. He's helped me developmentally edit it to a point I am now ready to work on my second draft.

It's a hard scifi novel with mild existential horror tones that is borne mostly of maths jokes. At one point the main character tries to escape the matrix (reality). But the matrix is defective, so the best way out was to orthogonalize the subspace and reduce the matrix to its eigenbasis instead. Most of the scenes are based on similar maths jokes.

Tentative name is Diagonalization of the Meta (I had previously called it The Metaverse).

chewxy··on IOGraphica
I did this 13 years ago: https://blog.chewxy.com/2011/08/31/one-point-five-hours/
chewxy··on APL Demonstration (1975) [video]
Not really, this is actually pretty readable
chewxy··on The case for not sanitising fairy tales
I told a variant of the original Little Mermaid story as part of a school outreach program. The kids came to the conclusion that God wasn't a fair being because he didn't give mermaids souls. I walked away satisfied that my little counterprogramming against catholic school indoctrination might have worked. I wasn't invited back (at least for school year 2024).
chewxy··on Mitochondrial signal transduction (2022)
there's also a good paper - Can a neuroscientist understand a microprocessor - by Kording's lab which is also an excellent read.
chewxy··on Why we no longer use LangChain for building our AI agents
Dana Angluin's group were studying chat systems way back in 1992. There even was a conference around conversational AI back then.
chewxy··on Eight years of organizing tech meetups (2023)
Haha I remember you and your SaltStack talks at SyPy. Shame you moved to Melbourne. Hey, coffee's better here now (https://www.timeout.com/sydney/news/experts-have-ranked-the-...). Wanna come back?
chewxy··on How it feels to get an AI email from a friend
Most LLMs out there are very American in their writing mannerisms. The article even alludes to this:

> The AI did, however, try to sound like someone. It was folksy and upbeat, talky and pretend-excited

I've not seen an LLM, even when fine tuned that doesn't actually do that (Chinese LLMs excepted). There's something just inherently American about the instruction datasets that these LLMs are instructed with.

chewxy··on Eight years of organizing tech meetups (2023)
I've been running GolangSyd for about 8 years now, and previously Sydney Python for about the same amount of time. I found finding sponsors to be one of the hardest things. One of the last Sydney Pythons in which I gave a talk on machine learning had ~250 people attending (PyConAU at that time had roughly as many people I think). So large that the pizza bill was way more than the host (who was the sponsor as well) had anticipated, so we were no longer welcome at that venue. And said host is a well known billion dollar Aussie company.

Thankfully for GolangSyd, my cohosts have been extremely talented with finding sponsors. This gives us a lot more opportunity to do weirder things like https://gogogogogo.casa .

Running meetups are hard work, full stop. On the other hand, I've gotten to know some people very well and some of my best collabs have been thru these meetups. Running a meetup was a way for me to overcome my own reluctance to socialize.

chewxy··on AI dev startups are struggling with one problem and I solved it
So.. sourcegraph?
chewxy··on I never stopped learning from Daniel Dennett
This is not a nice way to discover Dan Dennett died
chewxy··on After AI beat them, professional Go players got better and more creative
you have to admit that 3-3 invasion is pretty annoying to handle. AI is way too aggro and people are learning to be as aggressive.
chewxy··on Receive push notifications from your rice cooker
A good thing about living in a small house. I can hear when my rice cooker is done (plays a tune), when my washing machine is done (plays a tune), when my dryer is done (plays a tune). :) (note that that hasn't prevented me from hooking them up to smart plugs the way terrence did)
chewxy··on Ask HN: What Are You Learning?
Go (the game). I wrote a clone of AlphaGo in Go (the programming language) 8 years ago. Along the way I learned to play Go.

I've been using a combination of my own AI, LeelaZero and KataGo to teach myself in Go. For 8 years I've languished at the same level of play. Then I met a real human teacher who taught me Go at the end of 2023. And since then my game improved. I beat my own AI (which was intentionally trained to be impoverished in skill) for the first time in January.

Learning Go is teaching me all sorts of new ideas in pedagogy and putting a dampener in any enthusiasm that involves LLMs in education.

chewxy··on Of course AI is extractive, everything is lately
I agree. There are other types of AIs with different applications that do not need to be trained on the internet. The examples you have given however, are examples where the deep nets are extremely data hungry.

Take computer vision for example - a "hello world" version of object recognition would use ImageNet, which is 14 million hand annotated images. Or Cifar10 which is 80 million images. That of course but sets the stage for training data differentiation. Google's image recognition algorithm is far superior to other search engines'. Why? Because of Google's data set.

Any Tom Dick and Harry can go create their own image recognition AI and train it based on all the public datasets (COCO, CIFAR, ImageNet) but that's considered pretty baseline nowadays. The differentiator is what _other_ datasets you have.

Different datasets yield different results. It doesn't matter the network. More data is better (usually).

chewxy··on Of course AI is extractive, everything is lately
It was good enough in 2015/2016 for me to run a startup that allowed people to program in natural language. We even had paying clients though eventually none could stomach the $2000 per month for incremental/on-line training costs.

The only real difference between then and now is that OpenAI's models are significantly better than my models from 2015, and they have that because well, they can afford to pile on more data. TBH, I never even considered using a large proportion of the whole internet as a training set as even remotely possible due to the sheer mind boggling costs.

Even now, to go through about 10% of The Pile would cost me way too much money.

chewxy··on Of course AI is extractive, everything is lately
GP mentioned that the current slate of transformer based AIs are not transformative in the same way the Internet was. Rather it's more of a triumph of data engineering practices.

OP disagrees with GP. OP's main thesis is that AI enables a lot new applications. OP claims that GP is simply looking at it as if it were training data.

I stated that current AI techniques ARE indeed just reflections of the data used in training. I agree with GP that the current "AI"s are simply not transformative in the same way the Internet was.

If you change the training data for the current generation of AI, you get different behaviours. The training data forms a manifold - which you can think of as a landscape with features forming valleys and hills. What the current generation of AI does is that it tries to find a shape that fits the landscape - think of it like taking a very large sheet of cloth to cover a landscape. The stiffer the cloth, the less well the cloth fits to the landscape. The "stiffness" of the cloth is the amount of parameters that a neural network has. Modern deep nets are highly overparameterized - imagine a very soft pliable cloth - of course it fits to a landscape well.

So if you have a different training data - the neural network will fit to this different landscape as well. Hence the response will be different.

It's unfortunate that the training data is the entire internet for a few reasons:

1. Only the rich can train a vaguely competent AI. You're at the whims of those well-resourced enough. 2. There's no "alternate" training dataset anymore. (Though a clever thing people at OpenAI are doing are Mixture of Experts models, where you train multiple NNs using different subsets of the full training set, so you get multiple competencies)

chewxy··on Of course AI is extractive, everything is lately
Current AI/ML is in fact a reflection on training data. Change the manifold, and change the response. It's unfortunate that of course the data is the entire internet.
chewxy··on How I keep myself alive using Golang
I talked to Matt about not owning our own data after GopherConSG where he gave this talk. It was enlightening how complicated the issue is - there's a lot of legal liabilities on the end of the data provider (the company that monitors the glucose) so I can understand why larger corps are a bit hesitant to open up.

On the other hand, it seems quite heinous that users don't have access to data that is rightfully theirs that they can action on it.

chewxy··on Short-term Hebbian learning can implement transformer-like attention
The paper is about implementing transformer-like attention using simulated human neurones
chewxy··on Testing how hard it is to cheat with ChatGPT in interviews
The peppy, upbeat, ultra-American tone that the LLMs produce can be somewhat toned down with good prompting but ultimately, it does stink of the refinement test set.
chewxy··on I just wanted Emacs to look nice – Using 24-bit color in terminals
Meanwhile I want my emacs to look more like this https://imgur.com/a/h0jA1ro

(in case anyone thinks that's serious, it's a joke. I only use Cool Retro Term for presentations)

edit: apparently my emacs works with 24 bit colour out of the box: https://imgur.com/a/BM5OTxp. The syntax highlighting is a bit annoying though.

Page 1 of 34Next →