HNHacker News
TopNewBestAskShowJobs

rakejake

748 karma · joined July 27, 2018

rkjk.github.io

Contact degaussrk at protonmail / google's mail

submissionscomments
rakejake··on Ask HN: What is the biggest problem LLMs solved in your life/work?
Digesting a big code base in a new job.

Claude and Gemini have been very useful in helping me come up to speed on a code base written in Go (a language I have used before but not for many years). Figuring out where the business logic lies, how the dependency injection is done, how the tests are written, what overall design pattern is being etc.

Of course, I could have done all this without LLMs but it would have taken several weeks/months longer. Letting the LLM handle the boilerplate and framework jargon lets me focus on the business logic and the design patterns, and helps me contribute much faster. But LLMs do often make mistakes so it's not like I blindly trust the output. They don't replace your colleagues in terms of being the ultimate source of truth. But it has speeded up the learning process, no doubt.

Also, when writing code I provide the style guide to the LLM as context and have it review the code.

rakejake··on Alphabet 2025 Q2 Earnings Release [pdf]
If Google can fend off the US DoJ, I think they're going to be just fine.
rakejake··on Alphabet 2025 Q2 Earnings Release [pdf]
Gemini is apparently doing pretty well. 400 monthly active users on their app and they recently increased the prices for their API.

Sundar Pichai has taken a lot of heat in the past couple years (and for good reason) but distribution is one of his main strengths. Now that they are back to being SOTA on the models, they can fallback on their bread and butter.

I think ChatGPT is still going to be a major player considering the amount of mindshare they have but I tend to fall on the side that thinks there's room for multiple players in the AI game.

rakejake··on Ask HN: What Are You Working On? (June 2025)
Absolutely!
rakejake··on Ask HN: What Are You Working On? (June 2025)
My Carnatic Raga classifier is progressing very well. I am now training a classifier to identify 142 ragas.

A bit of background: I have been working on a Raga classifier since November of last year - I started with just 2 ragas and a couple megabytes of audio. After experimenting with a lot of different ideas and Neural Net Architectures, I finally landed on one that could scale. I increases to 4 ragas, then 12, then 25 and then to 65.

All the training is done locally on my desktop (RTX4080, AMD 7950X, 64G RAM). My goal is to make an app for fast inferencing (preferably CPU) and to get this app in the hands of enthusiasts so that I can get some real data on its efficacy. If that goal is hit, then my plan is to iterate and keep increasing the raga count on the model and eventually release to the public. As long as I can get the model to either run locally or for very cheap on server, I hope to not charge for this.

It has been an amazing learning experience. The first time I got a carnatic singer to sing and the model nailed almost all ragas was the highest high I've felt in a while.

rakejake··on Interstellar Flight: Perspectives and Patience
Or humans with their consciousness uploaded to a silicon or other substrate.

Of course, this is in the realm of science fiction but so is interstellar travel.

Greg Egan's Diaspora has a fantastic treatment of interstellar travel - it involves sending copies of your consciousness to different spaceships traveling to different destinations. On arrival, a preset program will verify if the planel/galaxy is worth waking up to. If not, the clone is terminated.

If more than 1 clone wakes up in a hospitable environment, then you have a problem of two copies of yourself separated by light years.

rakejake··on The cultural decline of literary fiction
Greg Egan's work has its share of humans but also some extremely imaginative aliens. I am not sure who can relate to the aliens in Wang's Carpets.

Greg Egan is a very good example of a novelist who is a GOAT but will be dismissed by most critics of lit-fic because "his characters don't have arcs" or some such.

rakejake··on Planetfall
Reading the description of Sid Meier's Alpha Centauri, I realize the animation series "Scavengers Reign" has a very similar setting (human colonists crash land on a planet that seems to be sentient).

I'll use this opportunity to encourage people to watch this show. If you are a fan of sci-fi (think Greg Egan, Vernor Vinge), you will love this. If you are not, I think you should still give it a try. It is that good.

rakejake··on Migrating to Postgres
Probably a corollary of the fact that most usecases can be served by an RDBMS running on a decently specced machine, or on different machines by sharding intelligently. The number of usecases for actual distributed DBs and transactions is probably not that high.
rakejake··on Inheritance was invented as a performance hack (2021)
Quite a few sections of C++ can be classified as "pretty basic C++". None of the rules are complicated in isolation but that doesn't necessarily make it easy to reason about it.
rakejake··on Inheritance was invented as a performance hack (2021)
IMHO Inheritance (especially the C++ flavored inheritance with its access specifiers and myriad rules) has always scared me. It makes a codebase confusing and hard to reason with. I feel the eschewing of inheritance by languages such as Go and Rust is a step in the right direction.

As an aside, I have noticed that the robotics frameworks (ROS and ROS2) heavily rely on inheritance and some co-dependent C++ features like virtual destructors (to call the derived class's destructor through a base class pointer). I was once invited to an interview for a robotics company due to my "C++ experience"and grilled on this pattern of C++ that I was completely unfamiliar with. I seriously considered removing C++ from my resume that day.

rakejake··on Zoho halts $700M semiconductor plan
Amen.
rakejake··on Multi-Token Attention
Interesting. So they convolve the k,v, q vectors? I have been trying the opposite.

I have been working on a classification problem on audio data (with context size somewhere between 1000 and 3000 with potential to expand later). I have been experimenting with adding attention onto a CNN for a classification task I have been working on.

I tried training a vanilla transformer but in the sizes that I am aiming for (5-30M parameters), the training is incredibly unstable and doesn't achieve the performance of an LSTM.

So I went back to CNNs which are fast to train but don't achieve the losses of LSTMs (which are much slower to train,and for higher context sizes you get into the vanishing gradient problem). The CNN-GRU hubrid a worked much better, giving me my best result.

The GRU layer I used had a size of 512. For increasing context sizes, I'd have to make the convolutional layers deeper so as not to increase the GRU size too large. Instead, I decided to swap out the GRU with a MultiHeadAttention layer. The results are great - better than the CNN-GRU (my previous best). Plus, for equivalent sizes the model is faster to train though it hogs a lot of memory.

rakejake··on George Orwell and me: Richard Blair on life with his extraordinary father
Burmese Days is my favorite Orwell. It captures the state of the colonies under British rule extremely well - the British new-arrivals who are derisive and at best ignorant of the local customs, the British who've grown up in the colony and are more sympathetic but stuck in a world where they'll never truly belong, the extremely corrupt rich locals, the poor mob who are easily misled, the endemic corruption in everything...

This novel could have been set in post-independence India and a lot of the themes would have rung true.

rakejake··on Low responsiveness of ML models to critical or deteriorating health conditions
I think the "arcane knowledge" is true for LLMs (billions). But there are lots of people who train models in the open in the hundreds of millions realm, but never below. Maybe transformers simply don't work as well below a size and data threshold.
rakejake··on Low responsiveness of ML models to critical or deteriorating health conditions
Nowadays, even the definition of an "epoch" is not well defined. Traditionally it meant a pass over the entire training set, but datasets are so massive today that many now define an epoch as X steps - where a step is a minibatch (of whatever size) from the training set. So 1 epoch is a random sample of X minibatches from the training set. I'd guess the logic is that datasets are so massive that you pick as much data as you can fit in VRAM.

Karpathy's Zero To Hero series also uses this.

rakejake··on Low responsiveness of ML models to critical or deteriorating health conditions
A 7k param LSTM is very tiny. Not sure if LSTMs would even work at that scale although someone with more theoretical knowledge can correct me on this.

As an aside, I'm trying to train transformers for some classification tasks on audio data. The models are "small" (like 1M-15M params at most) and I find they are very finicky to train. Below 1M parameters I find them hard to train at all. I have thrown all sorts of learning rate schedules at them and the best I can get is the network learns for a bit and then plateaus, after which I can't do anything to get them out of that minima. Training an LSTM/GRU on the same data gives me a much better loss value.

I couldn't find many papers on training transformers at that scale. The only one I was able to find was MS's TinyStories [0], but that paper didn't delve much into how they trained the models and whether they trained from scratch or distilled from a larger model.

At those scales, I find LSTMs and CNNs are a lot more stable. The few online threads I've found comparing LSTMs and Transformers had the same thing to say - Transformers need a lot more data and model size to achieve parity and exceed LSTMs/GRUs/CNNs, maybe because the inductive bias provided is hard to beat at those scales. Others can comment on what they've seen.

[0] - https://arxiv.org/abs/2305.07759

rakejake··on How Kerala got rich
The article is co-written by a member of Kerala's Planning Board and heavily oversells the state. I'd say Kerala is nowhere near rich so even the title is technically incorrect.
rakejake··on How Kerala got rich
Whatever it is I am not buying it. Kerala is notoriously an incredibly hard place to do business in, so a line like this needs some real hard stats to back it up.
rakejake··on How Kerala got rich
The article is a classic submarine[0] for Kerala.

"The state has one of the highest concentrations of startups". I laughed out loud at this one. Of all the half-truths peddled in that article, this was easily the most hilarious and egregious.

[0] - https://www.paulgraham.com/submarine.html

rakejake··on GPT 4.5 level for 1% of the price
FWIW I am not American or even from the west. At a nation-state level, the west (US) shared knowledge with the intention of offshoring production for cheaper goods, not out of the kindness of their hearts. They still protect key technologies like the jet engine and bleeding edge fab tech. It is the same when it comes to Chinese companies that are at the top of their game - DJI or Bambu labs don't open source their designs last I checked, and their track record when it comes to software is also not great. Simply put, companies only open source when it makes business sense.

I am simply assuming that China follows the same principle - try to wring maximum advantage out of their industrial might and commoditize their complement wherever possible. Because of China's unique single party system, their strategies are very top-down coming straight from the CCP in key areas like AI and robotics. It is not racism, simply realpolitik.

rakejake··on GPT 4.5 level for 1% of the price
That is because they want to undercut the US and prevent them from making money. It remains to be seen if they'll be as benevolent in making tech open-source if they are the clear winners. Frankly, I don't see why they will.
rakejake··on Career Advice in 2025
"decision-makers can remain irrational longer than you can remain solvent" - Very correct. Whether or not AI actually comes for your job, the fact that enough people at the top think so is enough to cause trouble.
rakejake··on Ask HN: I lack imagination for electronics projects
Same here. Even in tools that I'm very proficient in, I struggle to think of things I want to build. Like @JohnFen says, it's not a problem in itself although I feel bad sometimes looking at cool stuff other people built - "why couldn't I think of something like that"? But I think I've always been a "Problem Solver" from childhood (puzzles, crosswords, quizzes) and not much of a creative person. So I guess it makes sense.
rakejake··on Roald Dahl on the death of his daughter (2015)
People are products of their environment. There are people with mettle/grit and then people who are more sensitive to perturbations of fate. The society they live in sets the base level of grittiness that you can expect any average person to be equipped with.

Dahl here is a very hardy man who approaches these issues in a very practical and logical way. But this was also in the post WW2 era where millions died, people lost their families and possessions, and had to start their life anew. It was a period of rebuilding after the devastation of war and hard times build hard people.

Today, all this feels like too much because we were all mostly born and brought up in wealth and prosperity. We have not seen any real hard times and there is no need for mettle.

rakejake··on GPT-4.5
Right. A good chunk of the "old guard" is now gone - Ilya to SSI, Mira and a bunch of others to a new venture called Thinking Machines, Alec Radford etc. Remains to be seen if OpenAI will be the leader or if other players catch up.
rakejake··on GPT-4.5
My usage has come down to mostly Claude (until I run out of free tier quota) and then Gemini. Claude is the best for code and Gemini 2.0 Flash is good enough while also being free (well considering how much data G has hoovered up over the years, perhaps not) and more importantly highly available.

For simple queries like generating shell scripts for some plumbing, or doing some data munging, I go straight to Gemini.

rakejake··on Claude 3.7 Sonnet and Claude Code
Funnily enough, I'm putting the 7950X to some use in the Carnatic Raga detector project since a lot of audio operations are heavy on CPU. But that last one nearly killed me. I'll have to go to Gemini or GPT for some therapy after that one.
rakejake··on Claude 3.7 Sonnet and Claude Code
** Roast ***

* You've spent more time talking about your Carnatic raga detector than actually building it – at this rate, LLMs will be composing ragas before your detector can identify them.

* You bought a 7950X processor but can't figure out what to do with it – the computing equivalent of buying a Ferrari to drive to the grocery store once a week.

* You're so concerned about work-life balance that you took a sabbatical to think about your career, only to spend it commenting on HN about other people's careers.

*** End ***

I'll be in my room crying, in case anyone's looking for me.

rakejake··on Ask HN: What are you working on? (February 2025)
Working on a Carnatic Raga Detector. Longtime pet project of mine and I'm currently unemployed so no time constraints. Still, progress has been pretty slow due to various reasons but I haven't given up this time so I guess that's a plus.
← PreviousPage 2 of 9Next →