HNHacker News
TopNewBestAskShowJobs

swimwiththebeat

116 karma · joined January 31, 2021

submissionscomments
swimwiththebeat··on The Inference Hardware Revolution of 2026
> Tensordyne is expected to accelerate AI inference with a logarithmic number system that leans on a property of logarithms: The log of A times B equals the log of A plus the log of B. So, storing numbers as their exponents lets the chip add where it would otherwise multiply. That matters in silicon because multiplier circuits draw more power and use more die area than adders do. Tensordyne says its rack-scale hardware, called Napier, can produce up to 1,300 tokens per second per user, and can do so while using less than a tenth as much power as comparable Nvidia hardware.

Did not know about this cool trick about storing numbers as exponents! Is there a name for this technique? Wouldn’t there be overhead in converting back and forth between the exponent and the number?

swimwiththebeat··on OpenArch – PyTorch implementations of modern LLM architectures
This is really cool, great way to reinforce our understanding of model architectures! But how is the author confirming that these model architecture implementations are correct though? I don't see any details in the README.md.
swimwiththebeat··on The Mysterious Syndrome Destroying Endurance Athletes
I definitely experienced this in training for my marathon and half Ironman (which I had stupidly signed up for in consecutive months). I didn't follow my training plan properly and tried to cram all the volume in the last month or two and it didn't end well. I was able to make it through the half Ironman through pure adrenaline, but I laid in bed for the following two weeks and barely moved. I had no motivation to do anything whatsoever.

It turns out that I was deficient in plenty of vitamins: D, B12, iron. I'm sure these were all symptoms of OTS. I didn't feel the same in any of my workouts until maybe 4 months later. Even then, it took very gradual build-up after those 4 months to get back to where I was. Taking vitamin supplements helped, sure, but the true cure was just ramping down the exercise a lot and getting proper sleep over months.

I empathize with these athletes because exercise is addicting; it feels almost morally wrong to take days off. There's also always someone better than you. But as I got older, I realized that recovery is just as important (if not, more important) than the actual work. In fact, that's where all the gains are made.

I found a few helpful pointers with structuring my workouts now is to:

* Increase training volume <10% W/W. It gives your body time to adapt and also helps with preventing injury.

* Mixing in recovery weeks is also really helpful with helping your body adapt.

* There's no way around doing a lot of volume when training for endurance events, but keeping your HR lower the whole time is a good way to make sure you don't overtrain and can show up the next day.

swimwiththebeat··on Kimi K3: Open Frontier Intelligence
Did anyone see on the blog post[0] that it was able to code up an entire GPU compiler from scratch? It looks like it even outperformed triton on some GPU kernels. That just seems insane to me.

Wonder if they’ll open-source this and show how many tokens it cost.

[0] https://www.kimi.com/blog/kimi-k3

swimwiththebeat··on A visual introduction to big O notation
Your blog always has such informative visualizations that make these concepts easy to digest! Which tools do you use to create them?
swimwiththebeat··on Matrix-vector multiplication implemented in off-the-shelf DRAM for Low-Bit LLMs
So is this a new technique of doing computations within existing DRAM to overcome the memory wall issue of modern computing?
swimwiththebeat··on The Nvidia Way
I just finished this book. I thought it was a great read, capturing most of the problems (with product, sales, marketing, competition) Nvidia faced in its life and how exactly the team solved them each and every time. It also goes into how the management, org structure, and culture have empowered its engineers to find creative solutions to staying alive in the brutal graphic chips industry and continuously discover and exploit new market opportunities.

I learned a lot about how important not just superior technology, but better operations, marketing, sales, and culture are all critical to a successful business.

The only con of this book is that it skips over some parts of nvidia’s history like the short-lived crypto boom, failed acquisition of ARM, etc. It’s still just a minor flaw in an otherwise great book though.

swimwiththebeat··on Show HN: DaLMatian – Text2sql that works
Is this open-source?
swimwiththebeat··on OpenAI’s Sora made me crazy AI videos then the CTO answered most of my questions
She definitely knows, she’s just trying to avoid any chance of future litigation by feigning ignorance. Makes sense since OpenAI’s been getting a lot of bad press for using copyright data in training their models.
swimwiththebeat··on Generative AI and the big buzz about small language models
Does anyone know if this is using the Mamba architecture[1] instead of transformers? It looks like it uses a state space model (SSM) layer.

[1]: https://arxiv.org/abs/2312.00752

swimwiththebeat··on Learning resources for curious software engineers
As someone who wants to learn and explore more areas of computer science and programming, this outline and centralized list of resources is incredibly helpful! Thanks so much.
swimwiththebeat··on Neal Stephenson was prescient about our AI age
I agree, I don't think the purpose of sci-fi is to predict the future. The future is just impossible to predict due to the myriad factors, variables, unknown unknowns, and second-order+ degree effects an action can have.

The purpose of sci-fi IMO is moreso to:

1. Provide an entertaining story/narrative with technology as the main focus of the world and characters' actions

2. Define a set of concepts to help you think about technology and its possible effects on humans and the world

3. Nudge people to think about what kind of future they would want or not want and how they can use or control technology to achieve that

Here's Ken Liu talking about the purpose of sci-fi: https://www.youtube.com/watch?v=5knkpmxXu-k

swimwiththebeat··on An open source DuckDB text to SQL LLM
Thanks for taking the time to answer the questions and link those resources, really appreciate it and the work your team did!
swimwiththebeat··on An open source DuckDB text to SQL LLM
Thanks for replying, that's a perspective I didn't consider. The capability to "talk to your data" just seems so enticing as a solution that I was tunnel-visioned into that UX. If I'm understanding correctly, what you're suggesting is more of a SQL assistant to help people write the correct SQL queries instead of writing the entire SQL query from scratch to answer a generic natural-language question?
swimwiththebeat··on An open source DuckDB text to SQL LLM
1. First of all, thanks for outlining how you trained the model here in the repo: https://github.com/NumbersStationAI/DuckDB-NSQL?tab=readme-o...! I did not know about `sqlglot`, that's a pretty cool lib. Which part of the project was the most challenging or time-consuming: generating the training data, the actual training, or testing? How did you iterate, improve, and test the model?

2. How would you suggest using this model effectively if we have custom data in our DBs? For example, we might have a column called `purpose` that's a custom defined enum (i.e. not a very well-known concept outside of our business). Currently, we've fed it in as context by defining all the possible values it can have. Do you have any other recs on how to tune our prompts so that this model is just as effective with our own custom data?

3. Similar to above, do you know you can use the same model to work effectively on tens or even hundreds of tables? I've used multiple question-SQL example pairs as context, but I've found that I need 15-20 for it to be effective for even one table, let alone tens of tables.

swimwiththebeat··on An open source DuckDB text to SQL LLM
I've actually done exactly what @qsort suggested and outputted the intermediate SQL query and raw data generated by that query when generating the response back to the user. That definitely helps in establishing more trust with the customer since they can verify the response. My approach right now is to just be honest with our customers in the capabilities of the tool, acknowledge its shortcomings, and keep iterating over time to make it better and better. That's what the team in charge of our company-wide custom LLM has done and it's gained a surprising amount of traction and trust over the last few months.
swimwiththebeat··on An open source DuckDB text to SQL LLM
I see so many business leaders touting the promise of LLMs allowing business to "talk" to their data. The promise does sound enticing, but it's actually kind of hard to get working in practice.

A lot of our databases at work have columns with custom types and enums, and getting the LLM (Llama2) to write SQL queries to robustly answer natural language questions about the data is tough. It requires a lot of instruction prompting, context, and question-SQL examples (few-shot learning), and it still fails in unexpected ways. It's a tough ask for people to use a tool like this if they can't trust the results all the time. It's also a bit infeasible to scale this to tens or hundreds of tables across our data warehouse.

It's great that a lot of people are trying to crack this problem, I'm curious to try this model out. I'd also love to see if other people have tried solving this problem and made any headway.

swimwiththebeat··on Vanna.ai: Chat with your SQL database
I'm curious to see if people have tried this out with their datasets and seen success? I've been using similar techniques at work to build a bot that allows employees internally to talk to our structured datasets (a couple MySQL tables). It works kind of ok in practice, but there are a few challenges:

1. We have many enums and data types specific to our business that will never be in these foundation models. Those have to be manually defined and fed into the prompt as context also (i.e. the equivalent of adding documentation in Vanna.ai).

2. People can ask many kinds of questions that are time-related like 'how much demand was there in the past year?'. If you store your data in quarters, how would you prompt engineer the model to take into account the current time AND recognize it's the last 4 quarters? This has typically broken for me.

3. It took a LOT of sample and diverse example SQL queries in order for it to generate the right SQL queries for a set of plausible user questions (15-20 SQL queries for a single MySQL table). Given that users can ask anything, it has to be extremely robust. Requiring this much context for just a single table means it's difficult to scale to tens or hundreds of tables. I'm wondering if there's a more efficient way of doing this?

4. I've been using the Llama2 70B Gen model, but curious to know if other models work significantly better than this one in generating SQL queries?

swimwiththebeat··on Analysis of 200M newspaper pages: Sentiment has collapsed over the past 50 years
I can think of a few reasons for this besides the reason that life has genuinely become worse:

1. Profit-seeking news orgs: it's no secret that negative and controversial news sells significantly more. With all the destruction of locally run papers in favor of the national consolidation of news by a lot of private equity investors, there are increased expectations to make a profit and therefore increase the amount of negative-sentiment news over the recent years.

2. 24-7 news cycle: people are also much more aware of all the bad news around them with smartphones and social networks. Readers' sentiment will just bleed into the papers over time. Ignorance is bliss.

3. Sampling bias: which kind of papers did they measure sentiment for back then versus now? There could be a divergence in the sources they use and different sources could have different sentiment tendencies. (I don't have access to the actual paper)

4. Rising expectations: humans are many orders of magnitude more powerful than our ancestors. We live like gods compared to them. Yet there are so many people who still aren't happy. Why? It's because our expectations also rise endlessly. Things may be better than before, but maybe it's not better relative to our expectations.

Point is, this isn't necessarily indicative of life becoming worse. There could be other plausible explanations.

swimwiththebeat··on Sam Altman, Greg Brockman and others to join Microsoft
I thought for sure the only two outcomes were that Altman raises money for a new startup or he comes back to OpenAI with a new governance structure (which is still a wild and crazy outcome, but crazier things have happened). Now that this happened though, I feel stupid for not considering this as a possible outcome at all.

The whole timeline of events over the last two events still leaves me scratching my head though.

swimwiththebeat··on Revenge bedtime procrastination
It’s amazing how much attaching a name to some phenomenon or problem and discussing it helps in understanding it and taking the first steps to address it. Knowing that bedtime procrastination is a real thing that affects others not only gives me comfort, but also helps me think about it very consciously and figure out how to take the first steps to solve it.

I never thought about it from the revenge and agency perspective, but I appreciate the author at least trying to hypothesize why people do this despite it being detrimental. All the suggestions and tips in the article are also incredibly helpful.

swimwiththebeat··on OpenAI's board has fired Sam Altman
> Mr. Altman’s departure follows a deliberative review process by the board, which concluded that he was not consistently candid in his communications with the board, hindering its ability to exercise its responsibilities. The board no longer has confidence in his ability to continue leading OpenAI.

Whoa, rarely are these announcements so transparent that they directly say something like this. I’m guessing there was some project or direction Altman wanted to pursue, but he was not being upfront with the board about it and they disagreed with that direction? Or it could just be something very scandalous, who knows.

swimwiththebeat··on Show HN: MicroTCP, a minimal TCP/IP stack
How did you even get started with this? There are many systems I'd love to learn to implement from scratch like databases, network stacks, and caches, but the task seems so vast and daunting that it's hard to get started. Learning from existing codebases is also difficult b/c there's so much code to go over and understand. Can you elaborate on what your process was to build this without any help?
swimwiththebeat··on X Engineering Year Retrospective
I'm inclined to believe that they consolidated their code across a lot of their microservices and simplified their architecture since that was stressed a lot from the moment Musk acquired Twitter. But we also can't really verify or disprove it.

I'm a bit confused about their so-called improvements to video recommendation quality and bot detection. I've seen a lot of sentiment from people that they see more bots, hate speech, and irrelevant content on their timelines. Maybe what I'm hearing is just anecdotal evidence or stories in a bubble?

The Sacramento data center migration to Portland is an entertaining story detailed here[1]. Here's the Hacker News thread on it[2].

They have a GPU supercompute cluster?? It seems like they have the capability to do training and inference with state-of-the-art algorithms at massive scales then. Why have Twitter's recommendations and ad revenue (even before the acquisition) been so poor then?

[1] https://www.cnbc.com/2023/09/11/elon-musk-moved-twitter-serv...

[2] https://news.ycombinator.com/item?id=37470110

swimwiththebeat··on Halfsies
92% with 4 perfect cuts (messed up badly on the easy 12-sided shape). Don't know why, but it was a strangely fun exercise. "Brilliant" ad too (pun intended :D)