HNHacker News
TopNewBestAskShowJobs

LASR

2,507 karma · joined January 9, 2015

hn@sidbala.com
submissionscomments
LASR··on Show HN: FastGraphRAG – Better RAG using good old PageRank
It's slow. So we use hypothetical mostly for async experiences.

For live experiences like chat, we solved it with UX. As soon as you start typing the words of a question into the chat box, it does the FTS search and retrieves a set of documents that have word-matches, scored just using ES heuristics (eg: counting matching words etc)

These are presented as cards that expand when clicked. The user can see it's doing something.

While that's happening, also issue a full hyde flow in the background with a placeholder loading shimmer that loads in the full answer.

So there is some dead-time of about 10 seconds or so while it generates the hypothetical answers. After that, a short ~1 sec interval to load up the knowledge nodes, and then it starts streaming the answer.

This approach tested well with UXR participants and maintains acceptable accuracy.

A lot of the times, when looking for specific facts from a knowledge base, just the card UX gets an answer immediately. Eg: "What's the email for product support?"

LASR··on Llama 3.1 405B now runs at 969 tokens/s on Cerebras Inference
What you can do with current-gen models, along with RAG, multi-agent & code interpreters, the wall is very much model latency, and not accuracy any more.

There are so many interactive experiences that could be made possible at this level of token throughput from 405B class models.

LASR··on Show HN: FastGraphRAG – Better RAG using good old PageRank
So I've done a ton of work in this area.

Few learnings I've collected:

1. Lexical search with BM25 alone gives you very relevant results if you can do some work during ingestion time with an LLM.

2. Embeddings work well only when the size of the query is roughly on the same order of what you're actually storing in the embedding store.

3. Hypothetical answer generation from a query using an LLM, and then using that hypothetical answer to query for embeddings works really well.

So combining all 3 learnings, we landed on a knowledge decomposition and extraction step very similar to yours. But we stick a metaprompter to essentially auto-generate the domain / entity types.

LLMs are naively bad at identifying the correct level of granularity for the decomposed knowledge. One trick we found is to ask the LLM to output a mermaid.js mindmap to hierarchically break down the input into a tree. At the end of that output, ask the LLM to state which level is the appropriate root for a knowledge node.

Then the node is used to generate questions that could be answered from the knowledge contained in this node. We then index the text of these questions and also embed them.

You can directly match the user's query from these questions using purely BM25 and get good outputs. But a hybrid approach works even better, though not by that much.

Not using LLMs are query time also means we can hierarchically walk down the root into deeper and deeper nodes, using the embedding similiarity as a cost function for the traversal.

LASR··on OpenAI, Google and Anthropic are struggling to build more advanced AI
Question for the group here: do we honestly feel like we've exhausted the options for delivering value on top of the current generation of LLMs?

I lead a team exploring cutting edge LLM applications and end-user features. It's my intuition from experience that we have a LONG way to go.

GPT-4o / Claude 3.5 are the go-to models for my team. Every combination of technical investment + LLMs yields a new list of potential applications.

For example, combining a human-moderated knowledge graph with an LLM with RAG allows you to build "expert bots" that understand your business context / your codebase / your specific processes and act almost human-like similar to a coworker in your team.

If you now give it some predictive / simulation capability - eg: simulate the execution of a task or project like creating a github PR code change, and test against an expert bot above for code review, you can have LLMs create reasonable code changes, with automatic review / iteration etc.

Similarly there are many more capabilities that you can ladder on and expose into LLMs to give you increasingly productive outputs from them.

Chasing after model improvements and "GPT-5 will be PHD-level" is moot imo. When did you hire a PHD coworker and they were productive on day-0 ? You need to onboard them with human expertise, and then give them execution space / long-term memories etc to be productive.

Model vendors might struggle to build something more intelligent. But my point is that we already have so much intelligence and we don't know what to do with that. There is a LOT you can do with high-schooler level intelligence at super-human scale.

Take a naive example. 200k context windows are now available. Most people, through ChatGPT, type out maybe 1500 tokens. That's a huge amount of untapped capacity. No human is going to type out 200k of context. Hence why we need RAG, and additional forms of input (eg: simulation outcomes) to fully leverage that.

LASR··on New iMac with M4
I've been looking for an upgrade from my 2015 5k iMac 27.

Might finally pull the trigger on this version. What I will miss is the 27 inch 5k display.

Also, my use case is exactly what this iMac is meant for - shared family computer that takes up little space in the kitchen.

LASR··on Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
That's exactly it.

I've been peddling my vision of "AI automation" for the last several months to acquaintances of mine in various professional fields. In some cases, even building up prototypes and real-user testing. Invariably, none have really stuck.

This is not a technical problem that requires a technical solution. The problem is that it requires human behavior change.

In the context of AI automation, the promise is huge gains, but when you try to convince users / buyers, there is nothing wrong with their current solutions. Ie: There is no problem to solve. So essentially "why are you bothering me with this AI nonsense?"

Honestly, human behavior change might be the only real blocker to a world where AI automates most of the boring busy work currently done by people.

This approach essentially sidesteps the need to have effect a behavior change, at least in the short-term while AI can prove and solidify its value in the real-world.

LASR··on Computer use, a new Claude 3.5 Sonnet, and Claude 3.5 Haiku
This is actually a huge deal.

As someone building AI SaaS products, I used to have the position that directly integrating with APIs is going to get us most of the way there in terms of complete AI automation.

I wanted to take at stab at this problem and started researching some daily busineses and how they use software.

My brother-in-law (who is a doctor) showed me the bespoke software they use in his practice. Running on Windows. Using MFC forms.

My accountant showed me Cantax - a very powerful software package they use to prepare tax returns in Canada. Also on Windows.

I started to realize that pretty much most of the real world runs on software that directly interfaces with people, without clearly defined public APIs you can integrate into. Being in the SaaS space makes you believe that everyone ought to have client-server backend APIs etc.

Boy was I wrong.

I am glad they did this, since it is a powerful connector to these types of real-world business use cases that are super-hairy, and hence very worthwhile in automating.

LASR··on A step toward fully 3D-printed active electronics
I am a fan of 3D printing. And I think you can probably get some circuit traces 3D printed for some niche applications.

But active electronics? That's a huge stretch. But more importantly, the economics just doesn't make sense. Components already cost fractions of a cent. Small-run PCB prototyping is like <$25 for 5 boards or so.

"A step toward..."

Maybe. But why?

LASR··on Swarm, a new agent framework by OpenAI
The problem with agents is divergence. Very quickly, an ensemble of agents will start doing their own things and it’s impossible to get something that consistently gets to your desired state.

There are a whole class of problems that do not require low-latency. But not having consistency makes them pretty useless.

Frameworks don’t solve that. You’ll probably need some sort of ground-truth injection at every sub-agent level. Ie: you just need data.

Totally agree with you. Unreliability is the thing that needs solving first.

LASR··on Do AI companies work?
Yes. Numbers / math is pretty much instant hallucination.

But. Try this approach instead: have it generate python code, with print statements before every bit of math it performs. It will write pretty good code, which you then execute to generate the actual answer.

Simpler example: paste in a paragraph of text, ask it to count the number of words. The answer will be incorrect most of the time.

Instead, ask it to out each word in the text in a numbered list and then output the word count. It will be correct almost always.

My anecdotal learning from this:

LLMs are pretty human-like in their mental abilities. I wouldn't be able to simply look at some text and give you an accurate word count. I would point my finger / cursor to every word and count up.

The solutions above are basically giving LLMs some additional techniques or tools, very similar to how a human may use a calculator, or count words.

In the products we've built, there is an AI feature that generates aggregations of spreadsheet data. We have a dual unittest & aggregator loop to generate correct values.

The first step is to generate some unittests. And in order to generate correct numerical data for unittests, we ask it to write some code with math expressions first. We interpret the expressions, and paste it back into the unittest generator - which then writes the unittests with the correct inputs / outputs.

Then the aggregation generator then generates code until the generated unittests pass completely. Then we have the code for the aggregator function that we can run against the spreadsheet.

Takes a couple of minutes, but pretty bulletproof and also generalizable to other complex math calculations.

LASR··on Do AI companies work?
I lead an applied AI research team where I work - which is a mid-sized public enterprise products company. I've been saying this in my professional circles quite often.

We talk about scaling laws, superintelligence, AGI etc. But there is another threshold - the ability for humans to leverage super-intelligence. It's just incredibly hard to innovate on products that fully leverage superintelligence.

At some point, AI needs to connect with the real world to deliver economically valuable output. The ratelimiting step is there. Not smarter models.

In my mind, already with GPT-4, we're not generating ideas fast enough on how best to leverage it.

Getting AI to do work involves getting AI to understand what needs to be done from highly bandwidth constrained humans using mouse / keyboard / voice to communicate.

Anyone using a chatbot already has felt the frustration of "it doesn't get what I want". And also "I have to explain so much that I might as well just do it myself"

We're seeing much less of "it's making mistakes" these days.

If we have open-source models that match up to GPT-4 on AWS / Azure etc, not much point to go with players like OpenAI / Anthropic who may have even smarter models. We can't even use the dumber models fully.

LASR··on Tesla Full Self Driving requires human intervention every 13 miles
This is exactly how I feel about FSD in my Model 3.

It’s like supervising a teenager learning to drive, ready to take over at anytime.

It’s a lot less stressful to just drive myself.

LASR··on The Palletrone is a robotic hovercart for moving stuff anywhere
The only useful part of this might be the mid-air battery-swap tech.

A drone with an effectively unlimited runtime would be pretty useful in some narrow applications.

But as a pallet? Come on. Cool technical ideas. But terrible applications.

LASR··on Ask HN: Should I quit my project and move on?
> I’d die of embarrassment and I’d just be sad about wasting my time.

Welcome to entrepreneurship!

Doing something hard knowing very well that it's most likely is going to be a waste of time is just part of the game.

Every once in a while, you may strike gold and make it big. But a lot of the times, you strike silver and it's just promising enough to keep you going.

Striking silver is the exact "waste of time" failure mode you want to avoid. I speak from experience when I say that I blew ~3 years of my career chasing after a product idea that hovered just at breakeven. In hindsight, I should have given up in the first 3 months and moved on to something else.

But you won't know until you release something to users. Don't give up before you start, obviously.

LASR··on Leveraging AI for efficient incident response
We've shifted our oncall incident response over to mostly AI at this point. And it works quite well.

One of the main reasons why this works well is because we feed the models our incident playbooks and response knowledge bases.

These playbooks are very carefully written and maintained by people. The current generation of models are pretty much post-human in following them, performing reasoning and suggesting mitigations.

We tried indexing just a bunch of incident slack channels and result was not great. But with explicit documentation, it works well.

Kind of proves what we already know, garbage in, garbage out. But also, other functions, eg: PM, Design have tried automating their own workflows, but doesn't work as well.

LASR··on Intel Core Ultra 9 285K
It’s just harder to produce significant gains now. A 10% might be the result of all the chip giants giving it their absolute best.

I wouldn’t draw conclusions about the level of competition simply based on the spread in performance.

Remember that with Olympic runners the spread is 100ths of a second. And yet the competition is fierce.

You may not be excited about a 10% improvement. I am not either. But that’s a different matter.

LASR··on SanDisk introduces the first 8TB SD and 4TB microSD cards
Somewhat off-topic, but it blows me away every time it's brought up that the amount of information you can store in a volume scales by the surface area covering that volume and not the volume itself.
LASR··on SanDisk introduces the first 8TB SD and 4TB microSD cards
What do you mean?

Digital storage is kind of pointless if we cannot know exactly which bit came from what logical location.

ie: it's not like analog audio tape, where we get some signal, but don't really know its exact position.

LASR··on Microsoft says OpenAI is now a competitor in AI and search
With Meta giving out models for free, I wonder what competition even means in this space now.
LASR··on CrowdStrike global outage to cost US Fortune 500 companies $5.4B
I’m pretty certain CS has contracts that limit their liabilities in events like this.

Probably a refund is all they’ll be on the hook for.

Sadly, damage done like this is just chalked up to an accident, and swept under the rug.

LASR··on New Recovery Tool to help with CrowdStrike issue impacting Windows endpoints
Signing is meant only to verify the identity of the organization producing the signed artifact.

It’s not meant to signify that it’s bug-free.

LASR··on Ant Design – the second most popular React UI framework
I was on a different timezone on holiday, and I got woken up with some panic/anger calls direct from the CEO that we had been hacked. It took several hours of deep-diving, calling up security, backend, frontend engineers until we realized this behavior was due to the UI library.

The mind boggles, how someone thought this was a good idea.

LASR··on Rabbit data breach: all r1 responses ever given can be downloaded
You know what I am truly terrified of? When these AI services start keeping memories about your interactions with them to build up a full profile of you. Like chatGPT memories. But then it leaks due to a data breach.

It’s not even accurate sometimes and I definitely did not manually tell it things about me. But it made some incorrect assumptions and now it’s out there whether it’s true or not.

LASR··on Content Injection Attack on GitHub
So this opened in my GitHub iOS app at first and I was confused.
LASR··on I am sick of LeetCode-style interviews
I lead tech interviews for E5/6 level engs where I work.

Everyone knows how pointless it is. But it’s become more like paying the tax.

For the record, I typically give away the a-ha moment early on and see how the candidate is able to take the suggestion and work through a problem with me.

LASR··on Magic UI: UI Library for Design Engineers
Agreed. It's just so much easier to hire for React. It has become the lingua franca of client frameworks.

There are still many teams that choose to use something else. But experience in React is most easily transferrable. Even if I had a project in Vue, I would still look for React expertise.

LASR··on Show HN: Pls Fix – Hire big tech employees to appeal account suspensions
Yeah I don’t know about this. Looks like you’re opening yourself up for a TON of liability with some very powerful businesses.

But not only that, you’re likely endangering the jobs of the people who take up your offers.

Seems icky. I am all about generating side-income streams by connecting people with needs to those who can fulfill them. But this is not it. I would recommend you take whatever you’ve built and apply it to a different problem.

What about monetizing Reddit AMAs? Things like that are a lot more useful and apply to a much wider range of people.

LASR··on I want flexible queries, not RAG
I continue to hold the strong position that calling LLMs without injecting source truths is pointless.

LLMs are exceptionally powerful as a reasoning engine. It’s useless as a source of truths or facts.

We have chat bots, chat bots with automatic RAG etc. After the initial excitement wears off, you’re going to want a way to inspect and adjust the source queries yourself. In this case, being able to select what to search for in Google might be a good way for the cooking recipe usecase.

LASR··on Multi AI agent systems using OpenAI's assistants API
We’ve tried. A lot. Custom frameworks and all.

There is really no way to make the ensemble behave with an acceptable level of consistency.

Where we ended up is now having a frontier model generate a whole tree of possible execution plans, and then have the user select one of those path, and then we just run whatever the user chose in a plain sequence until the next decision point that needs user approval.

LASR··on Apple apologizes for iPad 'Crush' ad that 'missed the mark'
Yeah what was that ad even.

I am not a creative. But I do play the piano from time to time. It’s an old 15 year old Roland electric piano. I wouldn’t like to see it crushed. Even if it is obsolete. I bet a lot of actual creatives do have sentimental values attached to their tools.

Destroying things needlessly is very much off brand for Apple.

← PreviousPage 2 of 13Next →