Quite drastic to move away from a country for just this. Did you only do it for just this? Or was this simply one of the factors why you moved away?
938 karma · joined September 26, 2019
melvinroest <the fancy a> <Google's email brand> <the most popular TLD in the world>
Some hints are gmail, @ and .com ;-)
---
Random HN'ers are always welcome to join, feel free to email me!
——-
AI engineer, security enthusiast, product engineer and marketing (data) analyst.
Quite drastic to move away from a country for just this. Did you only do it for just this? Or was this simply one of the factors why you moved away?
In a high trust culture I get this. But what about if you're in a low trust culture with quite some "not so orderly behavior" (if you will).
I feel this. When I first got to learn what a workspace was (in Eclipse, of all things), and I really drilled down into the English meaning of it (I'm Dutch), I thought about exactly what you're describing right now.
Personally, I use LLMs for a lot of things. Oftentimes, I'm a think out loud type of person so even having something that feels like a rubber duck, but more competent, is already amazing for me. And LLMs are a lot more competent than a rubber duck.
But especially sometimes I've noticed that LLMs can be unbelievably stupid. It recently happened a few times with Fable 5.1 as well. Ultimately, I think it comes down to that LLMs can't think broadly. In software development one can usually see this too. For example, a whole app might be built by an LLM and it didn't spend a single token thinking about security because the prompter is at the level of "build a dating app for dogs, make no mistakes". Now you have a dating app for dogs that is insecure.
Since I prompt for almost everything in my life to have an LLM as a sounding board, I'm usually not an expert either. I've noticed LLMs are amazing at "bulk search engine information aggregation" (or whatever you want to call it). So if I need something from the Dutch government, I can find it way more quickly. But oftentimes I've noticed that going for a walk and thinking about a particular thing I'm facing is a more effective way of finding a good solution.
Other times times they are not incredibly stupid, but can't form a strong opinion. This usually happens when I'm tackling a wicked problem [1]. When that's the case, prepare for LLMs to sway with you for every small change in your opinion that you ever will experience.
So I agree: drop in replacement for knowledge workers? No. Rigorous specification is usually needed yes. Though, the small win here is that it doesn't always need to be as rigorous as programming is and it can happen in natural language. It depends on the topic/problem being tackled.
I really like them as UX tools though. Amazing for interactive prototyping and requirements elicitation. And that also corresponds with what the author is saying. Though I find it a bit of a disservice saying "just 3". You know how hard requirements elicitation is? It became a whole lot easier thanks to LLMs (I might change this opinion in a year, haha, but this is the opinion I hold now).
Because if kids can go this far. Then how far could I go? How far could we all go?
Also: I'm having a blast, quite literally. Love raving while mathing, it's my new favorite hobby. Shout out to Nigel Good. Nigel sure know's what's good when EDM is considered.
[1] Having 100+ hours of good EDM tracks apparently motivates me a lot. There are papers that mention it's not the best for focus, but I've noticed that I don't do any math at all if I don't feel the warm synths that I love with my favorite playlists [3] (I take a lot of care in constructing them). I'm curious to see if this will keep up since I've only been going hard at it around late August.
[3] My main playlist: https://open.spotify.com/playlist/3BUgKCDAWvcDa28MVToIFM?si=...
Don't get me wrong, it's interesting. But there is no technical discussion as to how they did it. It's simply: we did it and Mythos and Codex didn't.
It's good to know that it's possible, but I'd have already expected it. Put a base model versus a base model + harness + whatever else, and yea, if you do it right then you have a better system to find vulnerabilities.
> We then ran AISLE's autonomous AI system against curl.
They don't even mention what models the use under the hood. It wouldn't surprise me if they are from Anthropic and OpenAI.
I've been calling LLMs digital intelligence. That's in my opinion what they are and on the digital/conceptual realm they are better than humans since LLMs are better generalists.
But I'd never state their conscious. For me consciousness extends outward from myself, since I can't trust anything else. It starts with the question like: why can't I control the movement of another person? Why can't I transfer my mind to another body and theirs to mine? Why is it that whenever I sleep, and similar things, I end up back to this place called "reality"?
I come to the conclusion that I am conscious and I'm a conscious being in reality. I am human and I've been raised by humans. Now, technically, all humans besides me could be zombies. They could be non-conscious beings that simply can act like humans. I'm sure we'll be able to create them in a few decades. However, I'm incentivized for multiple reasons to believe they're not zombies. Therefore humans are conscious too. All of them. Can I know for sure? No since I can't even trust my own senses fully or my own thoughts (the whole Descartes thing) but I choose to trust those things too because I need some reasonable-ish foundation to work from.
If humans are conscious, why aren't animals like us? Surely great apes must be conscious at least to some extent. Hell, some of them even have better short-term memory than us [1]. Well, turns out, from an anatomy standpoint they look a lot like us. Fine, let's assume they're conscious too.
Now we get into the territory that anything that has a large enough brain must be conscious because that's how we govern our consciousness. So you can extend this to all kinds of living beings.
I did this from my perspective. Of course, you should do it from your perspective. The same argument could be made.
For as long as we can't manipulate conscious experience in some way shape or form, we have no clue whether anything that isn't like us can be conscious. We'd need to be able to merge and split conscious experiences or we'd need to be very strongly able to understand why that isn't possible. And I don't mean just at the biological level but at the experiential level. Currently conjoined twins, and what they tell us gives some insight (being able to sense the other part that's part of the other twin).
Other than that, I claim we have no clue what consciousness is other than that we're experiencing it. So to call LLMs conscious is way too big of a claim. But they are definitely intelligent because they come up with things that I wouldn't have and it's useful to me. Perhaps a practical characterization of intelligence but it works for me.
I think for anyone to say something useful from this you'd need the trifecta of deep neuroscience knowledge, deep philosophy knowledge and (at least) a strong understanding of how LLMs work.
AI text feels more like "this matters, not less." There's always an "it's x not y" pattern somewhere.
It could be written AI assisted but then look at his account. IMO every HN account that was created before ChatGPT came out has incentive to be written by a human since they have a paper trail before the whole LLM thing happened.
This is at least written by a human, not AI, totally human, (digital) pinky promise ;-)
It was such a fun birthday. It was basically a conference, haha.
I don't, but holy moly. That sounds insane!
Whatever it is, I'm sure it's load-bearing.
To be fair, vibecoding this memo app in Swift didn’t take too long. There were some tricks to it, using xcodegen helped a lot so that I don’t need to use the Xcode project.
It’s fun to see Swift code. I used to do some Objective-C back in the day.
[1] another thing I made. It’s a sequel to the Alice in Wonderland stories. It’s also a SQL course. I vibe engineered it, meaning I looked at the code and used AI-assisted development.
Except for the story though that’s almost all fully me. LLMs aren’t great storytellers. The same is true for the lesson scaffolding, that’s almost only me.
So basically a way to just go on an hour long walk with myself, spit everything from the top of my dome stream of consciousness style, and then have Claude structure whatever I said.
It's nice to have something that structures my thoughts by just thinking out loud.
I vibecoded it (it's approaching 20K lines including tests). It works quite well but there are some bugs, so will have to do some actual engineering. But the UX is working quite well.
Make usable software. Cheap code means that you can create a lot more prototypes to then perform usability tests by finding a user and sitting next to them. I mostly worked on internal apps lately, so perhaps it's much easier for me to do than it is for some others.
> Cognitive Debt, Like Technical Debt, Must Be Repaid
In quite a few circumstances, cognitive debt doesn't entirely need to be repaid. I personally found with multiple projects that certain directions aren't the one I want to go in. But I only found it out after fully fleshing it out with Claude Code and then by using my own app realizing that certain things that I thought would work, they don't.
For example, I created library.aliceindataland.com (a narrative driven SQL course). After a while, I noticed that the grading scheme was off and it needed to be rewritten. The same goes for how I wanted to implement the cheatsheet, or lessons not following the standard format. Of course, I need to understand the new code but I don't need to understand the old code.
With other small forms of code, I just don't really need to know how things work because it's that simple. For example, every 5 minutes I track to which wifi network I'm connected with. It's mostly useful to simply know whether I went to the office that day or not. A python script retrieves the data and when I look at it, I can recognize that it's correct. But doing it this way is sure a lot faster than active recall.
At work, I've had similar things. At my previous job I created SEO and SEA tools for marketing experts. So I remember creating this whole app that gave experts insights into SEO things that Ahrefs and similar sites don't, as it was tailored to the data of the company I worked at. The feedback I basically got was: the data is great, the insights are necessary, but the way the app works is unusuable for us. I was a bit perplexed as I personally didn't find it that complicated. But I also know that I'm not the one using it. Then I created a second version and that was way more usable. The second version assumed a completely different front-end app and front-end app architecture though. All the cognitive debt of V1? No payback needed.
The reason that this is the case, as it seems to me, fall under a few categories:
1. Experimenting with technologies. If you have certain assumptions about how a technology works but it turns out you're wrong, or you learn through the process that an adjacent technology works way better, then you need to redo it. Back when coding by hand was such a thing, I had this with a collaborative drawing project called Doodledocs (2019). I didn't know if browsers supported pressure sensitivity and to what extent it was easy to implement. It required a few programming experiments.
2. It's a small and simple script, not much more to it.
3. Experimenting with usability. A lot of the time, we don't know how usable our app is. In my experience, this seems to be either because (1) it's a hobby project or (2) the UX people have been fired years ago. In these cases, more often than not, UX becomes an afterthought. But with LLMs, delivering a 95% fully working version is usually done within a week for a greenfield project. This 95% fully working version is an amazing high fidelity interaction prototype (95% no less). Once you do that for a few iterations, you then understand what you really need. Once you understand what you really need, then you can start repaying the cognitive debt.
I've found it's usually category 3, sometimes 2 and rarely 1.
I'm having fun writing a sequel on the books that Lewis Carroll wrote and mixing it with a SQL course. My hope is that SQL will be more fun to learn that way. And it's fun to write a few pages that will hopefully evoke some narrative transportation and immersion vibes.
I'm still very much at the beginning though.
In the story Alice enters an Infinite Library. You (yes you!) are STAR: a magical sentient typewriter that can only write in SQL queries. When Alice finds you, you'll slowly both find out why this library exists and the secrets that it holds.
Course-wise: I'm trying to have tight lesson scaffolding, which is a fun challenge.
----
Full reaction:
Yes but perhaps not in a way you might expect. Qwen's reasoning ability isn't exactly groundbreaking. But it's good enough to weave a story, provided it has some solid facts or notes. GraphRAG is definitely a good way to get some good facts, provided your notes are valuable to you and/or contain some good facts.
So the added value is that you now have a super charged information retrieval system on your notes with an LLM that can stitch loose facts reasonably well together, like a librarian would. It's also very easy to see hallucinations, if you recognize your own writing well, which I do.
The second thing is that I have a hard time rereading all my notes. I write a lot of notes, and don't have the time to reread any of them. So oftentimes I forget my own advice. Now that I have a super charged information retrieval system on my notes, whenever I ask a question: the graphRAG + LLM search for the most relevant notes related to my question. I've found that 20% of what I wrote is incredibly useful and is stuff that I forgot.
And there are nuggets of wisdom in there that are quite nuanced. For me specifically, I've seen insights in how I relate to work that I should do more with. I'll probably forget most things again but I can reuse my system and at some point I'll remember what I actually need to remember. For example, one thing I read was that work doesn't feel like work for me if I get to dive in, zoom out, dive in, zoom out. Because in the way I work as a person: that means I'm always resting and always have energy for the task that I'm doing. Another thing that it got me to do was to reboot a small meditation practice by using implementation intentions (e.g. "if I wake up then I meditate for at least a brief amount of time").
What also helps is to have a bit of a back and forth with your notes and then copy/paste the whole conversation in Claude to see if Claude has anything in its training data that might give some extra insight. It could also be that it just helps with firing off 10 search queries and finds a blog post that is useful to the conversation that you've had with your local LLM.
Recently I built a graphRAG app with Qwen 3.5 4b for small tasks like classifying what type of question I am asking or the entity extraction process itself, as graphRAG depends on extracted triplets (entity1, relationship_to, entity2). I used Qwen 3.5 27b for actually answering my questions.
It works pretty well. I have to be a bit patient but that’s it. So in that particular use case, I would agree.
I used MLX and my M1 64GB device. I found that MLX definitely works faster when it comes to extracting entities and triplets in batches.