Afaik there are no open source projects that do this. AWS has a behemoth of a distributed system you can deploy in order to do something similar. But I made a Python script that does it in an afternoon with a couple of prompts.
Why? You don't trust a newly-created account that has not engaged with any of the comments to be anything but truthful?
I think people enjoy writing code for various reasons. Some people really enjoy the craft of programming and thus dislike AI-centric coding. Some people don't really enjoy programming but enjoy making money or affecting some change on the world with it, and they use them as a tool. And then some people just like tinkering and building things for the sake of making stuff, and they get a kick out of vibe coding because it lets them add more things to their things-i-built collection.
But the payoff for letting that go is huge.
You still have to do this, LLM's are still quite bad at choosing the right data structures and dealing with their interdependent relationships.
I have integrated Claude Code with a graph database to support an assistant with structured memory and many helpful capabilities.
I have clients. I automated a complicated data ingestion pipeline into a desktop app with a bulletproof process queue, localhost control panel and many features.
For another, I am writing an AI-specific app that is so cool. I wish I could tell you about it but it's definitely not a rushed remake of anything.
I hope that helps.
Is down. And the scoring one, no offense, seems like a project a junior would make to pad out their resume/portfolio. Nothing wrong with that of course, but I fail to see how this translates to all the hype being thrown around.
It's the kind of thing that would be hours of tedious work, then even more time to actually make all the changes to the account. Instead I just say "yeah do all of that" and it is done. Magic stuff. Thousands of lines of Python to hit the Amazon APIs that I've never even looked at.
I wouldn't trust thousands of lines of code from one of my co-workers without testing
And yes, I have occasionally run into compiler bugs in my career. That's one reason we test.
How did you verify that?
> prone to hallucination
You know humans can hallucinate?
> is perfectly deterministic
We agree then that you can verify, test, and trust the deterministic code an LLM produces without ever looking at it.
> That's one reason we test
That's one way we can trust and verify code produced by an LLM. You can't stop doing all the other things that aren't coding.
I get there's a difference. Shitty code can be produced by LLMs or humans. LLMs really can pump out the shitty code. I just think the argument that you cant trust code you haven't viewed is not a good argument. I very much trust a lot of code I've never seen, and yes I've been bitten by it too.
Not trying to be an ass, more trying to figure out how im going to deal for the next decade before retirement age. Uts going to be a lot of testing and verification I guess
The compiler works without an internet connection and requires too little resources to be secretly running a local model. (Also, you can’t inspect the source code.)
> You know humans can hallucinate?
We are talking about compilers…
> We agree then that you can verify, test, and trust the deterministic code an LLM produces without ever looking at it.
Unlike a compiler, an LLM does not produce code in a deterministic way, so it’s not guaranteed to do what the input tells it to.
but second of all, even when error rates were 20%, the time savings still meant A Viable Business. a much more viable business actually, a scarily crazy viable business with many annoyed customers getting slop of some sort, with a human in the loop correcting things from the LLM before it went out to consumers
agentic LLM coders are better than your co-workers. they can also write tests. they can do stress testing, load testing, end to end testing, and in my experience that's not even what course corrects LLMs that well, so we shouldn't even be trying to replicate processes made for humans with them. like a human, the LLM is prone to just correct the test as the test uses a deprecated assumption as opposed to product changes breaking a test to reveal a regression.
in my experience, type errors, compiler errors, logs on deployment and database entries have made the LLM correct its approach more than tests. Devops and Data science, more than QA.
Me? I use AI to write tests just as I use it to write everything else. I pay a lot of attention to what's being done including code quality but I am no more insecure about trusting those thousands of tested lines than I am about trusting the byte code generated from the 'strings of code'.
We have just moved up another level of abstraction, as we have done many times before. It will take time to perfect but it's already amazing.
So they don't know if it has the right behavior to begin with, or even if the tests are testing the right behavior.
This is what people are talking about. This is why nobody responsible wants to uberscale a serious app this way. It's ridiculous to see so much hype in this thread, people claiming they've built entire businesses without looking at any code. Keep your business away from me, then.
I was a product manager for 15 years. I helped sell products to customers who paid thousands or millions of dollars for them. I never looked at the code. Customers never looked at the code. The overwhelming majority of people in the world are constantly relying on code they've never looked at. It's mostly fine.
> How do you verify the end result?
That's the better question, and the answer is a few things. First, when it makes changes to my ad accounts, I spot check them in the UI. Second, I look at ad reporting pretty often, since it's a core part of running my business. If there were suddenly some enormous spike in spend, it wouldn't take me long to catch it.
https://hippich.github.io/minesweeper/ - no idea why but i had a couple weeks desire to play minesweeper. at some point i wanted to get a way to quickly estimate probability of the mine presence in each cell.. No problem - copilot coded both minesweeper and then added probabilities (hidden behind "Learn" checkbox) - Bonus, my wife now plays game "made" by me and not some random version from Play store.
another one made in a day - https://hippich.github.io/OpenCamber - I am putting together old car, so will need to align wheels on it at some point. There is Gyraline, but it is iOS only (I think because precision is not good enough on Android?). And it is not free. I have no idea how well it will work in practice, but I can try it, because the cost of trying it is so low now!
yes, both of these are not serious and fun projects. unlikely to have any impact. but it is _fun_! =)
- A "semantically enhanced" epub-to-markdown converter
- A web-based Markdown reader with integrated LLM reading guide generation (https://i.imgur.com/ledMTXw.png)
- A Zotero plugin for defining/clarifying selected words/sentences in context
- An epub-to-audiobook generator using Pocket TTS
- A Diddy Kong Racing model/texture extractor/viewer (https://i.imgur.com/jiTK8kI.png)
- A slimmed-down phpBB 2 "remake" in Bun.js/TypeScript
- An experimental SQLite extension for defining incremental materialized views
...And many more that are either too tiny, too idiosyncratic, or too day-job to name here. Some of these are one-off utilities, some are toys I'll never touch again, some are part of much bigger projects that I've been struggling to get any work done on, and so on.
I don't blame you for your cynicism, and I'm not blind to all of the criticism of LLMs and LLM code. I've had many times where I feel upset, skeptical, discouraged, and alienated because of these new developments. But also... it's a lot of fun and I can't stop coming up with ideas.
Every time I've asked people about what the hell they're actually doing with AI, they vanish into the ether. No one posts proof, they never post a link to a repo, they don't mention what they're doing at their job. The most I ever see is that someone managed to vibe code a basic website or a CRUD app that even a below-average engineer can whip up in a day or two.
Like this entire thread is just the equivalent of karma farming on Reddit or whatever nonsense people post on Facebook nowadays.
I have integrated Claude Code with a graph database to support an assistant with structured memory and many helpful capabilities.
I have a freelance gig with a startup adapting AI to their concept. I have one serious app under my belt and more on the way.
Concrete enough?
this site gets indexed
there are too many disincentives to cater specifically to your suspicion and cynicism