No server infra needed. Issues/Wiki/Discussions all built-in. Customizable web interface. Adapter to export to git. Single-file (also an SQLite database you can query).
It's obvious it's the superior stuff.
1,557 karma · joined May 8, 2012
No server infra needed. Issues/Wiki/Discussions all built-in. Customizable web interface. Adapter to export to git. Single-file (also an SQLite database you can query).
It's obvious it's the superior stuff.
Show me a social platform that is not complete bullshit. You can't.
I do agree with it, LinkedIn is bullshit. But c'mon, this is all the web is now.
We dislike _the presence_ of boilerplate, not the time spent writing it. If another thing writes it for you, it implies that _now boilerplate exists_, and it sucks.
It makes me unhappy when it exists. It makes me unhappy if it appears in seconds.
That said, there is some potential for using AI to reduce boilerplate and help create more meaningful software. However, that is definitely not the way things are shaping out to be.
It's feasible, and it makes more sense than separating and merging later (or keeping patches in memory then applying in bulk).
Why do I say it's feasible? We have the technology, right? The IDE knows which file the user has in focus, and can orient agents to use tooling that would inform them of that fact when they're running. Similarly, that same tooling could just spend a little bit of time planning focus to spread to multiple agents in a way they won't overlap.
Maybe big repos, monorepos and so on are a limitation. If we were on the previous "small-to-middle interlinked projects" era, that division would come in naturally. You only really need multiple agents in parallel on a single project if that thing is big enough to have more than one angle to work on. It's a push-and-pull that changes with the times, maybe we're heading to a more granular way of doing things.
The problem is, I always need more structure. Give me some YAML and time and I'll make hell (not a metaphor, I'll concoct hell itself on it).
Markdown keeps me honest.
Maybe we need a word that, when applied to mathematical concepts, describes how simple, easy to understand and generally useful a solution or idea is.
I wonder what that word could be.
I say "we need 100% coverage on that critical file". It runs for a while, tries to cover it, fails, then stops and say "Success! We covered 60% of the file (the rest is too hard). I added a comment.". 60% was the previous coverage before the LLM ran.
A math module that is not tested for division by zero. Classical LLM development.
The suite is mostly happy paths, which is consistent with what I've seen LLMs do.
Once you setup coverage, and tell it "there's a hidden branch that the report isn't able to display on line 95 that we need to cover", things get less fun.
So, where is that in the 2020s?
Yes, code is a detail (ideas too). It's a platform. It positions itself as the new thing. Does that platform allow upstarts? Or does it consolidate power?
The agent should look at my README.md, not a custom human-like text that is meant to be read by machines only.
It also should look at `Makefile`, my bash aliases and so on, and just use that.
In fact, many agents are quite good at this (Code Fast 1, Sonnet).
Issue is, we have a LONG debt around those. READMEs often suck, and build files often suck. We just need to make them better.
I see agents as an opportunity for making friendlier repos. The agent is a free usability tester in some sense. If it can't figure out by reading the human docs, then either the agent is not good enough or your docs aren't good enough.
Local models are a different thing than those cloud-based assistants and APIs.
I like rebasing! It works great for bisecting, reverting (squash messes that up), almost everything. It just doesn't play well with micro commits (which unfortunatelly have become the norm).
The force pushing to the PR branch is mostly a consequence of that rebase choice, in order to not pollute the main branch. Each change in main/master must be meaningful and atomic. Feature branches are other way to achieve this, but lots of steps involved.
I don't use it as inspiration. It's like I said: code that is not reviewed yet.
It takes the idea of 50 juniors working for you one step ahead. I manage the workflow in a way that they already made the code they wrote merge and build before I review it. When it doesn't, I delete it from the stash.
I could keep a branch for this. Or go even deeper on the temptation and keep multiple branches. But that's more of my throughput I have to spent on merging and ensuring things build after merging. It's only me. One branch, plus an extra "WIP". Stash is perfect for that.
Also, it's one level of stashing. It's stacked in the sense that it keeps growing, but it's not several `git stash pop`s that I do.
One thing that helps is that I already used this to keep stuff like automation for repos that I maintain. Stuff the owner doesn't want or isn't good enough to be reused. Sometimes it was hundreds of lines, now it's thousands.
Also, I stack the stash. When I vibe code, I pop it, let it work on its own mess, then I stash it again.
One project has almost 13.000 lines of vibe mess, all stashed.
One good thing, is that the stash builds. It's just that I don't want to release more code than I can read. It's a long review queue that is pre-merged somehow.
Once in a while I pick something from there, then I review it and integrate into the codebase more seriously. I don't have the throughput to review it all, and not all projects can be yolo'd.
https://github.com/alganet/PHL
---
Bootstrapping from an x86 image that is mostly source text (based on live-bootstrap):
https://github.com/alganet/abuild
---
Image with many shells, for testing script for portability:
[ Close with comment ]
I just explained that to you. Either we discuss this in terms of the imitation game thought experiment, or we don't.
It's not a psychology exercise, my dude.
You have two choices:
- It can potentially lash out in an alien-like way.
- It can potentially lash out in a human-like way.
Do you understand why this has no effect on the argument whatsoever? You are just introducing an irrelevant observation. I want the AI to behave like human always, no exceptions.
"What if it's a bad human"
Jesus. If people make an evil AI, then it doesn't matter anyway how it behaves, it's just bad even before we get to the discussion about how it fails. Even when it accomplishes tasks succesfully, it's bad.
It starts with "For the “do you detect an injected thought” prompt..."
If you Ctrl+F for that quote, you'll find it in the Appendix section. The subsection I'm questioning is explaining the grader prompts used to evaluate the experiment.
All the 4 criteria used by grader models are looking for a yes. It means Opus 4.1 never satisfied criterias 1 through 4.
This could have easily been arranged by trial and error, in combination with the selection of words, to make Opus perform better than competitors.
What I am proposing, is separating those grader prompts into two distinct protocols, instead of one that asks YES or NO and infers results based on "NO" responses.
Please note that these grader prompts use `{word}` as an evaluation step. They are looking for the specific word that was injected (or claimed to be injected but isn't). Refer to the list of words they chosen. A good researcher would also try to remove this bias, introducing a choice of words that is not under his control (the words from crosswords puzzles in all major newspapers in the last X weeks, as an example).
I can't just trust what they say, they need to show the work that proves that "Opus 4.1 never exhibits this behavior". I don't see it. Maybe I'm missing something.
That's why I decided to comment on the paper instead, which is supposed to outline how that conclusion was estabilished.
I could not find that in the actual paper. Can you point me to the part that explains this control experiment in more detail?