HNHacker News
TopNewBestAskShowJobs

kingcauchy

107 karma · joined December 30, 2024

submissionscomments
kingcauchy··on What I learnt co-leading an AI Safety bootcamp for legal and governance practit
link in the comments of the article is also interesting https://www.lesswrong.com/posts/bfwH88cFEm3r7PWfv/what-lawye...
kingcauchy··on I quit OpenAI because its culture is broken
Like farming, engines, computers before it.
kingcauchy··on We're going to need default hard budget caps on pretty much everything
Probably should also have spend controls for gambling too but seems like we're a long ways off from good legislation there.
kingcauchy··on We're going to need default hard budget caps on pretty much everything
It'd be nice if more than AI spend worked this way, autoscaling is almost a mixed blessing because unpredictable pricing can be worse than the cost savings...
kingcauchy··on We're going to need default hard budget caps on pretty much everything
What that's awesome? Yeah they definitely waited until the competitors did it first...

Edit: sadface

kingcauchy··on Federal judge calls Flock 'indiscriminate mass surveillance'
I'm pretty sure we're living in the prequel to Minority Report right now.
kingcauchy··on Getting the most out of Opus 5.5 in Claude and Claude Code
Limits have been decently comparable to 6.1 Sol anecdotally. Although it's subagent eagerness can make then comparison ebb/flow.

The notable difference to me is tokens/sec are still much higher on 6.1 Sol

kingcauchy··on Getting the most out of Opus 5.5 in Claude and Claude Code
I've had troubles with it getting stuck "waiting" for a day on some hook or something in CC that never completed and during a task that was waiting on an orphaned process. That maybe saves token money on checks waiting for long-running processes but makes it hard to trust for long-horizon work.

It's been amazing at making sure OOMs for multiple heavy builds on my machine don't happen, adding queues and locks to make sure performance measurements are isolated and gpu stays clean during experiments.

It's also way more able to execute subagent tasks all at once than GPT 6.1 I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.

kingcauchy··on Kolibri: A Sovereign Open-Weight Model
Yeah the pdf alone is awesome as a learning tool.
kingcauchy··on Apple Pass Designer
This seems so oddly specific, seems like an integration with canva would be a better direction.
kingcauchy··on Zig v0.17.0
I understand, maybe one day I'll have the chance to hear the blog posts' worth of thought on the matter!

https://codeberg.org/ziglang/zig/issues/37060 I opened the issue and can give you the agent generated RCA on the matter if you all want it in issue 604 on antfly's github but I understand that's against policy and totally respect that.

Appreciate all the work you guys do and have been following the whole Zig project since inception fwiw!

kingcauchy··on Zig v0.17.0
Might be my favorite comment ever now.
kingcauchy··on Zig v0.17.0
Thanks for the clarification here!

So for instance because of the size of our codebase our project has pushed Zig to some of the edges, specifically we end up hitting a bug when using llvm and zig on arm64 (Mac and Linux) where it seems to be caused by some configuration Zig passes through to LLVM. I’ve used codex and claude to help me diagnose and find the bug (we use nix’s glibc zig to circumvent the problem now). I now understand the root cause but am not sure what the proper fix would be. But I’ve not known whether or not even raising the issue would break the terms of contributing? Would raising the issue break the implicit agreement?

kingcauchy··on How to keep enjoying programming in a world of LLMs
Isn’t this true of literally every technological innovation at some point in time?
kingcauchy··on How to keep enjoying programming in a world of LLMs
What’s your evidence or argument that it’s not? Intelligence atrophy is not a given and we have no evidence that ai is making people stupider (skill atrophy is not the same thing).
kingcauchy··on Zig v0.17.0
Interesting to hear! We maintain the project Antfly (entirely zig) and have been nervous about bringing issues to the zig folks or asking questions because of our ai usage.

We’re quite knowledgeable and thoughtful folks fwiw

I think I’ve seen you are a core team member or contributor? I remember your tag?

kingcauchy··on Zig v0.17.0
I’ve not found the zig folks to ban dissent, they engage in a lot of thoughtful dialog. Just because they’ve made a different decision for how they take contributions than other people agree with doesn’t make them a cult?
kingcauchy··on Zig v0.17.0
Human reproductions are notoriously error prone whether ai assisted or not right? I’m not sure what the analogy is between “llms are error prone” and “thoughtlessly copy-pasting something from Claude” is.
kingcauchy··on Livenerf: Has Opus 5.5 been nerfed yet?
We don’t know the model architecture but there’s a lot of evidence to suggest load dependency on the hardware affect models quality (see for example how the original Google Translate models got worse depending on time of day). The GPUs at maximal utilization is what they’re shooting for with their pricing models so you’d need to measure model performance at peak loads to know the floor of performance I would think.
kingcauchy··on How to keep enjoying programming in a world of LLMs
It's not hard to use a drill or a hammer but learning how to use them as tools to make things was always the real challenge. Using LLMs in a useful way I would argue is a much more important skill than programming as was learning how to make programming useful was always the more important skill, but it depends on if you view programming as a means to an end or the end in and of itself.
kingcauchy··on I don't want to read what you didn't write
I don't disagree with the main points of the article. But I feel like soon with all the writing that's been hating on AI writing recently on HN, the LLMs are going to be really good at writing articles about how bad AI is for writing...
kingcauchy··on Data Protection Commission fines Google €403M over processing of location data
It's like taxes with extra steps, except you get taxed for someone polluting you!
kingcauchy··on A search-and-inference database from scratch in pure Zig
First principles are your foundational beliefs, orientations, ideas upon which your values are construction, mission is built, direction is set, decisions are made for a business usually (I think the idea applies in general but personally I see the terminology more in technology/business). It's most important for being the deciding factor between two important values or difficult decisions say. For example, if your principles first and foremost say people have a right to their opinion, and your values say people should be respectful to each other and that people should honest, there can be conflict when someone is honestly disrespectful. But first principles would say (in a company at least) that a person has a right to their opinion and that's more important that they were honest about it than necessarily respectful. Contrived example but I hope that communicates the idea! For technology it more means the principles upon which product direct and tradeoffs are based
kingcauchy··on A search-and-inference database from scratch in pure Zig
DuckDB is a good one, we're working on that especially for the serverless/lakehouse stuff we've got planned for the next release! I believe we originally had qdrant in our benchmarks but ran into an explosion of testing requirements for each provider and different licensing checks for each but I can drum up those numbers!
kingcauchy··on A search-and-inference database from scratch in pure Zig
It can't be hosted as a cloud service correct (see ValKey by Google, OpenSearch by Amazon), there's a disclaimer on the GitHub about how and why as well.
kingcauchy··on A search-and-inference database from scratch in pure Zig
I think of perfect from two perspectives, one being "finding things I wanted to find", the other being "findings things I didn't know I wanted to find". I think Claude is great if the data isn't proprietary, secret (an all open-source project) but for dealing with Tax documents on my local machine I would hope that a search for my W2 would also find my 1099 I had forgotten I had, it'd be nice if I didn't have to allow the big AIs into everything to do that.
kingcauchy··on A search-and-inference database from scratch in pure Zig
We definitely were combining the rewrite with the opportunity to lay foundation for a more performant architecture, for instance index management and indexing autosharding could be resourced together in the new world with slightly different semantics in the apis. So in general if the traces disagree, we can count on the new version being correct (unless the spec was covered by a TLA spec)!

At the moment the reverse is true though, the simulator and what we've captured as ground truth for the desired design has been refined enough in tests and specs that the code is often the one implicated, and most of the bugs have been in code related to caching correctness and are only exposed through soak testing.

In opposition to Anthropic/Bun, we mostly used a hands-on approach to the rewrite and took the opportunity to capture the original design of Antfly into specs and any missing tests one subsystem at a time so we didn't strive to be as hands-off as "let Claude hill-climb on the tests". Especially since the system as a whole is far more dynamic and depends more on scalability, distributed systems stuff than Bun required!

kingcauchy··on A search-and-inference database from scratch in pure Zig
We rewrote Antfly, which I introduced to the world a little bit back https://news.ycombinator.com/item?id=47414291, from Go to Zig.

Thought it is interesting to juxtapose to the Bun rewrite from Anthropic and wanted to talk about why we went the other way! Would love to talk about our process or the technology!

Benchmarks against are linked in the article but here they are again for posterity https://antfly.io/releases/v0.2

kingcauchy··on AI is code – and can't be prompted into being smarter
I wonder if we'll see a new sort of "role" in the training (user, system, assistant) for unstrusted sources, I'm a little surprised we haven't already. In fact it would probably make sense to have an arbitrary number of entity roles and to be able to configure the chat calls with truth values. Interesting article though.

That being said AI is not code, it's a statistical algorithm with non-determinism baked in. You can write code to run them but it's nothing without the evolution of the model weights from the training process. And you can absolutely make the model weights better aligned with intent.

kingcauchy··on Anthropic apologizes for invisible Claude Fable guardrails
How much of the apology was written by Claude? How much of the release note process was written by Claude? Will they have better prompts going forward to make sure Claude doesn't write upsetting things into the release notes for devs like silent nerfing? Spooky times.
Page 1 of 2Next →