HNHacker News
TopNewBestAskShowJobs

caust1c

3,293 karma · joined February 13, 2012

Programmer. Hacker. Building RunReveal, the Security Data Platform

    https://runreveal.com
    https://twitter.com/Caust1c
    https://www.abraithwaite.net
submissionscomments
caust1c··on Ask HN: Who is hiring? (June 2025)
You got this, best of luck! Engineers with ClickHouse knowledge are def in short supply right now.

To job seekers, this is your sign. Learn about how ClickHouse works. If you know SQL, you're halfway there. The other half is learning ClickHouse's data storage model which is what gives it it's efficiency.

caust1c··on Poll: How often do you bathe?
HN has polls?
caust1c··on Is Winter Coming? (2024)
This anecdote corroborates my theory that it will still be critical to become an expert in your field. Everyone is treating AI like it's a zero-sum game with regards to jobs being "lost" to AI, but the reality is that the best results will come from experts in the field who have the vocabulary and knowledge to get the best answers.

My fear is that people treat AI like an oracle when they should be treating it just like any other human being.

caust1c··on Build your own Siri locally and on-device
So build your own crappy agent-assistant?

In earnest though, I'm certain we'll see a community replacement of Siri by end-of-year if the iPhone permissions model allows it or there's some workaround. IDK what the limitations are here but I'm eagerly awaiting the community to step in where Siri has failed.

caust1c··on Show HN: GlassFlow – OSS streaming dedup and joins from Kafka to ClickHouse
Thanks for clarifying, best of luck!
caust1c··on Show HN: GlassFlow – OSS streaming dedup and joins from Kafka to ClickHouse
How does the deduplication itself work? The blog didn't have many details.

I'm curious because it's no small feat to do scalable deduplication in any system. You have to worry about network latencies if your deduplication mechanism is not on localhost, the partitioning/sharding of data in the source streams, and handling failures writing to the destination successfully, all of which cripples throughput.

I helped maintain the Segmentio deduplication pipeline so I tend to be somewhat skeptical of dedupe systems that are light on details.

https://www.glassflow.dev/blog/Part-5-How-GlassFlow-will-sol...

https://segment.com/blog/exactly-once-delivery/

caust1c··on Fivetran to acquire Census
Check out redpanda connect / warpstream bento (depending on your license needs). Both came out of what was benthos.

https://github.com/redpanda-data/connect

https://github.com/warpstreamlabs/bento

caust1c··on The Agent2Agent Protocol (A2A)
Something else I just thought about:

Agent2Agent in an unsupervised environment could easily lead to the first Agentic worms. It's not hard to imagine a few agents talking to one another with the right prompt injection attacks that could end up spreading to other agents via A2A.

This is of course just speculation but I could definitely see this as being a big enabler of that possibility.

caust1c··on The Agent2Agent Protocol (A2A)
This Agent2Agent Protocol (a "compliment" to Anthropic's Model Context Protocol?) seems to me to just be an attempt at a land grab in the line-protocol AI communication ecosystem.

If I'm reading it correctly, A2A is similar to MCP in that they both use JSONRPC but extends the capabilities for agents to be able to communicate with one another, potentially using separate backend models. MCP simply exposes applications data and workflows to a model itself and is not attempting to make agents communicate with one another.

The fact that A2A wasn't proposed as an extension to MCP seems disingenuous at best. To me, it looks like Google (among the other AI giants) is trying to create their own repository of agents, controlling the protocol, thereby enabling them to become the de-facto source for finding trusted agents.

Further, it comes off to me as a defense against the existential threat that AI poses to google's search and ads monopoly.

The problem is, as a consumer of AI, I don't want multiple agents communicating with one another. What I want is one model that communicates with non-agentic services. Making AI work well and understanding what it's doing is hard enough. You now want to pull in multiple models and companies into the picture? Talk about a risk management nightmare.

Shadow IT SaaS is already a massive problem for companies. Now imagine Shadow Agents doing work for your business using A2A to connect dozens of different unsupervised work for the company. No thanks!

For the inevitable defense of A2A "But it's open source and Apache licensed!". That's just bait. If you control the protocol, you control the ecosystem. See: Android, VSCode, Chromium, Java, Kubernetes, etc.

For me? I like my single-model audibility pulling in context using MCP. A2A just seems like an insane attempt at a land grab in the AI agent wars.

caust1c··on Breaking up with vibe coding
I find reports on how others are using LLMs to code interesting to read.

I've personally found LLMs to have recently crossed the "uncanny valley" of programming for me, meaning that I'm much much more productive than without them.

I find that if you're really good at describing a problem and the constraints you want to solve, using a language it knows well (like Go) and following well known patterns in that language, you can describe thousands of lines of code and get accurate results.

Maybe this isn't vibe coding speicfically, but I actually review every line of code that the LLM puts out. It doesn't take long if you know what you're reading, and the LLM make a weird solution to the problem. Usually if I'm specific about how I want it solved, it does it well.

I also found it useful to say "Please tell me your plan to implement the solutions and ask me about any ambiguities that need clarification." In other words, don't make your own assumptions.

The results are incredible. Thousands of lines of code that maybe not stylistically like mine, but are structurally very accurate to what I'm looking for. Giving function/interface signatures and code examples works wonders.

caust1c··on Leaking the email of any YouTube user for $10k
It's everywhere and it's the worst. I sometimes ponder whether or not the volume of protobuf bytes represented as b64 encoded protobuf in JSON exceeds that of actual protobuf bytes sent over the wires of the internet, and then I pour one out for myself.
caust1c··on Teen on Musk's DOGE team graduated from 'The Com'
Discord is very popular with skiddies and real criminal organizations alike. It's got pretty basic KYC controls in place, meaning essentially anyone with just an email can sign up. It can be accessed from behind VPNs without any issues, so effectively it doesn't matter that it's not e2e encrypted.

I feel that discord the company probably let's it slide because:

1. Moderation at scale is incredibly difficult. 2. They work with law enforcement agencies to execute warrants and subpoenas.

caust1c··on Three Observations
The crux of the challenges AGI presents is hardly mentioned as a footnote in this blog:

> In particular, it does seem like the balance of power between capital and labor could easily get messed up, and this may require early intervention. We are open to strange-sounding ideas like giving some “compute budget” to enable everyone on Earth to use a lot of AI, but we can also see a lot of ways where just relentlessly driving the cost of intelligence as low as possible has the desired effect.

The primary challenge in my opinion is that access to AGI will dramatically accelerate wealth inequality. Driving costs lower will not magically enable the less educated to be better able to educate themselves using AGI, particularly if they're already at risk or on the edge of economic uncertainty.

I want to know how people like sama are thinking about the economics of access to AGI in more broad terms, more than just a footnote in a utopian fluff piece.

edit: I am an optimist when it comes to the applications around AI, but I have no doubt that we're in for a rough time as the world copes with the economic implications of it's applications. Globally, the highest paying jobs are knowledge workers and we're on the verge (relatively speaking) of making that work go the way that blue collar work did in post-war United States. There's a lot of hard problems ahead and it bothers me when people sweep them under the rug in the name of progress.

caust1c··on Exposed DeepSeek database leaking sensitive information, including chat history
Seems that you're right! Also, not that I doubted they were using OpenAI, but searching for `"finish_reason"` on the web all point to openai docs. Personally, I wouldn't say it's a very common attribute to see in logs generally.

https://platform.openai.com/docs/api-reference/introduction

Right there in the docs:

> Now that you've generated your first chat completion, let's break down the response object. We can see the finish_reason is stop which means the API returned the full chat completion generated by the model without running into any limits.

Regarding how training data ends up in logs, it's not that far fetched to create a trace span to see how long prompts + replies take, and as such it makes sense to record attributes like the finish_reason for observability purposes. However the message being incuded itself is just amateur, but common nonetheless.

caust1c··on Exposed DeepSeek database leaking sensitive information, including chat history
Interesting to note:

- Dev infra, observability database (open telemetry spans)

- Logs of course contain chat data, because that's what happens with logging inevitably

The startling rocket building prompt screenshot that was shared is meant to be shocking of course, but most probably was training data to prevent deepseek from completing such prompts, evidenced by the `"finish_reason":"stop"` included in the span attributes.

Still pretty bad obviously and could have easily led to further compromise but I'm guessing Wiz wanted to ride the current media wave with this post instead of seeing how far they could take it. Glad to see it was disclosed and patched quickly.

caust1c··on New speculative attacks on Apple CPUs
Blame society. Businesses won't value security unless the fear of getting attacked is sufficiently strong and the losses significant. Otherwise why invest in it at all?

Definitely not just hardware exploits though. Look at heartbleed for example. It's been going on a long time. Hardware exploits are just so much more widely applicable hence the interest to researchers.

caust1c··on Go 1.24's go tool is one of the best additions to the ecosystem in years
Thanks for the explanation and contribution! Very much appreciated :-)
caust1c··on Go 1.24's go tool is one of the best additions to the ecosystem in years
Having not looked at it deeply yet, why require building every time it's invoked? Is the idea to get it working then add build caching later? Seems like a pretty big drawback (bigger than the go.mod pollution, for me). Github runners are sllooooow so build times matter to me.
caust1c··on Show HN: Houseplant – Database Migrations for ClickHouse
In my experience, my pain hasn't been from managing the migration files themselves.

My pain comes from needing to modify the schema in such a way it requires a data migration, reprocessing data to some degree, and managing the migration of that data in a sane way.

You can't just run a simple `INSERT INTO table SELECT * FROM old_table`, or anything like that because if the data is large, it takes forever and a failure in the middle could be fairly difficult to recover from.

So what I do is I split the migration into time-based chunks, because nearly every table has a time component that is immutable after writing, but I really want a migration tool that can figure out what that column is, what those chunks are, and incrementally apply a data migration in batches so that if one batch fails I can go in there to investigate and know exactly how much progress on the data migration has been made.

caust1c··on A Tour of WebAuthn
Adam Langley is probably one of the most gifted teachers when it comes to explaining cryptography concepts. Very clear, concise, precise, and makes it simple enough for me to follow without getting my neurons all knotted up.
caust1c··on Show HN: K8s Cleaner – Roomba for Kubernetes
So is this the first instance of a Cloud C-Cleaner then? You could call it CCCleaner!
caust1c··on Egoless Engineering
This is great! The best teams I've worked on have worked towards the following:

Pizza teams that own the whole stack, and for the roles that don't need a full-time individual, specialists that come in and advise but also make it possible to DIY the things they do.

The best examples of specialists are Designers and Security teams as this talk highlights. They can make the tools and the means for other teams to self-service those needs. For example, security teams implementing CI tools and designers building design frameworks that are easy to apply. Conversely, they can feel free to make changes themselves and are empowered to at the best organizations.

Everyone else in product development is a generalist, including the managers, and everyone is on-call. When everyone is on-call then it results in far fewer alerts going off because when there is an issue, it's taken very seriously and remediated quickly in the following days & weeks.

I think GTM teams could also benefit from this same kind of process, but instead melding Marketing, Sales and Support roles and responsibilities.

My theory on why this wasn't more common in the past was that the work was too complex and specialized and that the tools and knowledge to do the job weren't as easy to acquire as it is today. LLMs have certainly leveled the playing field immensely in this area and I'm truly excited to see the future of work myself.

caust1c··on Ask HN: Why did no one save the Living Computers museum in Seattle?
The truly sad part is that there wasn't a portion of the inheritance earmarked for an endowment to keep the museum going. I've heard from unreliable sources that Paul might have assumed this would happen via the executors of his will but the executors decided otherwise. :(

An endowment of 200M at a pittance 6% interest rate would more than cover the costs of running the museum, but they probably wouldn't have even needed that much after working on figuring out how to make it generate a bit more revenue.

I'm sad that I never got to attend, but would have easily paid close to $100 to visit.

caust1c··on Stephen King to shut down his 3 radio stations in Maine
So happy they expanded to the Bay Area after living in Seattle for a few years!
caust1c··on Reweb: Visual website builder for Next.js and Tailwind
I hope this is the next standard for CMS-style sites. I'm eagerly waiting for the thing that's easy enough for non-coders to use that doesn't make the code-capable maintainers want to cry mercy.

I was pretty optimistic about netlify-cms the approach they took just missed the mark on some technical things that NextJS handles as part of the framework.

Best of luck! IMO there's still lots of opportunity in visual site builders, and this one looks like it has a lot of potential.

caust1c··on Go-Safeweb
[flagged]
caust1c··on How the Unchecked Power of Companies Is Destabilizing Governance
Yes, good point. That's pretty similar to scientific falsifiability at least!

I think for topics that are not as easily provable as reproducible builds though, trust gets murky.

There was something else I read (can't find it now) that made a similar analogy for a web of trust.

A web of trust (e.g. PGP) will have de-facto authorities since there will be a tendency for more people to sign individuals presumed to be trustworthy based on their history. It follows that the system runs into issues if their key gets compromised or if a false individual is subject to a sybil attack, producing the illusion of trust. See also: github stars, cryptocurrency, social media follower counts.

caust1c··on A new book shows how the power of companies is destabilizing governance
I've thought about this a lot too and I think about a few things:

- The barrier to creating and distributing content was higher (it still had to capture people's attention).

- We didn't have all the tools to artificially create content, just our imaginations.

I'm no doomer by any means, and I think it's useful to look back at history for clues as to how to manage it but it's hard to find clues when the situation is so different.

I still believe education and critical thinking are the best antidote for disinformation, but higher education in the US has continued to come under attack (and perhaps rightfully so with the costs rising extremely out of proportion to inflation).

caust1c··on A new book shows how the power of companies is destabilizing governance
I think trust and identity are fundamentally intertwined, for sure. I think it's also why trust can be easily gamed by posing as certain identity groups. It's a dark path taken to the extreme.

I just read supercommunicators by Charles Duhigg and one of the best take aways from it was to remind people that they have multiple identities that they hold dear, and some of them may be in conflict. It gets people to think more and not regress to knee-jerk beliefs they think they're supposed to hold based on their identity on a given topic.

caust1c··on A new book shows how the power of companies is destabilizing governance
Orthogonally related to the article, I think it gets to the deeper issue at hand with regulating technology:

The internet connects everyone and allows for free-flow of information, free-flow information is eroding people's trust.

We want free speech, but people use words to deceive and coerce. You can't make rules to stop this - people will always find ways around them.

Ken Thompson wrote "Reflections on Trusting Trust" in 1984 (fitting as it may be). The conclusion being that we can't rely on computers to build trust. But we need trust to live in a society.

It's human instinct trust one another. But falsehoods spread fast online, and after being fooled so many times, people are losing their natural trust in others.

What's the way forward? I'm curious what this crowd thinks.

← PreviousPage 2 of 11Next →