I moderate /r/kafka; people mistake it as a subreddit about kafka the product
twitter.com
twitter.com
I remember /r/crypto having predictable problems, given its a fairly small community going back long before cryptocurrency gained mainstream attention.
It's most interesting when people post asking about hacks (cheats) for Rust (the game) and get sent articles about hacking (programming) using Rust (the language)
Do they not see it coming? Or do they not care?
People can be excused for that type of name in the early era of the internet, but if you do it today... I just don't understand.
The foremost in my mind is the Go programming language. Several people refer to it as Golang because of it's poor name.
It's extremely non-obvious for a beginning programmer (or an intermediate programmer who hasn't used go before (like me)), and just leads to more confusion and frustration.
That and my URL purchase went from a potential $40k or more to $8.
I can't imagine a less searchable term than "I"
[0] https://www.reddit.com/r/Gunners/comments/2blrik/how_are_lon...
to be honest, i'm totally looking forward to arsenal's future -- young manager, absurdly young and talented team mixed with good veterans.
a front of Saka, ESR, Ødegaard and Martinelli looks so good!
The sub for the woody perennials ("real trees") is resigned to r/sfwtrees.
/r/johncena is about potato salads.
That's hardly a justification for either side, though. The Kafkas could have went with "/r/franz_kafka" and "/r/apache_kafka" and there wouldn't be any problem.
> I remember /r/crypto having predictable problems, given its a fairly small community going back long before cryptocurrency gained mainstream attention.
/r/peloton is about professional cycling, but many people think it's about Peloton trainers
I just learned Peloton is actually a word with a meaning for cyclists, not just a brand.
https://www.reddit.com//r/suberbowl
> This subreddit was banned due to being unmoderated.
Aww, it sounded funny. It's a shame reddit couldn't at least leave a read-only archive up. I tried archive.org but there's not much there.
This is not a new problem, I remember around 25 years ago when the usenet news group comp.windows.news, which was dedicated to Sun's "Network Extensible Window System", became overwhelmed by posts from people who thought it was for news about Microsoft Windows.
It’s rare a piece of tech has a more fitting name! “Is your orgs politics so complicated that direct team-to-team communication has broken down? Is your business process subject to unannounced violent change? Bogged down by consistent DB schemas and API versioning? Tired of retries on failed messages? Introducing Kafka by Apache: an easy to use, highly configurable, persistently stored ephemeral centralized federated event based messaging data analytics Single Source of Truth as a Service that can scale with any enterprise-sized dysfunction!”
Seriously, the one time I was in a situation where much of the team seemed hellbent on this "just put all in Kafka" idea (without really understanding why, exactly) the arguments they came up with were not too dissimilar from what you've shared with us above. It all seemed to come down to "OMG databases are hard, schemas are hard, our customers don't understand the data they're shoving at us. But Kafka will take care of all of that for us. Because, you know, shiny."
That said I'd still like to have a more ... balanced understanding of why Kafka may not necessarily be The Answer, and/or have more hidden complexity or other negative tradeoffs than we may have bargained for.
I worked for a high profile recently-failed project from a company that rhymes with Brillo, and our data was just beginning to be too big for google sheets (!). However, we were also having organizational problems because the higher ups were seeing the failing project losing money so they of course decided to hire 100 extra engineers. Our communications (both human and programmatic) were failing and the confluent salespeople began circling like buzzards. Of course by the time it was suggested we we use it the project was already 6 months past the point of no return.
My advice is that if your data fits in a database, use a database. Anyone who says that isn’t scalable should have to tell you the actual reason it doesn’t scale and the number of requests/users/GBs/uptime/ etc that is the bottleneck.
E.g., Confluent Replicator vs. Mirror Maker 2, Confluent Platform's tiered storage has been available for quite some time (right now a bunch of people from AirBnB are doing a stellar job bringing tiered storage to FOSS Kafka, I'm hoping 3.1 or 3.2).
Actually, easiest thing to do to see the differences is grep all the subpages of this link for properties that start with "confluent":
https://docs.confluent.io/platform/current/installation/conf...
Of course, when have a few megabytes of data and you route it through Kafka, then all you get is an opaque message queue where you can't see which message went from where to where. Good luck debugging any issues. But, hey, you got to use Kafka.
There's many ways to answer that using data streamed over Kafka - ingest it into your preferred query engine, go query it.
Kafka is a distributed log, that's it.
Kafka sits at roughly the same tier as HTTP, but lacks a lot of the convention we have around HTTP. There's a lot of convention around HTTP that allows people to build generic tooling for any apps that use HTTP. Think visibility, metrics, logging, etc, etc. Those are all things you effectively get for free with HTTP in most languages. Afaict, most of that doesn't exist for Kafka in a terribly helpful. You can absolutely build something that will do distributed tracing for Kafka messages, but I'm not aware of a plug-and-play version like there are for most languages.
The fact that Kafka messages are effectively stateless (in the UDP sense, not the application sense) also trips up a lot of people. If you want to publish a message, and you care what happens to that message downstream, things get complicated. I've seen people do RPC over event buses where they actually want a response back, and it became this complicated system of creating new topics so the host that sent the request would get the response back. Again, in HTTP land, you'd just slap a loadbalancer in front of the app and be done. HTTP is stateful, and lends itself to stateful connections.
Another issues it that when you tell people that they can adjust their schema more often, they tend to go nuts. Schemas start changing left and right, and suddenly you now need a product to orchestrate these schema changes and ensuring you're using the right parser for the right message. Schema validation starts to become a significant hurdle.
It's also architecturally complicated to replace HTTP. An HTTP app can be just a single daemon, or a few daemons with a load balancer or two in front. Kafka is, at minimum, your app, a Kafka daemon, and a Zookeeper daemon (nb I'm not entirely sure Zookeeper is still required). You also have to deal with eventual consistency, which can make coding and reasoning about bugs dramatically harder than it needs to be. What happens when Kafka double-delivers a message?
My pitch is always that you shouldn't use Kafka unless it becomes architecturally simpler than the alternatives. There are problems to which Kafka is a better solution than HTTP, but they don't start with unstable schemas or databases being difficult. Huge volumes of data is a good reason to me, not being sure what your downstreams might be is an option. There are probably more, I'm not an expert.
> our customers don't understand the data they're shoving at us. But Kafka will take care of all of that for us
Kafka isn't going to help with this at all. If your HTTP app can't parse it, neither will your Kafka app. Kafka does have the ability to do replays, but so does shoving the requests in S3 or a databases for processing later. I promise you that "SELECT * FROM requests WHERE status='failed'" is drastically simpler than any Kafka alternative. It is neat that Kafka lets you "roll back time" like that, but you have to very carefully consider the prospect of re-processing the messages that already succeeded. It's very easy to get a bug where you have double entries in databases or other APIs because you're reprocessing a request.
HTTP definitely has the edge when it comes to library support. In fact, Confluent et al offer HTTP endpoints for Kafka so that you don't have to deal with the vagaries of actually connecting to a broker yourself (the default timeout in python for an unresponsive broker is _criminal_ for consumers. You will spend several minutes wondering when the message will arrive.) We use an in-house one. But that introduces HTTP's problems back into the process; you need to worry about overwhelming your endpoint again...
Regarding application patterns, ideally you're writing applications that read data from one topic (or receive messages, parse a file, etc) and write to another topic. Treating it as a request that will somehow be responded to later in time scares me and I wouldn't do it. What if your application needs to be restarted while some things are in-flight?
I think the biggest drawback to HTTP in this space is that there's typically no coordination between clients and the server. Clients send requests when they want and the server has to respond immediately.
That becomes a big issue when you have an outage and all your clients are in retry loops, spiking your requests per second to 3x what they would normally be, on top of whatever the actual issue is.
Most of the retry stuff seems largely shared; i.e. your code should still have handlers for when Kafka isn't responding right. Kafka will only preserve messages on the queue, it won't help if you lose network connectivity, or your ACLs get messed up, or etc, etc.
> Regarding application patterns, ideally you're writing applications that read data from one topic (or receive messages, parse a file, etc) and write to another topic. Treating it as a request that will somehow be responded to later in time scares me and I wouldn't do it. What if your application needs to be restarted while some things are in-flight?
The pattern I've seen is to make the processing itself idempotent, and only ack messages once they've been successfully processed. So if you restart the app while it's processing, the message will sit there in Kafka as claimed until it hits the ack timeout, and then Kafka will give it to a new node.
As far as RPC, I'm not advocating that it's a good idea, but you could implement timeouts and retries on top of an event bus. Edge cases will abound, and I wouldn't want to be in charge of it, but you could shove that square block into the round hole if you push hard enough.
Redis is an order of magnitude easier to work with but struggles under loads that Kafka has no problem with. Also every once in a while our Amazon managed Redis queue will have a bad failover or melt down because someone runs a bad command on it, but our Amazon managed kafka has been rock solid since we switched to it. When we ran Kafka ourselves though we definitely watched it melt down a few times because we threw too much at one broker or we made obscure config mistakes. And figuring out why a consumer isn't getting messages is always a pain, whereas redis is always a dream to use.
Definitely agree. The basic concept of Kafka is that the publisher doesn't care, so long as data isn't lost. If you need the producer to redo stuff if the consumer failed, then Kafka is the square peg in your round hole.
And yeah, the best use case for Kafka is, IMO, "I have to shift terabytes or more of data daily without risking data loss, and I want to decouple consumers from producers".
But then there are blog posts saying kafka is a terrible job queue because you can only have one worker per partition and it's hard to get more partitions dynamically.
A very basic rule of thumb is, on an X broker cluster, have N partitions, where N / X = 0.
There's no harm in choosing something like 20 - 30 partitions for a topic, and increasing that when you need to scale consumers horizontally.
Dropping partitions is harder, but again, they're cheap, you won't need to for most use cases.
Only caveat to increasing partition count is when you're relying on absolute ordering per partition - key hashing can point to different partitions when you have 10 vs 50. It can still be done, but it requires a careful approach.
The protagonist gets the number of a bureaucrat and wants to contact him at his office in the castle.
The fact that the number he’s given causes all phones in all offices in the castle to ring, and the people who answer don’t know or care about what he’s after, just adds a layer of confusion and difficulty to his goal of getting to the castle.
> hninsight -q kafka
https://www.google.com/search?q=site:news.ycombinator.com%20...
it almost always turns up good stuff.
(Incidentally, at least in my opinion his three novels are his finest work, but they seem to attract much less popular attention than his short stories and novellas.)
The one thing I remember RMS saying after all this time is that “the whole point of writing software is so that you can give it a funny name”.
All things considered, I still think that’s good advice.
Why can't people follow the simple rule of not reusing any word or name that's already in use?
They don't seem to have a problem with racehorses.
In fact, why not get the most comprehensive list of racehorse names to date, and start using them for new software projects - then nobody will have to be creative until they run out.
That's it :(
If you haven't heard Elder, listen to this masterpiece:
Epic.
If all subreddits had to transition to this. Cool. Obviously that isn’t going to happen.
A forum/ a semi-anonymous board would provide the same utility as reddit does, without all the negative drawbacks(which are and people ignore them).The funny thing is that the solutions for these kinds of interactions multiplied compared to 5-10+ years ago where you had 2-3 forum options.I don't think it's a naming issue, it's a "swipe-down i'm lazy" issue from the low attention span tiktok generation.On one hand people grown accustomed to not try to hack the software at all and ask for help at every step of the way, and on the other hand, that help better be asked really easily.
I hated looking at threads on Microsoft or Intel's forums. Many replies to postings were clearly level 1 outsourced agents in India or some other faraway place reading from a script. It was plainly obvious because many times their response would not have much to do at all with what the OP was asking after, and all it did was just aggravate people because they felt they were not being heard. Lots of noise, no signal.
On reddit, one doesn't have this infestation of script-reading call-centre agents and can often find the information you need. Of course, reddit is not perfect either. Rules are arbitrarily interpreted by many mods, and there is no avenue for review or appeal to a higher authority.
That said, I do like classic web forums. They're great, everything is subject-specific. But one downside to them is that you have to travel to them to view them as opposed to being on an aggregator like reddit where people are already there and might have subscribed to a sub.
A lot of companies use Facebook for their customer interactions because they don't need to do much.. the business page template is pre-defined.
Have you tried any alternatives to a forum? I see similar sentiment to your post a lot on HN, and the impression I always get is that people haven't given alternatives an honest chance. There is a reason the forum is dying. The alternatives are simply better.
And i did not necessarily say full-fledged forums.Something like what spotify has, disqus, mastodon,etc. There are a lot of options, and i understand the "i don't want to make an account for that issue" argument, i truly do, but on the other hand i think if the issue is so important maybe another step doesn't hurt. Also naming clashing, spam/bots, eventually censorship issues are arguably way harder to appear.
Again, i think this is better for actual product/services where we're talking about official support channels, and not something that you rely on the community: free software: editors, games, etc.
No way I'm going to create a forum account, confirm the email, save the password, just for a on off.
If there are two things people think of as "kafka" (even if kafka is one of multiple names, identifiers, or euphemisms something is or has been known by!!), the descriptors should be allowed to all co-exist and users should discover the actual referents (which at that point might be identified by a UUID) via contextual search, direct links, or disambiguation pages. Projects that understand this at a deep level and reject the premise of simultaneously permanent unique human grokkable identifiers--projects such as Discord, which at least suffixes usernames with numbers to disambiguate conflicts--deserve our unending respect, as it is just so easy to not give a shit and build a system with usernames: no one was ever fired for being part of this problem, and people will defend to their dying breath how important these namespaces are without ever addressing the practical world of what happens when there are thousands of unrelated namespaces attempting to serve tens of billions of users (and no: this isn't a wild exaggeration, as products that exist for decades can serve more unique humans than were alive at any given moment).
A similar bot (using a naive bayesian classifier and a few examples) should be easy to build for the /r/kafka reddit.
When I google "elixir strings" I also only get results for Elixir guitar strings instead of the hex doc page I'm looking for :)
We added a note about it: https://lobste.rs/about#michaelbolton People still make the assumption but at least they now usually get that link as a rebuttal.
You might say it's tongue-in-cheek, but you never know on Twitter.
I'm not sure if there's a term for it but it's like linguistic debris. The namespace is being polluted; eventually it won't be obvious what any particular word refers to.
AWS is probably the worst in this regard...so many services with names that give no indication about what the service actually does.
Meaning the long and awkward "descriptive" name becomes a lie over time. But the "code" name has no meaning, so it doesn't become wrong, you just need to maybe re-learn the meaning of the otherwise meaningless word.
See also: How do you descriptively describe the replacement/rewrite that does the exact same thing, but exists in some form while the other one does, also?
How often is it that you have a service that drastically differs in purpose from it's original intent? Can you give an example of how that comes about?
And usually not too long after I've coined a fun name for a project, I wind up feeling like it's cringy, and start to loathe when other people talk about it.
Also choosing names with non-obvious pronunciation is also a peeve, sometimes it can become a marker between those who know how to pronounce it and those who don't. For example Godot.
That's not a naming problem as much as it's a search engine problem. Search engines ignore anything that's not remotely alphabetic so for the longest time Google saw "C++ C#" as "C C". More and more characters are taken into account when searching these days, but it's still very error prone today.
If anything, C# was a mistake because it was developed after the first search engines entered the web. Then again, this was made by the same company that developed frameworks called "COM", "COM+" and ".NET". The names C and C++ were chosen way before this could have ever been perceived as a problem, because only lunatics and tinfoil hats could have predicted the internet in its current form all the way back in 1983.
In my experience, adding a revision (C99, C11, C17) to a search query often gets more relevant results specifically for C actual. For C++ the search term cpp always works for me. It's a bit of a pain to teach yourself to use, but after a while you start to get used to it.
Yes, and nerfed the + search operator in the process.
And some of the names are reused in Android even though they are in standard Java libraries.
It's probably invisible to people who have been using it all along, but for me, it's absurd and infuriating because almost nothing has a name that places it in a larger context.
Google seems to have a particular issue with naming things.
If you can't name something in a meaningful way, the next best thing is to name it something completely unique, and not just a generic synonym.
Alsoalso why do we need "factories" for abstract objects that relate to other abstract objects that are some sort of "thing"?
I have to wonder if some of this has to do with really smart people coming to work at Google, adding something without ever really grokking the whole mess, and then leaving after a couple of years because they got their tour of duty for their resume.
One can't be expected to get names right on the first draft.
Actually, one of the best named products I remember (for devs) was Big Worlds. It was an engine for making MMORPGs. I think the name had a clever mystique.
It's freeing to "break the rules" and code with single letter variable names, first letter that comes to mind. ;)
And yet Project Zero exists :-)
Googling your product will always be quite nice
It has always disturbed me that magician’s names of power in books are far too short: I presume there is more encoding than just the letters - the intonation perhaps contains a lot of entropy. Otherwise a name like Ged could be brute forced (or would that make a good short fantasy story?)
Edit: Related is the short story “The nine billion names of god” - the story is archived a bit incorrectly, it starts in middle of the page with no heading, with the line “This is a slightly unusual request”: https://archive.org/stream/ninebillionnames00clar/ninebillio...
The biggest tech companies on earth get their heads together and come up with those beautiful names. I mean, alphabet & meta are a thing...
You'd expect that the people who chose these "names" have dogs named "dog" and their childs have names like boy" or "girl" ...
Something like a place that's more accessible to search engines, better moderator tools, etc.
Reddit isn't bad, but it doesn't have a level of excellence.
I really do think we need better community tools. It's something I've thought about for a while now.
Curation markets and DIDs are something that I think will be vital of scale is to be achieved.
Many people joined obviously thinking it was a gay right activist group.
The tribe that they named after that helicopter Mr.Kafka flew back in Vietnam and with that he killed Apache, the famous Viet Cong Sniper?
The world is a quite strange place...
And there is no escape.
As an NLP AI, I really wouldn't want do the job. Allthough, you could learn a lot.
Pythonesque I would say. After all, we're on Hackerne.ws.
This is the kind of Sailor's yarn, you sometimes catch a mermaid.