HNHacker News
TopNewBestAskShowJobs

cangencer

1,343 karma · joined January 12, 2012

Working on distributed systems
submissionscomments
cangencer··on CrowdStrike's outage should not have happened
The article is extremely shallow beyond saying "Formal verification a la SPARK" should have been used, while not offering how this could actually work in the real world - I don't think the author has any experience working on any similar piece of software either.

While such techniques are available, would they be really applicable in a very dynamic environment such as with millions of PCs running various windows versions, needing continuous / real-time updates.

And yes, we of course know that QA and testing magically removes all possible failure modes/bugs.

cangencer··on Preliminary Post Incident Review
The thread is still wrong, since it was a OOB memory read, not a missing null pointer check as claimed. 0x9c is likely the value that just happened to be in the OOB read.
cangencer··on DHH: The luxury of working without metrics
Yet another contrarian nonsense from DHH - one medium successful product doesn’t give them so much credibility to dismiss what everyone else does.
cangencer··on Computing Performance 2022: What's on the Horizon
https://www.youtube.com/watch?v=zGSQdN2X_k0
cangencer··on Google IoT Core will be discontinued on Aug. 16, 2023
Like the good old Monty Python cheese shop :) https://youtu.be/Hz1JWzyvv8A
cangencer··on Don't start with microservices – monoliths are your friend
There are actually some patterns to deal with this, such as Saga - I'm actually working on a project (not open-source yet) related to this specific problem. You can reach me at can@hazelcast.com if you want to learn more.
cangencer··on Forget Twitter Threads; Write a Blog Post Instead
Threads are similar to slideshows - you can put a few bullet points that sound reasonably correct, but lacking in detail and context and you mostly forget about it after you read it, it's disposable.

A long-form narrative that is convincing and made to last and read repeatedly is much harder to write.

cangencer··on “Location-Based Pay” – Who are we to complain?
Your salary is only partially based on "the value created by employee" - most of it is the market forces of supply and demand. When you're remote, you're competing with a much larger number of people for the same positions.
cangencer··on TeaVM: Build Fast, Modern Web Apps in Java
I'll do a shameless plug of our modern Java benchmarking story here for the doubters: https://jet-start.sh/blog/2020/06/09/jdk-gc-benchmarks-part1
cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
The paper is part of the standardization effort but is not the final authority on the process. It is a very good reference on how to approach streaming SQL, even though the Jet model will have a few differences to the presented paper.

p.s.: sorry for the late reply (somehow I wrote and didn't publish).

cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
This is indeed a question that we get asked a lot. We have so far not though about adding more advanced scheduling capabilities for the cooperative threads. With the slot system, if you have 48 core available in the cluster, and running 8 jobs, each job will only use 6 cores each. With cooperative threading, each job runs on all the 48 cores. We have tested something like 5,000 concurrent jobs on same cluster, but essentially they may be competing for the same resources, so you'll need to do your capacity planning accordingly. Simple way to work around that would be to create separate Jet nodes (a Jet node is very lightweight) so you could have separate execution pools.
cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
Do you mean for a single job/pipeline? This wouldn't be possible at the moment. Our current focus has shifted from Beam a little bit - as we found out the beam threading model didn't play nicely at all with Jet's green threads (there is no way to distinguish between blocking and non-blocking calls).
cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
It's not very different than normal SQL. Imagine that instead of a finite result set, you instead have a never ending result set. You can also roughly map operations like windowed aggregation into SQL with some additional syntax. This paper gives a pretty good overview, even though we don't fully agree with the model presented here: https://arxiv.org/abs/1905.12133
cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
This is very true. Stream processing is both old and new and I think it takes time for technology like this to really mature. There's currently a standardisation effort around Streaming SQL which may bear some fruit, but probably still many years away. Right even if you want to use some standard language like SQL to describe streaming queries there's differences in each tool both in syntax and semantics.
cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
While Flink is a fully-featured stream processing framework I think there's some notable differences. Off the top of my mind:

- Flink uses Zookeeper for metadata and coordination, Jet doesn't require any external systems for resilience.

- Flink uses RocksDB and HDFS for checkpointing/snapshotting, Jet stores it in distributed, replicated in-memory store.

- Flink allocates operators to slots, while Jet uses green threads/cooperative multi-threading. This means you can run many concurrent streaming jobs on the same cluster, with very low overhead.

- Jet is basically a single, self-contained JAR. It's all you need to run a production-grade service (+ some connectors, if you'd like)

- Jet can scale up/down with very little friction. You start a couple of processes and they will form a cluster automatically. Kill a couple of the processes, and the cluster goes on.

That said, Flink have a great set of overall features, especially around persistence and huge states. This is another area we're currently investing in as well as SQL support.

cangencer··on Show HN: Jet – in-memory, fault-tolerant, distributed stream processing
The license is meant to prevent service-wrapping by cloud providers, other than that it doesn't have any implications for standard usage. The core library / server is Apache 2 and the rest of the connectors are community license. You can use and embed both the core module and the connectors for free.

The license itself is similar to the licenses from Confluent, Elastic among many others. You can read more about it here: https://hazelcast.org/blog/announcing-the-hazelcast-communit...

cangencer··on The battle to invent the automatic rice cooker
+1, in countries like Turkey, Persia, India, Pakistan rice cookers are pretty uncommon. In my home country of Turkey where rice is eaten very commonly, I've never heard of a single person use it. Further, rice is typically first washed several times, fried with a bit of oil and perhaps some spices before being simmered. I don't know if it's possible to do this with a Japanese style rice-cooker.
cangencer··on Performance of modern Java on data-heavy workloads
Follow-up post: https://jet-start.sh/blog/2020/06/23/jdk-gc-benchmarks-remat...
cangencer··on The Low-Latency Rematch: Performance of modern Java on data-heavy workloads
Original article and discussion: https://news.ycombinator.com/item?id=23465660
cangencer··on Avro Arrow – The record-breaking jet which still haunts a country
exactly!
cangencer··on Ask HN: Who is hiring? (May 2020)
Hazelcast | Fully Remote, European Time Zones | QA/Quality Lead | https://hazelcast.com/

I'm a Director of Engineering at Hazelcast. We build distributed systems at scale. We're looking for QA Leads/Engineers to help test our distributed storage and compute engines. If you want to do exciting work on distributed systems (think jepsen), find consistency and concurrency issues and do performance testing at scale, Hazelcast might be the right fit for you! Contact me directly at can@hazelcast.com

cangencer··on Design of a Distributed Stream Processing Engine
I'm one of the core devs and happy to answer any questions as well.
cangencer··on Amazon CodeGuru – Preview
I wanted to try it on a single repository, but it requested access to all repositories, public or private and also needs admin access for webhooks. No thanks.
cangencer··on How Did WeWork’s Adam Neumann Build a $47B Company?
I've only been to WeWork in London, yes I agree the vibe is a little different. You could argue they're both faking it being hip but WeWork does a slightly more convincing job while Spaces feels bit more corporate.

That said, I found the WeWork in London extremely noisy and tight compared to my current space. It might be just a side effect of real estate prices in respective cities, though.

cangencer··on How Did WeWork’s Adam Neumann Build a $47B Company?
See https://www.spacesworks.com/ which is basically Regus with a different brand name.

I am using their co-working in Copenhagen since a few months and quite happy so far.

cangencer··on Distributed Locks Are Dead; Long Live Distributed Locks
There was also a follow up blog post again by Basri regarding how it's tested using Jepsen for those interested: https://hazelcast.com/blog/testing-the-cp-subsystem-with-jep...
cangencer··on The tasting of surströmming
I also think the "disgusting" aspect is way overrated. It's basically fermented fish. The smell is very concentrated because it's been canned. A good video on how to eat it: https://www.youtube.com/watch?v=AGRyr8yIo9w
cangencer··on Do We Need Distributed Stream Processing?
We are adding something called a "rolling aggregation", where you receive a record, accumulate it and then emit the current accumulated value. I'm not sure if this matches what you want.
cangencer··on Do We Need Distributed Stream Processing?
I meant that products like Esper, StreamBase, InfoSphere have been around for a long time, which have a _very_ rich set of features [1], and are mostly designed around single process usage. Lot of the type of queries they support are not possible to implement in a performant way in a distributed system. Though nowadays Esper claim to have horizontal scalability - it was originally designed as a single threaded system. They do also have a passive/active type solution as you mentioned.

Stream processing frameworks originally evolved to offer "big scale" through data partitioning compared to the traditional CEP systems. But CEP engines have been able to deal with windowing and similar concepts since many years ago - the main difference of the stream processing frameworks _is_ the distribution and scalability aspect.

My point was that the systems linked in the original article seem to match closely to the limitations of what distributed stream processing frameworks are able to do, but only run on a single node.

[1] http://www.espertech.com/esper/

cangencer··on Do We Need Distributed Stream Processing?
Single process "streaming" or basically CEP engines have been around for a very long time and used to be the norm before distributed stream processing engines came around (such as Storm). CEP engines have much richer functionality than distributed stream processing engines because they don't need to deal with partitioning or data distribution. I'm not quite sure I see the appeal of making a non-distributed engine but with the same limitations of a distributed engine.

I work on Hazelcast Jet [1], which is a Java based distributed stream processing engine. The core engine is fast enough that it can be used with very good throughput on a single node (several times faster compared to Flink or Spark) but usually several nodes are not only needed strictly for parallelization but also for tolerating node failures and being able to restart where you left off. As others have pointed out, not every computation can be parallelised efficiently. Jet also offers in memory storage, so adding more nodes also increases your storage capacity.

Since the core of Jet is small enough (~400kb JAR), we also considered making a non-distributed version that runs strictly in process. Mainly for lightweight usage or embedding but would also offer a path to distributed execution, if it was ever needed.

[1] https://github.com/hazelcast/hazelcast-jet

Page 1 of 2Next →