HNHacker News
TopNewBestAskShowJobs

dikei

1,116 karma · joined April 1, 2013

submissionscomments
dikei··on Open table formats are inevitable for analytical datasets
That comparison blog seems biased toward Hudi.
dikei··on We switched to Java 21 virtual threads and got a deadlock in TPC-C for Postgres
The first issue is your second code is not Java (no await/async literal for Java yet)

The second issue is they're not completely equivalent. In the second case, you'd need extra memory for the `reentrantLock`, while `synchronized` works with any object. Furthermore, if you need to use `wait/notify`, then there need to be an extra `Condition` object to use in combination with the `ReentrantLock`. For sure, developers can rewrite most `synchronized` to use `ReentrantLock` and `Condition`, but javac won't do it automatically for you.

dikei··on We switched to Java 21 virtual threads and got a deadlock in TPC-C for Postgres
The default is 256, way higher than 10.

But of course, when you have thousands of Virtual Threads all deliberately pinning the carrier thread, you quickly run out.

dikei··on We switched to Java 21 virtual threads and got a deadlock in TPC-C for Postgres
No, synchronized is a very primitive lock implementation compared to what's available in java.util.concurrent.Locks.

However, it's built directly into the JVM specification, so it's difficult to change while keeping compatibility, while j.u.c.Locks is just a library. In other words, they can't change synchronized schematic, so they created j.u.c.Locks as a replacement.

dikei··on We switched to Java 21 virtual threads and got a deadlock in TPC-C for Postgres
You can code in async/await style in Java too, using `CompletableFuture.supplyAsync` and `CompletableFuture.get`

Using Virtual Threads is a choice, it's not forced on to you.

dikei··on We switched to Java 21 virtual threads and got a deadlock in TPC-C for Postgres
I wonder if HikariCP, currently the best Java DB Connection Pooling library, suffer the same issue as c3p0.
dikei··on An overview of distributed Postgres architectures
Oracle RAC is a shared storage system, every database instances write to same datastore, so as long as a transaction is committed by one node, it'll be seen by other nodes.

Of course, the issue with all shared-storage systems is how much it costs to have a reliable and fast shared storage.

dikei··on Python 3.13 Gets a JIT
It's pretty clear to me.

>> JIT, or “Just in Time” is a compilation design that implies that compilation happens on demand when the code is run the first time.

>> What people tend to mean when they say a JIT compiler, is a compiler that emits machine code.

A JIT compiler is a compiler that emits machine code the first time that code is run, vs an AOT compiler which emits machine code when the code is built.

dikei··on An overview of distributed Postgres architectures
Redshift is a data warehouse, it's not suitable for OLTP use case.
dikei··on An overview of distributed Postgres architectures
>> The free alternative would be Mysql/Mariadb + Galera Cluster. Not as solid as proprietary ones, but far easier to use and less buggy than Postgres + tons of tools.

Until someone accidentally run an expensive DDL on your Galera Cluster: now your cluster is down for hours without anyway to cancel that query except nuking the entire database and restore from backup.

dikei··on An overview of distributed Postgres architectures
>> First time seeing someone call Spanner, CockroachDB, and YugabyteDB a "distributed key-value store with SQL"

That was the first thing come to my mind when I read the paper on Spanner and CockroachDB (haven't read the paper on YugabyteDB yet) though, and surely I'm not the only one.

dikei··on Where Have All the Websites Gone?
Isn't this how we classify web "generation"?

* Web 1.0: independent websites, self-maintained, hard

* Web 2.0: hosted platforms like blogs and social networks, easy, but rely on providers

* Web 3.0: promise to free users from the platform providers but are mostly crypto scam currently.

dikei··on Show HN: Quickwit – OSS Alternative to Elasticsearch, Splunk, Datadog
> To give you a concrete example, one company is ingesting hundreds of terabytes of logs daily and migrating from Elasticsearch to Quickwit. They divided their compute costs by 5x and storage costs by 2x while increasing retention from 3 to 30 days

I guess that's to be expected. Almost anything is more storage-efficient than Elasticsearch, FTS is so expensive.

dikei··on Apollo 11 vs. USB-C Chargers (2020)
Good point.
dikei··on Apollo 11 vs. USB-C Chargers (2020)
It's more complicated than that, in the link, they described it better:

>> The microcontrollers, running on PowerPC processors, received three commands from the three flight strings. They act as a judge to choose the correct course of actions. If all three strings are in agreement the microcontroller executes the command, but if 1 of the 3 is bad, it will go with the strings that have previously been correct.

This is a variation of Byzantine Tolerant Concensus, with a tie-braker to guarantee progress in case of absent voter.

dikei··on Mitchell reflects as he departs HashiCorp
> Cloudflare for cdn, GCP for load balancing, aws for compute

Isn't this cause you to pay twice the already expensive Egress fee ?

dikei··on datetime.utcnow() is now deprecated
So is Java with `LocalDateTime` and `ZonedDateTime`, a very useful example of making your typesystem work for you.
dikei··on Privacy is priceless, but Signal is expensive
> Is that true at scale? If I tell the telecoms that I want to send a billion messages per year it seems like they might be willing to take a lump sum instead of setting up the systems to bill based on usage.

In most of the world, SMS is billed per-message, so it's basically no extra effort on the Telecoms side at all. In fact, Telecoms' online charging systems are fast enough to calculate users' data usage by seconds in real time, so they don't even blink at counting SMS.

dikei··on Hacking Google Bard – From Prompt Injection to Data Exfiltration
Even human cannot reliably distinguish instructions from data 100% of the time. That's why there're communication protocol for critical situations like Air Traffic Control, or Military Radio, etc...

However, most of the time, we are fine with a bit of ambiguity. One of the amazing points of the current LLMs is how they can communicate almost like human, enforcing a rigid structure in command and data would be a step back in term of UX.

dikei··on Show HN: Jeeves – A Pythonic Alternative to GNU Make
Apparently, Fabric v1 is replaced by Invoke now.
dikei··on Show HN: Jeeves – A Pythonic Alternative to GNU Make
Remind me of

* Fabric: https://www.fabfile.org/

* Invoke: https://www.pyinvoke.org/faq.html

dikei··on Things I've learned about building CLI tools in Python
By default, pipx use the python it's installed with, but you can change the python interpreter for each installed application.
dikei··on Star Citizen's Squadron 42 is feature-complete, ten years after being announced
> now how do you handle updating all those in realtime is what amazed me, they even have bullets go through locations

Probably just as other distributed systems implement replication: you have a leader to broadcast instruction and other followers applied the instruction to their local copy of the state.

However, there'll always be delay/stale data between the leader and followers. So, to improve performance, you can allow the followers to perform speculative decision based on their own calculation, but these local decisions will be overwritten / rollback by instructions received from the leader.

dikei··on OpenTelemetry at Scale: Using Kafka to handle bursty traffic
Yeah, and to use S3 efficiently you also need to batch your messages into large blobs of at least 10s of MB, which further complicates the matter, especially if you don't want to lose those messages buffers.
dikei··on Ruffle: Flash Player Emulator
Back when being a student, I remembered following GNU Gnash effort to support for AS2 and AS3, they took years, but in the end, still could only make it work partially. Flash was still dominant in the browsers at the time, yet nobody managed to port to Gnash before it died.

I wonder how Ruffle get it working so fast.

dikei··on Counter Strike 2 is here
> I don't know if anyone is actually that good

All pro-gamers train everyday to build this spray pattern countering movement into muscle memory. Of course, some are better than others.

dikei··on Choose Postgres queue technology
IIRC, LISTEN/NOTIFY needs to pin a PostgreSQL connection to the client, so you won't be able to use transaction-level pooling with it.
dikei··on CERN swaps out databases to feed its petabyte-a-day habit
https://github.com/VictoriaMetrics/VictoriaMetrics#cardinali...

If I understand this correctly, it deals with high cardinality by dropping data. The operators need to monitor for this and adjust their data to lower the cardinality.

dikei··on NSO group iPhone zero-click, zero-day exploit captured in the wild
It seems that even turning off iMessage is not enough ?

This a zero-click exploit, which means you don't even have to open the message to get hacked.

dikei··on Designing a new concurrent data structure
> I think the lib could postpone the log replaying until the next publish. Chances are all readers will have the new epoch by than = the writer thread won't have to wait at all.

I think you can't even do a new publish if all readers haven't finished switching to a new epoch after a previous publish. Otherwise, you risk corrupting readers that are still on the original epoch.

← PreviousPage 3 of 17Next →