HNHacker News
TopNewBestAskShowJobs

ilovesoup

15 karma · joined August 4, 2017

submissionscomments
ilovesoup··on We Chose a Distributed SQL Database to Complement MySQL
Well, its a lot easier to use TiFlash than replicating your OLTP data to a read-optimized database (whatever it is).

1. Add TiFlash nodes to your TiDB cluster (you might need to config some ports and ip address just like other databases and then start the nodes)

2. ALTER TABLE your_table set TiFlash replica 1;

That's it. You won't need all data in memory and replicating details are taken care of silently (schema change, fault-tolerance, load-balance, consistency and etc).

I believe heterogeneous replicating even with minutes latency tolerance will not be as easy as above :)

PS: the good part of keep TP/AP consistent is that user can use them in a single application without any surprise: it's just a different forms of a single logical data.

PS2: I'm a dev of PingCAP.

ilovesoup··on TiDB 4.0 GA Release
The reason for ClickHouse is simple: it's fast. And we need to have it function like a TiKV coprocessor which support filtering and aggregation mainly and ClickHouse is good at aggregation and filtering. Also it might take more time and dirtier to do seamless / compatible integration to TiDB if Impala or presto on top of MPP layer. But the price we paid is implementing MPP layer by ourselves now.

Almost all the data lake based products loss full control over storage system. It makes them very hard to build delta-main engine we need. To make HTAP storage transparent to query layer, TiFlash need a lot more control over storage engine than data lake can provide.

ilovesoup··on TiDB 4.0 GA Release
I'm the product owner of TiFlash. Yes. We used ClickHouse as the compute engine for TiFlash. The project started as a modification of ClickHouse (more or less it still is). It was like "pushdown query to an actual OLAP database" style 2 years ago. But later on we made a lot tighter integration (raft instead of binlog sync from TP or data-replication for TiFlash itself, implemented same type system, txn, online ddl, Coprocessor interface as TiKV and etc) to make it more "transparent" for query layer.

We will have more detail explained recently.

It will be open sourced in a year or two. For us we need to make the code open-source ready instead of just turn on github settings.

ilovesoup··on Modern Data Lakes Overview
(I'm a dev of TiDB so I might be biased.) Yes and no. The yes part is that TiDB still rely on TiSpark for large join query as well as bridging big-data world. TiDB itself cannot shuffle data like MPP database yet. On the other hand, TiDB without TiSpark is still comfortable of those dimensional aggregation queries (which are typical analytical queries as well). The no part is, TiDB now has a columnar engine (TiFlash) for analytics and providing workload isolation. TiFlash can keep up to date (latest and consistent data to be more specific) with row store in real-time in separated nodes via raft. IMO, HTAP should be TP and AP at the same time instead of just "TP or AP you choose one". In such cases, workload interference is real deal. Especially when you are talking about transactions for banking instead of streaming in logs. In such sense, very few, if any, "newsql" systems achieved what I considered true HTAP. For more details: https://pingcap.com/blog/delivering-real-time-analytics-and-...

Welcome to try it in March with TiDB 3.1.

ilovesoup··on TiSpark sits Spark SQL on top of a storage engine to answer complex OLAP queries
I'm the main dev of TiSpark. I totally agree. For now we allow trx only in TiDB. I believe one day those big-data stuff will be unified onto one platform. With a full-featured distributed db storage layer underneath, there might be tons of tricks to play comparing to data on hdfs. Ultimately, we plan to put a mysql layer on top of Spark SQL (maybe or something else as mpp engine), as you said, to make user not aware of existence of Spark SQL underneath.