I tried to check the kind of flights they flew in the world's dangerous airport (Lukla, Nepal)[0] and found they use ATR-72 series.
[0] https://adsb.exposed/?dataset=Planes&zoom=12&lat=27.7136&lng...
1,573 karma · joined November 4, 2013
I tried to check the kind of flights they flew in the world's dangerous airport (Lukla, Nepal)[0] and found they use ATR-72 series.
[0] https://adsb.exposed/?dataset=Planes&zoom=12&lat=27.7136&lng...
Can you elaborate more please? I would love if you can say what all can be improved to make debian package up to standards.
clickhouse local
ClickHouse local version 25.4.1.1143 (official build).
:)
There are few benefits of using clickhouse-local since ClickHouse can just do lot more than DuckDB. One such example is handling compressed files. ClickHouse can handle compressed files with formats ranging from zstd, lz4, snappy, gz, xz, bz2, zip, tar, 7zip. clickhouse local --query "SELECT count() FROM file('top-1m-2018-01-10.csv.zip :: *.csv')"
1000000
Also clickhouse-local is much more efficient in handling big csv files[0] clickhouse local --file medias.csv --query "SELECT edito, count() AS count from table group by all order by count FORMAT PrettyCompact"
┌─edito──────┬─count─┐
│ agence │ 1 │
│ agrégateur │ 10 │
│ plateforme │ 14 │
│ individu │ 30 │
│ media │ 423 │
└────────────┴───────┘
With clickhouse-local, I can do lot more as I can leverage full power of clickhouse.[0] https://clickhouse.com/docs/en/sql-reference/table-functions...
[1] https://clickhouse.com/docs/en/engines/table-engines/integra...
Overall I think Postgres adoption and integrations and thus community is much more wider than MySQL which gives it major advantage over MySQL. Also looking at the number of database-as-a-service companies of Postgres vs those of MySQL we can immediately acknowledges that Postgres is much widely adopted.
As the author mentioned, I completely agree with this statement. In fact, many companies like Cloudflare are built with exactly this approach and it has scaled them pretty well without the need of any third database.
> Another reason I suggest checking out ClickHouse is that it is a joy to operate - deployment, scaling, backups and so on are well documented - even down to setting the right CPU governor is covered.
Another point mentioned by author which is worth highlighting is the ease of deployment. Most distributed databases aren't so easy to run at scale, ClickHouse is much much easier and it has become even more easier with efficient storage-compute separation.
[0] https://clickhouse.com/blog/asynchronous-data-inserts-in-cli...
[0] https://clickhouse.com/docs/en/interfaces/formats#tabseparat...
https://clickhouse.com/docs/en/guides/developer/alternative-...
Apart from that, it provides various other feature:
- Dynamic datatype [0] which are very useful for semi-structured fields which generally logs contains very often.
- You can configure column's & table's TTL [1] which provides efficient way to configure retention.
At my previous job (Cloudflare), we migrated from Elasticsearch to ClickHouse and saved nearly 10x reduction in data size and got 5x perf improvement. You can read more about it [2] and watch the recording here [3]
Recently, ClickHouse engineers published a wondering detailed blog about their logging pipeline [4]
[0] https://clickhouse.com/docs/en/sql-reference/data-types/dyna...
[1] https://clickhouse.com/docs/en/engines/table-engines/mergetr...
[2] https://blog.cloudflare.com/log-analytics-using-clickhouse
[3] https://vimeo.com/730379928
[4] https://clickhouse.com/blog/building-a-logging-platform-with...
I work for ClickHouse and available for any queries or help you need.
[0]: https://clickhouse.com/ [1]: https://clickhouse.com/docs/en/about-us/adopters
- https://clickhouse.com/blog/apache-parquet-clickhouse-local-...
- https://clickhouse.com/blog/apache-parquet-clickhouse-local-...
Best part of ClickHouse native data format is I can use the same ClickHouse queries and can run in local or remote server/cluster and let ClickHouse to decide the available resources in the most performant way.
ClickHouse has a native and the fastest integration with Parquet so i can:
- Query local/s3 parquet data from command line using clickhouse-local.
- Query large amount of local/s3 data programmatically by offloading it to clickhouse server/cluster which can do processing in distributed fashion.
Having served as both ClickHouse and Postgres SRE, I don't agree with this statement.
- Minimal downtime major version upgrades in PostgreSQL is very challenging.
- glibc version upgrade breaks postgres indices. This basically prevents from upgrading linux OS.
And there are other things which makes postgres operationally difficult.
Any database with primary-replica architecture is operationally difficult IMO.
- https://clickhouse.com/blog/clickhouse-fully-supports-joins-...
- https://clickhouse.com/blog/clickhouse-fully-supports-joins-...
- https://clickhouse.com/blog/clickhouse-fully-supports-joins-...
- https://clickhouse.com/blog/clickhouse-fully-supports-joins-...
- https://clickhouse.com/blog/clickhouse-fully-supports-joins-...
- For querying csv data from command line, I use clickhouse-local.
- For querying csv data programmatically using a library, I use chdb (embedded version of clickhouse)
- For querying large amount of csv data programmatically, I offload it to clickhouse cluster which can do processing in distributed fashion.
If you are looking from query performance perspective, this blog is useful: https://www.vantage.sh/blog/clickhouse-local-vs-duckdb
- Whether the database vendor is a lock-in. It would be a straight "NO" for me if the database isn't open-source with proper license, since I can't self-host it in case, I need to move away from their SaaS offering due to various unexpected foreseen reasons: vendor decided to increase pricing, vendor has hidden pricing, vendor has reliability issues etc etc.
- How big is the community behind the database. Check their public forums to understand how community feels about the database and how their requirements are considered.
- Don't believe any random benchmarking post online but by doing my own benchmarking for my use-case.
- Check which other companies have adopted that database and read their experiences.
- It works with every data format which I come across in my daily job. It support data formats like Protobuf, Avro, Cap'n Proto which regular tools don't support. Funny thing, it can even read mysql dumps.
- I can read the data stored in local or remote location like http, s3, gcs, azure and whatever location I can think of.
- I can SQL queries on the raw data and improve my SQL skills on daily basis.