() -> System.out.println(1)
I lost hope in escape analysis quite frankly.Rnd is something I have written because Java's Random is slow and clunky.
https://github.com/questdb/questdb/blob/master/core/src/main...
1,042 karma · joined May 13, 2016
() -> System.out.println(1)
I lost hope in escape analysis quite frankly.Rnd is something I have written because Java's Random is slow and clunky.
https://github.com/questdb/questdb/blob/master/core/src/main...
@State(Scope.Thread)
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
public class EscBenchmark {
Rnd rnd = new Rnd();
public static void main(String[] args) throws RunnerException {
Options opt = new OptionsBuilder()
.include(EscBenchmark.class.getSimpleName())
.warmupIterations(5)
.measurementIterations(5)
.forks(1)
.addProfiler(GCProfiler.class)
.build();
new Runner(opt).run();
}
@Benchmark
public int testEscapeAnalysis() {
int[] tuple = {0, 2}; // esc analysis? where are you?
return tuple[rnd.nextPositiveInt() % 2];
}
}
And the output of GC profiler: Benchmark Mode Cnt Score Error Units
EscBenchmark.testEscapeAnalysis avgt 5 8.234 ± 0.029 ns/op
EscBenchmark.testEscapeAnalysis:·gc.alloc.rate avgt 5 2647.216 ± 9.275 MB/sec
EscBenchmark.testEscapeAnalysis:·gc.alloc.rate.norm avgt 5 24.000 ± 0.001 B/op
EscBenchmark.testEscapeAnalysis:·gc.churn.G1_Eden_Space avgt 5 2643.140 ± 177.137 MB/sec
EscBenchmark.testEscapeAnalysis:·gc.churn.G1_Eden_Space.norm avgt 5 23.963 ± 1.613 B/op
EscBenchmark.testEscapeAnalysis:·gc.count avgt 5 157.000 counts
EscBenchmark.testEscapeAnalysis:·gc.time avgt 5 103.000 msAt QuestDB we help developers handle explosive amounts of data while getting them started in just a few minutes with the simplest and most accessible time series database.
We are looking for an experienced Technical Content Writer to join our fast growing team. You will be writing technical documentation for developers along with explainers and tutorials. You should have demonstrable experience working closely with engineers and product managers to understand and document features along with programming experience and knowledge of database technologies. The role requires a good command of English, and great communication and self-motivation to succeed in a remote environment. You will report to the CTO and work closely with the engineering and product teams.
Find out more about the position @ https://questdb.io/careers/technical-content-writer/
Author here. QuestDB is a fast SQL open source database for time series. About a month ago we launched on HackerNews [1].
Today, I am excited to announce QuestDB 5.0.3 [2]. This new release includes our changes in memory mapping strategy, giving us better performance, as well as some major changes to the Postgres wire protocol support.
In this blog post, we explain our journey to improve QuestDB's performance, especially regarding memory management. Relying as much as possible on the kernel and avoiding extra layers turned out to be very successful.
Thanks
Vlad
Author here. QuestDB is a fast SQL open source database for time series. About a month ago we launched on HackerNews [1].
Today, I am excited to announce QuestDB 5.0.3 [2]. This new release includes our changes in memory mapping strategy, giving us better performance, as well as some major changes to the Postgres wire protocol support.
In this blog post, we explain our journey to improve QuestDB's performance, especially regarding memory management. Relying as much as possible on the kernel and avoiding extra layers turned out to be very successful.
Thanks
Vlad
I have written this piece of code: https://github.com/questdb/questdb/blob/master/core/src/main...
This sums 64bit values and using AVX2 it will sum 1Bn in 0.26s. Incrementing conditionally will not be as fast and will throw vectorization out of the window too.
SELECT cab_type, count(*)
FROM trips
GROUP BY cab_type;
From execution time it seems to me that this is a straight sum() of 32-bit integers. "cab_type" has two distinct values and if stored 32bit value for "green" is 0 and "yellow" is 1, straight sum of these integers will produce the desired outcome and explain performance. That said the same performance will not extend to key that has three or more distinct values.We also need to experiment with hugepages. The beauty is that if read and write are separated - there is no issue with writes. They can still use 4k pages!
Eventually both. We are starting with baby steps, e.g. get data from A to B quickly and reliably. Replication/HA will be first of course. Then we want to scale queries across multiple hosts. Since all nodes have the same data - they may as well all participate. Sharding will be last. We are thinking of taking a route of virtualizing tables. Each shard can be its own table and SQL optimiser can use them as partitions of single virtual table. We already take single table and partition it for execution. Sharding seems almost like a natural fit.
There is an option to store set of dimensions separately as asof/splice join separate tables.
- TCP-based replication for WAN - UDP-based replication for LAN and high traffic environments
We are currently building foundation elements of this replication, such as column-first and parallel writes. These will go into and always be part of QuestDB. TCP-replication will go on top of this foundation and also part of QuestDB. UDP-based replication will be a part of a different product we are building that will be named Pulsar.
- replication is in the works, this is going to be both TCP and UDP based, column-first, very fast.
- yes, benchmarks are indeed are done on second pass over the mmaped pages. First pass would trigger IO, which is OS-driven and dependant on disk speed. We've seen well over 1.5Gb/s on disks that support this speed. Columns are mapped into memory separately and they are lazy accessed. So the memory footprint depends on what data your SQLs actually lift. We go quite far to minimize false disk reads by working with rowids as much and possible. For example 'order by' will need memory for 8 x row_count bytes in most cases.
- durability is something we want user to have control over. Under the hood we have these commit modes:
https://github.com/questdb/questdb/blob/master/core/src/main...
NOSYNC = means OS flushes memory whenever. That said, we use sliding 16MB memory window when writing. Flushes will trigger by unmapping pages. ASYNC = we call msync(async) SYNC = we call msync(sync)
What makes QuestDB different from other tools is the performance we aim to offer. We are completely open on how we achieve this performance and we serve community first and foremost.