HNHacker News
TopNewBestAskShowJobs

mikaelronstrom

34 karma · joined April 9, 2018

submissionscomments
mikaelronstrom··on Scalable Revisioned Graph Database
Dydra using RonDB as backend scales to many billions of items in a graph database with revisioned history.
mikaelronstrom··on How to design a DBMS for Telco requirements
Not sure how it is today, but in the 90s and early 00s when I worked at Ericsson Mnesia had a number of systems that it was used in. There was interestingly 4 DBMSs developed at Ericsson in the 90s and 00s. Mnesia and NDB Cluster as mentioned and two internal DBMSs, TelORB and DBS, I had interesting discussions with all of them and worked with DBS before developing NDB Cluster. We were all situated within 5 minutes walking distance :)
mikaelronstrom··on How to reach 100M Key lookups using REST server with Python clients
This provides the How-To explanation of challenges one can meet when building highly scalable Key-value stores using a REST server and Python clients.
mikaelronstrom··on Migrating from AWS to a European Cloud – How We Cut Costs by 62%
The reliability of OVH has so far not been a problem. Performance is mostly a factor of the VM instances that you use, AWS has much more instances to choose from, but obviously at a much higher cost, so performance per dollar is obviously better with OVH.
mikaelronstrom··on Migrating from AWS to a European Cloud – How We Cut Costs by 62%
There are lots of alternatives for database solutions with Kubernetes, in Hopsworks we use RonDB which is highly available and has a MySQL interface but also supports a REST API and even a prototype Redis service.
mikaelronstrom··on Migrating from AWS to a European Cloud – How We Cut Costs by 62%
One obvious benefit of OVH compared to Hetzner is the S3 storage, the Kubernetes framework and a number of other services provided by OVH, don't think Hetzner would provide those services.
mikaelronstrom··on Migrating from AWS to a European Cloud – How We Cut Costs by 62%
Hopsworks has a platform with a high-availability database (RonDB), a highly available distributed file system (HopsFS) and the services are also redundant. Hopsworks even supports failover to another region. So the software is very reliable. But given that, OVH is a proper cloud vendor and not only a budget hoster, so it has fault tolerant services as well. We use other European companies for budget hosting of our development systems.
mikaelronstrom··on [dead]
Intriguingly simple implementation of fibers in Linux and Mac OS X can be used to experiment on the use of fibers in the RonDB database engine.
mikaelronstrom··on [dead]
New VM types using Intel/AMD VMs of the 4th generation in GCP is compared to 2nd generation VMs. Benchmark is Sysbench and database is RonDB (www.rondb.com).
mikaelronstrom··on [dead]
Some playing with numbers for vacations.
mikaelronstrom··on The end of a myth: Distributed transactions can scale
RonDB is a KV store with SQL capabilities (the MySQL storage engine NDB is maintained by Oracle as part of MySQL NDB Cluster). Hopsworks is on top of this adding a new REST API (REST + gRPC) service that is already in the github tree and will ready for production usage in a few months.
mikaelronstrom··on The return of Desktops for development environment
Most developers today use laptops for all development work. For the last 15 years I have considered desktops and laptops to be very similar in performance and use cases. This is no longer the case as I will discuss in this blog.
mikaelronstrom··on [dead]
LATS stands for low Latency, high Availability, high Throughput and scalable Storage. When testing an OLTP DBMS it is important to look at all those aspects. This means that the benchmark should test how the DBMS works in scenarios where data fits in memory, where data doesn't fit in memory. In addition tests should run measuring both throughput and latency. Finally it isn't enough to run the benchmarks while the DBMS operates in normal operation. There should also be tests that verify the performance when node fails and when nodes rejoin the cluster.
mikaelronstrom··on Automatic memory management in RonDB
After 5 years of development all steps in the new memory management architecture in RonDB is completed. This has been introduced in 4 steps and now the final step is completed. Thus it is a good time to describe the architecture and how it works.

The original architecture was using fixed size memory structures that was very fast, but not so flexible. We have worked hard during this transformation to retain the performance while adding a much more flexible memory architecture.

This type of automatic memory management is normally only found in expensive closed source DBMSs. Now RonDB brings this type of automatic memory management to open source DBMSs as well.

mikaelronstrom··on Ask HN: What does database (internals) development look like?
General optimisations are not so common to work on, but new HW requires new thinking about database architecture.

A few examples, Scaling to more CPU cores, Scaling to bigger RAM, bigger disks, scaling to larger systems. Introduction of new HW such as Intel Persistent Memory, new GPUs that can be used for various purposes in a database such as compression and encryption.

Every product has weaknesses that needs to be addressed, what these are is obviously product dependent.

Personally I spent a very significant amount of time the last 5 years to automate algorithms such that they automatically adapt to load, memory sizes, VM size and so forth.

Masters are definitely ok, but my Ph.D studies have certainly helped since that made me to do a deep dive into all database algorithms. So masters is sufficient to be a database developer, but I would say a Ph.D is a good idea if you aim for a database development architect role down the road. Best of luck in your new tasks.

mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
Regarding memory a very quick formula (more details exists in docs.rondb.com and in blogs) is around 25 bytes of overhead per row plus 15 bytes of overhead per primary key index and an additional 10 bytes per row per ordered index. Non-indexed columns can be disk-based and thus use SSDs or NVMe drives or networked storage. This is decided when creating the table.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
RonDB Community is GPL v2 and there is no limitations to its use. Our business model is to provide the managed service of operating RonDB and providing support for that.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
RonDB have an interpreter that can execute a set of simple things. It is mostly used to push filtering, to push increments/decrements. There is also a pushdown join processor in it. It wouldn't be very hard to build more functionality into the interpreter. The interpreted programs is created by the NDB API and executed by the data owner. Thus the intermediate parts like transaction handler has no idea what it is passing along.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
RonDB has the capabilities to both scale on VM size and number of VMs as online operations. The auto part requires that we support in the managed version as well. Hopsworks have added auto-scaling to AI worker nodes. So auto-scaling in RonDB seems like a natural progression to this. Thx for the suggestion.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
Yep, this is a key difference between traditional SQL databases and key-value databases. SQL databases optimise specific queries and have a high overhead per query. Key-value databases have a low overhead per query and optimise on flows of queries instead of on a single query. RonDB is a key-value database with SQL capabilities, so has a bit of both.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
For the moment managed RonDB will provide the availability of the cloud, but remember that the managed version is still in development. The steps to go 6 9s requires 1) Integrate cloud APIs such that we know when the cloud provider will freeze the instances 2) Provide global replication between cloud regions and failover handling of this. As mentioned reaching 6 9s requires both RonDB SW that is capable of reaching 6 9s as well as operational competence to actually deliver it. This is what we're aiming at, to make this availability reachable for normal users without this operational competence to deliver 6 9s.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
Correct, Availability is the amount of time you are available to read and write data. Durability is the amount of time your data is not lost. The metrics mentioned in this post are about availability. Most cloud vendors provide SLAs of 99.95% availability. One problem to solve when working with a cloud vendor is that they need to upgrade their OS images every now and then, so to get the highest availability one must integrate with the cloud APIs announcing those changes.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
We are integrating benchmarks into RonDB to make it easy to compare to other products, currently Sysbench (standard open source database benchmark) and DBT2 (TPC-C open source variant) available, will later add more internal and standard benchmarks.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
YCSB is one benchmark that is commonly used by key-value stores. See http://mikaelronstrom.blogspot.com/2020/10/ycsb-disk-data-be..., http://mikaelronstrom.blogspot.com/2020/02/ndb-cluster-world...
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
RonDB supports read-what-you-write consistency which is actually more than any eventual consistency database provides. Thus when you have written something into RonDB you can trust that it is seen by you and others. RonDB provides row locking, this means that the application can provide a stricter guarantee if desirable. Concurrency control and consistency in a database is too complex to handle here, for an in-depth coverage of RonDB consistency, see https://docs.rondb.com/rondb_concepts/
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
It is based on MySQL NDB Cluster which is GPL v2 licensed. Thus so is all changes made to RonDB.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
Assuming sacrificing: The availability numbers requires instant failover, this requires updating all replicas synchronously. Throughput and latency are both coming from using an asynchronous programming model which have been refined over the years. Todays new blog on this topic is here: http://mikaelronstrom.blogspot.com/2021/05/research-on-threa...
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
Yep, 30 seconds is easier to remember as is 5 minutes for 5 9s :)
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
In order to achieve 6 9s you need a two levels of replication, you need synchronous replication with instant failover, plus asynchronous replication to handle site failover. RonDB provides both. NDB is used in lots of telecom services where you can't make a phone call unless NDB is up. These services definitely can at times run for decades without downtime. RonDB is built on top of NDB.
mikaelronstrom··on Show HN: RonDB – fast key-value database in the cloud
These numbers are based on NDB customer experiences from operating tens of thousands of NDB clusters for more than 10 years. Obviously to achieve 99.9999% uptime requires an operational competence as well as the software to achieve it. This is why we are building this operational competence to make those numbers accessible to anyone.
Page 1 of 2Next →