We benchmarked it, i was much, much faster.
Technically though once that multi-gig file becomes many hundreds of gigs, my computer would loose by a huge margin.
We benchmarked it, i was much, much faster.
Technically though once that multi-gig file becomes many hundreds of gigs, my computer would loose by a huge margin.
https://web.archive.org/web/20200414235857/https://adamdrake...
I have tried this, I literally wanted to create a simple web app that is powered by the cheapest solution possible, but it had to serve from a database that cannot be smaller than 150GB. SQLite failed. Even Postgres by itself was very hard! In the end I now launch redshift for a couple days, process all the data, then pipe it to Postgres running on a lightsail vps via dblink. Haven’t found a better solution.
The PostgreSQL database for a CMS project I work on weighs about 250GB (all assets are binary in the database), and we have no problem at all serving a boatload of requests (with the replicated database and the serving CMS running on each live server, with 8GB of RAM).
To me, it smells like you've lacked some indices or ran on a rpi?
Anyway, I’m curious what kind of data the op is trying to process.
But this means there are a hundred million entities, publishing 3x number of papers and a bunch of metadata associated. On redshift I can get all of this loaded in minutes and takes like 100G but Postgres loads are pathetic comparatively.
And I have no intention of spending more than 30 bucks a month! So hard problem for sure! Suggestions welcome!
By default you get a commit after each INSERT which slows things down by a lot.
For a 32 core processor, that means that it can process a data set of 100G in the order of 30 seconds. For some types of tasks, it can be slower, and if the processing is either light or something that lets you leverage specialized hardware (such as a GPU), it can be much faster. But if you start to take hours to process a dataset of this size (and you are not doing some kind of heavy math), you may want to look at your software stack before starting to scale out. Not only to save on hardware resources, but also because it may require less of your time to optimize a single node than to manage a cluster.
This is a great phrase that I'm going to use more.
Yes, you can. Without indexes to slow you down (you can create them afterwards), it isn't even much different than any other DB, if not faster.
>Even Postgres by itself was very hard!
Probably depends on your setup. I've worked with multi-TB sized Postgres single databases (heck, we had 100GB in a single table without partitions). Then again the machine had TB sized RAM.
*Obviously not really. But very very many things do, even doing useful jobs in production, as long as you have high enough specs.