How to move from Amazon RDS to a dedicated PostgreSQL server
layer0.authentise.com
layer0.authentise.com
PostgreSQL does (if unicode character set was specified when PostgreSQL database cluster is initialized first time); it's Ubuntu OS which creates PostgreSQL database cluster with ASCII character set (encoding) after PostgreSQL's installation.
> Here is the snippet that I had to use: ...
Instead of resorting to those hacks, follow the PostgreSQL documentation[1] to do it the right way. The simplest way is to initialize your PostgreSQL database cluster as:
initdb --encoding=UTF8 --pgdata=<database-cluster-dir>
If you want to use an existing database cluster, PostgreSQL documentation shows how you can do it.
--
[1] - https://www.postgresql.org/docs/9.5/static/multibyte.html
> Using the distro-supplied packages for PG
> (or any database server)
Or almost any of your application environment, really.newly setup:
List of databases
Name | Owner | Encoding | Collate | Ctype | Access privileges
-----------+----------+----------+-------------+-------------+-----------------------
postgres | postgres | UTF8 | en_US.UTF-8 | en_US.UTF-8 |
template0 | postgres | UTF8 | en_US.UTF-8 | en_US.UTF-8 | =c/postgres +
| | | | | postgres=CTc/postgres
template1 | postgres | UTF8 | en_US.UTF-8 | en_US.UTF-8 | =c/postgres +
| | | | | postgres=CTc/postgres
(3 rows)So LANG=de_DE.UTF-8 will load a de_DE.UTF-8 Database, while my LANG previously was en_US.UTF-8, i can change LANG to something different and it will create the cluster to the value I've set my LANG to.
This is actually poor advice. The right way on Ubuntu is to use the packages provided by PostgreSQL Development Group (https://wiki.postgresql.org/wiki/Apt), and in that setup you use pg_createcluster instead of initdb.
The snippet in the original article was so hacky because it fixed an existing cluster instead of destroying it and creating a new one in UTF-8.
initdb is part of standard PostgreSQL utilities/tools[1] for initializing your PostgreSQL server that's why I suggested using it. On the other hand, I don't see any reference to pg_createcluster command in the official PostgreSQL documentation. So I don't know why you think using a non-standard PostgreSQL tool (in place of a standard one i.e. initdb) is the right way of initializing a PostgreSQL database cluster.
[1] https://www.postgresql.org/docs/9.5/static/reference-server....
PGCLUSTER=9.4/main pg_dump foo
PGCLUSTER=9.5/main pg_dump fooI've been trying to compare RDS for Postgres to other offerings like Compose or Heroku but have come up surprisingly dry on comparisons.
You're looking at $2700/mo for the IO provisioning and $750/mo for the 6TB of storage. Double those if you want Multi-AZ. Then you get to price the server size which I imagine is one of the more expensive options if you need 30k/6TB.
Also FYI, your homepage has a 2015 copyright notice on it. No big deal, but thought you might like to know since it could put off particularly "nit picky" types of customer :)
If that's when the work was created, then that's correct. Copyright notices aren't there to tell you what year it is today. They are there to tell you when the work was created. If they change it to 2016 when the work was really created in 2015, then that's an invalid copyright notice and equivalent to no notice at all.
What if the page is dynamically created on the fly? What copyright should I have? If page contains snippets/work created in different years?
What traits are you comparing on? What kind of a thing are you building? What features (cost, performance, ease of use, time) are the most important to you?
https://blog.codeship.com/heroku-postgresql-versus-amazon-rd...
https://www.runabove.com/PaaSDBPGSQL.xml
No affiliation, just a happy customer of OVH's hardware.
My company has a bunch of microservices that each need a Postgres database, but none of them are particularly high traffic. We do however need high availability guarantees for all of them, and so Heroku's pricing starts at $200/mo per database.
With RDS, we get adequate performance for a similar total cost (around $300/mo), but get to run about a dozen separate high-availability databases on that same instance, each with their own usernames and passwords. That means significant savings.
(of course you could just share a single Heroku Postgres database and user account between these services, perhaps separating them by scheme, but that wasn't really to my taste.)
Aiven services are available on UpCloud and we're friends with the UpCloud people as we're both based in Helsinki and have met at various events, but there's no other connection between the companies.
One tool they did miss for continuous archiving is WAL-E, which tends to be the one most used including by us at Citus Cloud and Heroku Postgres - https://github.com/wal-e/wal-e
*disclaimer: I wrote that code while at heroku
* "AWS DMS doesn't support change processing on Amazon RDS for PostgreSQL. You can only do a full load."
* "AWS DMS doesn't map some PostgreSQL data types, including the JSON data type. The JSON is converted to CLOB."
http://docs.aws.amazon.com/dms/latest/userguide/CHAP_Source....
Hopefully it improves soon, since it's an otherwise awesome tool.
I wrote up my experiences and the reasons for the switch here: https://www.theguardian.com/info/developer-blog/2016/feb/04/... .
Do people do this for security / policy purposes, or are they motivated by cost savings? Does removing all that virtualization buy improved performance?
DMS isn't just for moving two RDS it can also be used to move off or RDS or even move from one non-RDS server to another non-RDS server that isn't even on AWS.
It can also do Postres to MySQL, MySQL to Postgres, etc.
It also does continuous replication so once the copy is done it will keep up to date (yes, even if it is Postgres -> mySQL).
One gotcha, though, it does not bring over secondary keys so make sure to recreate them before sending traffic over.
But this article does provide some insight on how to configure Postgres to be similar to RDS in terms of functionality.
Even if you did I probably wouldn't use it. I don't want a 3rd party dependency on the most critical part of my infrastructure.
edit: and if it's an hosted service, I'm assuming decent networking options (ipsec et al)?
The main thing we are trying to figure out - if there is a need for service that has significantly higher tears then what Azure, AWS offers.
Just a little input: Amazon and Microsoft are well known companies. I (and my superiors) think they can support us. Your company that I haven't heard of is not.
It may be dumb sounding, but I won't get in trouble if I used AWS and something went wrong. How could I have known?
With an unknown company, if something goes wrong, it's on me because I selected that vendor.
Stupid perhaps, but important to understand.
Think about it this way: OCZ has on paper SSDs that are simultaneously both faster and cheaper than, say, Intel. But I wouldn't put a OCZ drive into the lowest budged ricer gaming PC I could imagine (let's just say that have a bit of a reputation, to put it lightly), and I should be rightfully fired if I suggested putting one in a server.
Yes, your service might be massively faster, but... at 10-15% cheaper? No way! Not until you've been running for years with a great reputation. Now, if you enter the managed database space and pull a Digital Ocean ($5-$40 pricing vs comparable EC2 at $80-$400 AND better raw performance on top), then you can essentially create a new market niche. That's my suggestion to focus on, because businesses who operate 50TB databases aren't looking to save a few dollars. But people who use DO, need a database to go with it that is both faster and cheaper than AWS.
Honestly, performance ranks pretty low on the things I've found people look for when choosing a managed database service. You can serve a million users per day and average a whopping 12 IOPS, which might just about push a Raspberry Pi with sqlite slightly. Got more than a million users? Congrats, you now officially qualify to throw money at problems, the kind of problems RDS and Amazon loves to solve by injecting money.
Anyway enterprise SSD's have so much overall better performance that consumer grade SSD's is something I consider a joke. A joke even for consumer applications, web-surfing and email-clients. Remember that a lot of applications out there use SQLite, and poor DB performance means poor application performance.
I highly recommend reading this: http://www.sebastien-han.fr/blog/2014/10/10/ceph-how-to-test...
Reliability of a database is always no.1, without it, nothing else matters. So the customer must have enough access to the machine to setup replication and failover themselves (eg. using something like repmgr). Or have damn good guarantees that you're doing it well for them. ie. multiple replicas, verifiable backups, offsite backups, etc.
Secondly is size. This is a major sticking point for me when looking at hosted postgres solutions - they're mostly for relatively small databases <1Tb or so.
Closely in third is price. Even when you do find a service that provides a decent size like when RDS eventually went up to 6Tb (still not quite enough for me), the price they charge for it (and the ram) over buying dedicated iron is exorbitant. Given that 6-8Tb drives are mainstream now, it's a real head scratcher why they'll charge $750-1500/mth for a 6Tb drive that costs $250 outright. And then charge you for the IOPS on top of that... Even factoring in multiple redundant drives/replicas, SSD vs HDD, it's still very unfavourable. I can install 5 machines (at least) in a cluster for the price of 1 on RDS; giving me not just price savings, but my no.1 need - reliability.
Finally performance. The number of high IOPS workloads out there is rather small. You're far more likely to have a high storage need with small IOPS, than high IOPS, smaller storage. Everyone overestimates the amount of traffic they'll get, and generally if you have an IOPS problem, it's better off fixed elsewhere in the app. Usually some bad code or ORM is thrashing the DB.
Take a simple login DB as a thought experiment, just 10,000 IOPS is upto 864 million users logging in per day. Whereas the storage for that number of users, at a very generous 1Kb per user, is 864Gb. Size is way more important than IOPS. And if you've got that number of users, you've also got your own datacentre :)
You might also run into the problem that the kind of organisations that need high storage + high iops are also the kind that wouldn't use the cloud for it in a million years (banks, high freq. traders, etc).
In summary, I think the main thing that would attract me is high capacity storage for a decent price. Reliability is a given, but it must be easy to manage. Ideally I'd like a mixed storage system, with the ability to arrange the tablespaces across the drives as I need. 50Tb is great, but I need the ability to put the high IOPS tables on an SSD, then spread the infrequently accessed tables onto large HDDs.
> [pg_dump ...] The main disadvantage of this method is that it will not provide high reliability.
Pretty sure the author means "availability" not "reliabilitiy". pg_dump is completely reliable, arguably more so than ANY other backup mechanism as it creates logical machine independent backups.
Best intro tip regarding pg_dump: use -Fc (custom format)