Securing a Postgres Database
goteleport.com
goteleport.com
This post seems to outright state that by default postgres is listening to everyone via TCP for connection.
This is not true.
Unless you edit pg_ident.conf, your postgres install will not listen for connections outside of on localhost. So, while it's correct to say that it listens to TCP port 5432, there is a very narrow limit to whom it's listening for. Namely, the same machine.
Postgres is pretty secure by default. It doesn't allow external connections. It also requires a username and password to connect with permissions on the database you're connecting to.
Compare that to something like redis. They at least used to listen for connections external by default and not even have a password to connect. I can imagine it's still very easy to find all kinds of interesting stuff snooping around on port 6379.
I don't know what the defaults are, but pg_ident.conf has absolutely nothing to do with this. The main configuration file (I think postgresql.conf usually) has listen_addresses, which controls the addresses on which postgres listens, as you might guess.
pg_hba.conf (not pg_ident.conf) controls the authentication methods the server asks from the client, depending on how they're connecting.
(I think I recall exactly one in the history of PostgreSQL since I started using it, but it is what it is.)
Not anymore since version 3.2.0 [0]
>Unfortunately many users fail to protect Redis instances from being accessed from external networks. Many instances are simply left exposed on the internet with public IPs. For this reasons since version 3.2.0, when Redis is executed with the default configuration (binding all the interfaces) and without any password in order to access it, it enters a special mode called protected mode. In this mode Redis only replies to queries from the loopback interfaces, and reply to other clients connecting from other addresses with an error, explaining what is happening and how to configure Redis properly.
What is the advantage of listening on localhost compared to using a socket, with free access control?
Can install and use right away locally without figuring out where your distro puts the socket at.
Edit: Also no need to play with permissions of the socket in such a case.
Put another way, any user can become root through privilege escalation, so access control is pointless, since any untrusted user can take over the machine.
The real unit of security is the whole OS (VM), not its internal user boundaries.
Also, the loopback is used as a networking interconnect or guest->host channel for sandboxed containers and VMs, so it's security sensitive in this way.
Layers of security exist for this exact reason. If one fails, hopefully something else stops it.
Now, you probably want to allow services to connect without needing to speak ssh - so you probably do want a bind whitelist, an IP whitelist and ssl - but the article is off to a bad start by being wrong about defaults.
https://github.com/postgres/postgres/blob/master/src/backend...
IE; the 'postgres' unix user is required to access databases as the 'postgres' database user.
This is the case for the official docker image at least, and I'm fairly certain also true of distro package managers
Is a reverse tunnel really air gapped?
I thought AG meant one had to physically touch the device and transfer using devices without any network capability, such as a flash drive?
Perhaps we need another somewhat similar term. It's like null-routing or firewalling devices on your network.. they're technically "connected" but if they cannot dial out they're in some ways gapped. (This is handy for dubious quality IoT devices, they can't phone home, auto-patch to drop features, or share your usage information with $corp, but still respond to local network commands).
To some extent tunneling is security through obscurity (an SSH tunnel has moved the port you need to secure from 5432 to 22)
Not sure what they mean by "out of the box", but you can make the listen_addresses list empty:
"listen_addresses (string) ...If the list is empty, the server does not listen on any IP interface at all, in which case only Unix-domain sockets can be used to connect to it..."
https://www.postgresql.org/docs/9.3/runtime-config-connectio...
I disagree with the author and agree that indeed a fresh install supports (i.e. allows one to configure without rebuilding) not talking over a network at all.
The article appears to say that a database/DBMS in the ideal world is not accessible over network at all. That is, apparently only accessible by users who have physical access to the machine it is on.
I then use a bastion host when I need to access ssh on the instance. The bastion host remains off and inaccessible except for when I need to perform maintenance.
The advantages of this is that there is no "always-open" access to the instance.
Not sure why the author does not advocate this.
Email is one such way.
I'm sorry but could you clarify? You mean that you monitor and/or collect data from hosts inside of a network via email somehow?
Your inbox needs to be an automatable inbox, say controlled by bleeder or another flowengine. From there you can build a full messaging dashboard and have it done in something that could pass a very stringent audit.
Bastion<->app host<->DB
Bastion can talk to the net.
App host can talk outbound but inbound only accepts bastion and DB.
DB can only talk with app host.
Obviously, you harden everything appropriately... But with this arrangement, it's very difficult to penetrate this sort of network. Think of it as a network that as a whole is default-deny.
Probably because the blog post is an advertisement for their product, which already allows you to implement bastion hosts as you describe.
From the bottom:
> Databases do not need to be exposed on the public Internet and can safely operate in air-gapped environments using Teleport’s built-in reverse tunnel subsystem.
Not sure what there reverse tunnel product is, but a bastion host is super easy to implement, just spin up an ec2 and walla.
Curious as to what value they are providing
Haven't used them myself but I wouldn't be against trying it if in the market for something like that.
The same is true for server applications that have weird 3rd party dependencies that may go down when you least suspect it.
https://www.collinsdictionary.com/dictionary/french-english/...
In cloud environments, it's straightforward to update network firewall rules as IP addresses change. Residential and office IP addresses don't change much so it's not much of a hassle in my experience. That said, it can get annoying if you find your self working on a network that rotates your IP address frequently (e.g., a hotspot).
Edit: To add some details - using ProxyJump you don’t have to expose anything to the jump host and instead just proxy through it.
You could then use a bastion to access servers in the private subnet, or use something like AWS Session Manager which provides command line access via web browser in lieu of a bastion.
Adding more mechanisms on top is pointless when the effort could be invested in, for example, automated auditing of SGs, which is vastly more potent from a hardening perspective than adding additional layers of technical redundancy that are still exposed to the same flawed human processes.
When you reach a team of 10-20 folk on a project, stuff tends to get confusing and/or lazy with elaborate configurations. Security design therefore is about more about managing that outcome through simplicity and process hardening than.. well.. I don't even know what threats a separate VPC protects against
From the perspective of maintaining security in the least confusing way possible in a team setting, I could see a scenario where you have a VPC that only has private subnets, and no public subnets at all.
You could call it “Database VPC” and assuming your team doesn’t reconfigure the VPC to add public subnets, you can be comfortable with your team adding more EC2 servers / databases / whatever in that VPC since they would all be in private subnets inaccessible by the public internet.
And then you could have a 2nd VPC with private/public subnets that are less locked down than the database VPC, which might contain your load balancers, application servers, etc.
I suppose the benefit would be the logical separation of databases into a VPC without any public subnets. And the only way to gain access to subnets in that VPC would be through VPC peering.
(The above assumes you’re using AWS). Although after typing all that, I still think you can accomplish a comparably secure (on a network level) architecture using security groups, or even separate subnets, without separate VPCs.
In general I try to avoid multiple VPCs because VPC peering in AWS can get tricky (or impossible if both VPCs have overlapping CIDR blocks).
I also limit access to just one IP, which is a VPN server hosted on a different cloud provider.
You can run a bulletproof VPN server easily using something like algo (https://github.com/trailofbits/algo)
People can make out-of-band updates to the query, like making it more efficient or migrating it, without requiring any changes to the all.
That’s not a rag on stored procedures more generally. Just that specific use case scales poorly to a large number of operations as its inherently cpu bound.
A simple sha256(lower(email)) is equally secure as a complete random salt, the only requirement on a salt is to be unique.
It is quite possible to test stored procedures, it's just requires some additional work/infrastructure.
It's not as simple, you still have to make sure public backend user can't access web logs, that may reveal session id of an backoffice admin account or other information useful for breaking in other parts of the website.
Eventually, the attacker will gain access, but this is useful to slow him down enough, so there's a chance he'll give up, or you notice something suspicious.
For example pgbouncer allows you to connect via multiple roles/credentials to a database. It's just about how you configure your pool.
Of course, this is a security through obscurity type of approach, and you'd want to secure your database whether or not somebody knew where it was running. But there's a difference between somebody seeing that you've just created `staging-psql.foo.com` that you might still be configuring, and the passive background noise of internet port scanning that's a little less targeted.
My (our) problem is that we use a lof ot AWS lambda functions that read and write to the database, and they always execute from different (dynamic) IPs, so what is the best solution in this case?
For blog post masquerading as a how-to on securing Postgres, the fact the article only fleetingly mentions the word "function" once is not cool.
One of the biggest things you can do for Postgres (or any database for that matter) is enforce the use of stored procedures ("functions" in PG-speak) rather than direct SQL queries.
SQL Injection attacks are common as muck. Stored procedures are a quick and easy way to mitigate them.
Stored procedures also have the added bonus of allowing the DBA to remain in control and ensure a higher quality of SQL query rather than upstream devs sending all manner of unoptimised SQL queries.
The application pods ingest the db credentials from Vault. The biggest concern we have today is automating credential rotation.
Curious if anyone else has a similar setup or thoughts on ours?
https://www.hashicorp.com/resources/securing-databases-with-...
Usually DB servers are "always on" and always accept connections, so the reverse tunnel also needs to be always up.
Regarding row-level security: that sounds quite awesome, but in the form it's described in the article, very limited in the number of use cases.
Quite often you have a web app that talks to the DB and that uses a service account. So as with, with row-level security you can just allow or disallow things to the service account, not to the user logged into the web application.
Is there a way to drop from the service account into a less-privileged role inside a transaction or so?
This work was done because the Teleport users who used it for SSH kept asking for the same access for their databases. The reasoning goes like:
1. Setting up a single proxy gives you the same benefits for N databases as they come online. No need to manage additional endpoints (public IPs, ports, etc).
2. You have the same centralized place to manage auth/authz for all users.
3. This allows to connect to databases on the edge, where there isn't an opportunity to have a permanent public IP and locations frequently go online/offline.
4. Finally, it's nice to have unified visibility into what's available (for users) and centralized logging/audit for the security team.
As always, all of this is possible with other tools. The world of open source is vast and full of options, but we were hoping to make it simpler, with less configuration and moving parts.
Obviously you need to trust the service account enough to do that.
(1) When is operating your own PostgreSQL instance desirable?
(2) Isn’t this equivalent to running an RDS instance in its own VPC with public access turned off and only allowing comms with app VPC?
For example: how much of this would I need not worry about if I use a managed postgres database (like from digitalocean or aws or any provider actually)
Speaking of monitoring: there are far fewer options if you're on RDS. Personally my favorite is munin, where you can instantly see all kinds of stats & history at multiple levels of abstraction. There are many excellent Postgres plugins to report on transactions, locking, etc., and it's easy to write your own.
On RDS you can't install custom extensions. They have a whitelist of the most commonly-used ones, but if you find a different one (or build your own), you're out of luck. This really hampers you if you want to get the most from your database. I will say though, building custom extensions can also block you from using CI/CD solutions (or at least make them harder to set up), since they may have similar restrictions. Writing a custom extension is sort of a last restort, but it can be a huge boost for certain problems.
RDS also doesn't grant you direct access to the WAL. That means you can't use WAL-E/WAL-G (a really nice incremental backup solution) or many other helpful tools. You can't do replication except via AWS's own black-box features. (This can be especially annoying when you do upgrades.) It also matters because RDS only gives you 30 days of backups. Tons of businesses need more than that, and it's hard to achieve without using pg_dump. But for large databases, pg_dump can take hours and impact performance of other connections.
On RDS you also have to deal with EBS expense and performance limitations. PIOPS are very expensive. Running on local disks is a lot faster. On plain EC2 you can do more to work around all that. Ephemeral storage gives you real disks, but they might not be big enough (and they don't scale independently of the instance size). A better approach is RAIDing over gp2 volumes. I've heard they did this at Reddit and Citus. I've set it up before and it has worked great. You might be able to find some details in my comment history. Go back a few years. . . . Of course if running in your own datacenter is an option, that's even simpler (in some dimensions anyway). To me the sweet spot is renting dedicated machines, e.g. from Wholesale Internet.
RDS is sooo easy though. I have to admit it's hard not to recommend it for early-stage ventures. The limitations probably won't bite you until you're far along.
If our Enterprise pricing is too high, take a look Teleport's open core version:
Do still use SSL and password authentication too; a VPN alone isn't a complete solution.
Is this something people really do? Wouldn't these logs be enormous and/or potentially leak unwanted information?
So we use it in development but not in production.