A cloud firewall broke every deploy, and the fix was to stop dialing Postgres
10 min read
A cloud firewall went onto one of the servers Railyard deploys to, and every deploy to that server started failing before it reached the app, with the same line in the deploy log: connection to <server> port 5432 timed out.
Nothing about the apps had changed. The control plane had been reaching into the server’s Postgres over the public internet on every deploy, and the firewall had just closed that door.
How each app gets its database
Every app on Railyard that asks for Postgres gets its own role and its own database on the server it runs on. The server runs one managed Postgres container, and each app gets a slice of it: a login role with a password, a database that role owns, and grants so it can create tables in it. That carving runs on every deploy, so a redeploy can repair a slice that drifted.
I wrote the first version of that carving as Ruby in the control plane. It opened a TCP connection from wherever the control plane runs (on Fly, at the moment) to the customer’s server on port 5432, logged in as the Postgres admin, and ran CREATE ROLE and CREATE DATABASE. MySQL worked the same way on 3306. It worked in my local Compose stack, it kept working on real servers, and I stopped looking at it.
The firewall had been added for a good reason: bots scan the whole IPv4 space for open database ports, and an exposed 5432 starts collecting login attempts soon after it appears. The owner did the right thing, and Railyard punished them for it.
Two firewalls
A server on a cloud provider can sit behind two different firewalls, and the difference explains the error.
A host firewall runs inside the machine. On Ubuntu that’s usually ufw, a friendly front end for the kernel’s packet filter (iptables or nftables). It sees packets after they arrive at the server’s network interface and decides what to accept. Anyone with root on the box can change it, and so can any program running as root.
A cloud firewall runs in the provider’s network, before traffic reaches the machine. DigitalOcean, Hetzner and Vultr all offer one. You configure it in the provider’s dashboard or API, and the server has no way to see it. A packet it blocks is dropped silently, with no reply at all. So the client doesn’t get “connection refused”, which is what a closed port on a reachable host sends back. It waits for a SYN-ACK that never comes, and eventually reports a timeout. That’s why the error read like a slow network and not a closed door.
Railyard’s agent already manages the host firewall. On every boot it makes sure ufw allows its baseline ports (SSH, HTTP, HTTPS and its own gRPC port), plus any extra ports the operator opened from the dashboard. I had built the database path assuming ufw was the only firewall in play. The cloud firewall is a layer the agent can’t see and Railyard doesn’t manage, and a customer is entitled to put one there.
Why a database port shouldn’t face the internet
If the control plane can reach Postgres from the internet, so can everyone else. A password becomes the only thing between a stranger and the data, every Postgres release that fixes a bug in connection handling or authentication becomes urgent, and the logs fill with failed logins from scanners, which buries the one failed login that would matter.
Docker adds a trap here that catches a lot of people. When a container publishes a port (-p 5432:5432), Docker writes its own iptables rules. It rewrites the packet’s destination to the container’s address in the nat table, which sends the packet through the FORWARD chain, and it inserts its own chains (DOCKER-USER, then DOCKER) at the top of FORWARD, ahead of the chains ufw adds. The packet is accepted before ufw’s rules are consulted. Docker’s own documentation warns that the two don’t combine the way people expect. So “ufw denies everything” does not mean a published container port is closed. A cloud firewall doesn’t have this problem, because it sits outside the machine entirely.
The apps didn’t need the public port either. They connected to their own database through the server’s public address, which is a strange path for two containers on the same machine.
Sending the SQL through the agent
The agent is a Go program that runs on every server and already owns Docker there. The control plane reaches it over one gRPC channel on one port, with mutual TLS and a bearer token. Deploys, restarts, logs and shells already go through that channel, and its port is in the agent’s own firewall baseline. So the fix was to make database provisioning take the same road.
I added one RPC to the agent’s contract:
service Agent {
// ...deploy, logs, exec, restart...
rpc ExecDatabaseCommand(DatabaseCommandRequest) returns (DatabaseCommandResult);
}
message DatabaseCommandRequest {
string engine = 1; // "postgres" | "mysql"
string admin_password = 2;
string database = 3;
repeated string statements = 4; // run in order, stop at the first error
}
message DatabaseCommandResult {
bool ok = 1;
string detail = 2; // stderr on failure, a short summary on success
}
The control plane renders the SQL and ships it. The agent checks that the datastore container is running, then runs the statements inside it with docker exec:
func (s *Server) ExecDatabaseCommand(ctx context.Context, req *pb.DatabaseCommandRequest) (*pb.DatabaseCommandResult, error) {
if err := s.authorize(ctx); err != nil {
return nil, err
}
if exec.CommandContext(ctx, "docker", "inspect", "-f", "{{.State.Running}}", "postgres").Run() != nil {
return &pb.DatabaseCommandResult{Ok: false, Detail: "postgres is not running on this host"}, nil
}
runCtx, cancel := context.WithTimeout(ctx, 60*time.Second)
defer cancel()
cmd := exec.CommandContext(runCtx, "docker", "exec", "-i",
"-e", "PGPASSWORD="+req.GetAdminPassword(),
"postgres",
"psql", "-X", "-q", "-v", "ON_ERROR_STOP=1",
"-U", "postgres", "-d", req.GetDatabase(), "-f", "-")
cmd.Stdin = strings.NewReader(strings.Join(req.GetStatements(), "\n") + "\n")
out, err := cmd.CombinedOutput()
if err != nil {
return &pb.DatabaseCommandResult{Ok: false, Detail: strings.TrimSpace(string(out))}, nil
}
return &pb.DatabaseCommandResult{Ok: true, Detail: fmt.Sprintf("ok, %d statement(s)", len(req.GetStatements()))}, nil
}
A few choices in there matter. The statements go in on stdin, as a script. The admin password goes in as an environment variable of the exec’d process, never as a command-line argument, so it doesn’t show up in ps on the box. ON_ERROR_STOP=1 makes psql stop at the first failing statement and exit non-zero, and whatever it printed comes back in detail and lands in the deploy log where the user can read it. A failed batch is a normal result with ok: false; gRPC errors are kept for “the call itself didn’t work”, such as a bad token or an unknown engine.
The connection psql makes is local to the container. Nothing crosses the network except the gRPC call the agent was already accepting. MySQL takes the same path with the mysql client and MYSQL_PWD.
Idempotent SQL from the Ruby side
Because provisioning runs on every deploy, every statement has to be safe to run twice. Roles are easy: a DO block checks pg_roles and either creates or alters. Databases are awkward, because CREATE DATABASE can’t run inside a transaction block or a function, so it can’t go in a DO block. psql has a meta-command for this, \gexec, which runs a query and then executes each value in its result as a statement of its own. If the database exists, the SELECT returns no rows and nothing runs.
This is the control plane’s side, trimmed:
def provision_postgres(cred)
u, db, pw = cred.username, cred.db_name, cred.password
exec_database_command("postgres", database: "postgres", statements: [
<<~SQL.strip,
DO $do$ BEGIN
IF EXISTS (SELECT FROM pg_roles WHERE rolname = '#{u}') THEN
ALTER ROLE "#{u}" WITH LOGIN PASSWORD '#{pw}';
ELSE
CREATE ROLE "#{u}" WITH LOGIN PASSWORD '#{pw}';
END IF;
END $do$;
SQL
"SELECT format('CREATE DATABASE %I OWNER %I', '#{db}', '#{u}')\n" \
"WHERE NOT EXISTS (SELECT FROM pg_database WHERE datname = '#{db}')\n\\gexec",
%(REVOKE CONNECT ON DATABASE "#{db}" FROM PUBLIC;),
%(GRANT CONNECT ON DATABASE "#{db}" TO "#{u}";),
])
end
def exec_database_command(engine, statements:, database:)
req = Railyard::V1::DatabaseCommandRequest.new(engine:, database:, statements:,
admin_password: server.datastore_password)
res = agent.stub.exec_database_command(req, metadata: agent.auth_metadata, deadline: agent.deadline)
raise Error, "#{engine} addon: #{res.detail}" unless res.ok
res
end
The role and database names are app_ plus the app’s numeric id and the passwords are random alphanumeric strings, all generated by Railyard and never typed by a user, which is what makes building the SQL as a string acceptable here. A second call against the new database grants the role everything on the public schema and sets default privileges, so tables the admin role creates there later are usable by the app too.
\gexec only works because the agent hands psql a script. A Ruby database driver sends statements over the wire protocol and has no meta-commands, so the direct-connection version had to check and branch in Ruby. Moving the work onto the box made the SQL simpler as well.
Apps reach the database on the box
The second half of the change was the connection string each app gets.
Every server Railyard manages has one user-defined Docker network, and the app containers and datastore containers all join it. A user-defined network is a private virtual network on the host. Containers on it reach each other directly, and Docker runs an embedded DNS resolver for it, so a container’s name works as a hostname. (The default bridge network doesn’t resolve container names; only user-defined networks do.) The Postgres container is named postgres on that network, so the app’s database URL now says exactly that:
DATABASE_URL=postgres://app_42:<password>@postgres:5432/app_42
Traffic between the app and its database never leaves the machine and never touches the public interface. Redis got the same change: its URL points at the Redis container’s name on the same network. Existing apps pick up the new URL on their next deploy.
One door into the server
The rule this made explicit: a deploy should never need the control plane to reach the services on a customer’s server. It needs one door, the agent’s, and the agent does the work behind it.
There are other ways to build this. Kamal and Capistrano SSH into the server and run commands, which is a fine model for a tool that runs on your laptop. Railyard uses SSH once, to install the agent on a fresh server; after that, the agent is the management channel. A long-lived control plane holding SSH keys to every customer’s servers is a much bigger thing to protect than one authenticated gRPC port, and a shell hands back strings to parse where gRPC hands back typed results.
The model I had slipped into, the control plane dialing each service directly, couples every deploy to the network path between my infrastructure and theirs. Each new kind of service would mean another port to open and another rule the customer’s cloud firewall has to allow, and every deploy would break the moment they tightened security. Heroku, Render and Fly avoid this because they own the network; Fly apps reach Postgres over a private network the platform runs. On servers the customer owns, the equivalent is to keep app-to-database traffic on the box and keep the control plane outside it, talking to one agent.
The agent also knows things the control plane can only guess. It checks that the Postgres container is running before talking to it and says so plainly in the deploy log if it isn’t. It already had root and Docker on the box, so the database work added no new privileges.
A deploy now needs no inbound database port at all, so the owner who added that cloud firewall can keep it exactly as it is.