Skip to content
A Bekkour
Writing

Moving Railyard's control plane from Fly.io to one EC2 box

10 min read

Yesterday I moved Railyard’s control plane off Fly.io and onto a single small EC2 instance with a managed Postgres next to it. It had been on Fly since 30 August, long enough for one machine to grow into four plus a dedicated IP address.

Some context on what moved. The control plane is the Rails app that serves the dashboard, owns the database and decides what runs where. Customer apps run on the customer’s own servers, where a small Go agent heartbeats to the control plane over gRPC and the control plane dials it back to run deploys (the split is in two programs and a contract). If the control plane goes down, customer apps keep serving. Deploys and the dashboard stop until it comes back.

One app, four machines

The plan on 30 August was one Fly machine doing three jobs: Puma serving the dashboard on port 3000, the gRPC server for agents on port 50051, and the Solid Queue job worker running inside Puma so I would not pay for a second machine. The fly.toml from that first day says so in a comment: one machine plays three roles. It lasted about a day.

The reason was the grpc gem. It wraps grpc-core, a large C library, and two of its properties matter here.

The first is fork(). When a process forks, the child gets a copy of the parent’s memory but only the thread that called fork. A library that has started background threads, opened sockets and taken locks hands the child a copy of all that state with none of the threads that maintain it. grpc-core does not survive this. Solid Queue’s supervisor starts its workers with fork(), so the job worker can never live in a process that has started the gRPC server.

The second is how it fails. When grpc-core crashes, it segfaults, and a segfault takes down the whole process. No Ruby rescue can catch it. With gRPC embedded in Puma, each crash took the dashboard with it, and on Fly that showed up as a 502 on refresh until the machine restarted. The commit that split it out on 31 August says the trade plainly:

A grpc crash now restarts a ~50MB sidecar; the dashboard stays up.

So the processes separated: one for Puma, one for the gRPC server alone on its main thread, one for bin/jobs. Two days later a fourth joined, a small daemon that forwards logs, which had been a recurring job and worked better as a long-running loop. Same Docker image, four commands.

On Fly a process group is a machine, with its own size, its own restarts and its own line on the invoice. For a typical web app that model is clean and sensible. For mine it meant three 2 GB machines and one 512 MB machine, most of them idle most of the time.

Why gRPC needed its own IPv4

The second cost was the network. Fly’s shared IPv4 addresses route incoming connections by hostname, which the edge reads from the HTTP Host header or the TLS server name. The agent connection carries neither in a form the edge routes on: it is plaintext HTTP/2 on its own TCP port, 50051, authenticated by a bearer token in gRPC metadata. With no hostname to route by, the only way to expose a raw TCP service on Fly is to allocate a dedicated IPv4 to the app, which is billed separately.

Fly was good to me for the first month. Release commands, health checks, private networking and fly deploy building remotely saved real time while the product took shape. My workload was the odd one out: four small processes that cannot share an OS process, plus a non-HTTP port that every agent dials. Each new process meant a new machine, and that is the part I wanted to stop paying for.

What it runs on now

One EC2 instance on Graviton, AWS’s ARM processors, with 2 vCPUs and 2 GB of RAM, running Ubuntu 24.04. Docker Compose on that box runs the same four processes from the same image:

services:
  web:
    image: control-plane:latest
    restart: unless-stopped
    command: ["./bin/rails", "server", "-b", "0.0.0.0", "-p", "3000"]
    env_file: .env.app
    ports: ["127.0.0.1:3000:3000"]
  rpc:
    image: control-plane:latest
    restart: unless-stopped
    command: ["./bin/grpc-server"]
    env_file: .env.app
    ports: ["50051:50051"]
  jobs:
    image: control-plane:latest
    restart: unless-stopped
    command: ["./bin/jobs"]
    env_file: .env.app
  logship:
    image: control-plane:latest
    restart: unless-stopped
    command: ["./bin/rails", "runner", "LogShipper.run"]
    env_file: .env.app

The process split came along unchanged. It was never a Fly requirement; it comes from grpc-core and fork(), and those follow the code to any host. The difference is that four containers on one box cost the same as one. Puma binds only to localhost. The gRPC port is public, because every agent anywhere has to reach it.

The IP problem went away too. An Elastic IP attached to a running instance carries any port, so agents dial the same address as the dashboard, on 50051.

Postgres lives in RDS on the smallest Graviton instance class, single availability zone, with a security group that accepts connections only from the app box. A Postgres container on the same box would have been cheaper, and I wrote it down as a fallback. The control plane’s database is the one thing I cannot rebuild from git, so managed backups and a clean upgrade path won.

Caddy runs on the host and terminates TLS in front of Puma. An AWS load balancer in front of one instance balances nothing, Caddy fetches Let’s Encrypt certificates on its own, and Railyard already runs Caddy on every customer server. The marketing site moved onto the same box as static files:

railyard.run, www.railyard.run {
	root * /var/www/website
	file_server
	try_files {path} {path}/index.html /404.html
}

app.railyard.run {
	reverse_proxy 127.0.0.1:3000
}

http://{$PUBLIC_IP} {
	reverse_proxy 127.0.0.1:3000
}

The last block took three attempts. Before DNS pointed at the box, I wanted the dashboard reachable over HTTPS at the bare IP. Caddy’s tls internal (a certificate from Caddy’s own local CA) tries to install its root certificate through sudo, which fails with no terminal attached, and handshakes died with tlsv1 alert internal error. A plain openssl self-signed certificate worked alone, but a manually configured catch-all on 443 did not coexist reliably with the automatic-HTTPS blocks for the named domains. I tried a bare :443, https:// plus the IP, and the IP with :443. What worked was dropping TLS from the IP fallback entirely and serving it over plain HTTP. It only exists for the minutes when DNS is wrong.

ARM was an easy choice: it costs less for the same memory, and both the Rails image and the Go agent already built for arm64. The image builds on the box itself from an rsynced copy of the code, which mirrors what fly deploy did and needs no registry. A 2 GB machine running bundle install and asset precompilation needs a 2 GB swapfile, and the build competes with the live app for CPU while it runs. That is the choice I will revisit first. Railyard already supports a dedicated build server for customers, and the same feature applies here.

Cost was half the reason. The other half is that Railyard tells customers to bring their own cloud account and keep their machines and data. The control plane now does the same, on a plain VM in an account I own.

The cutover

Fly secrets are write-only, which is the right design for a secret store: fly secrets list shows names and digests, never values, and no API returns one. Most of the values also lived in my local environment file, but a few API tokens existed only in Fly’s store. I read them from the running machine’s process environment before destroying anything. Had I torn the app down first, each one would have needed regenerating at its source.

The database backup failed in a predictable way. Running pg_dump inside the app machine hit a version mismatch: the server was Postgres 18.3, the client bundled in my image was 17.11, and pg_dump refuses to dump a server newer than itself. Tunneling to the database with fly proxy and dumping from a local postgres:18-alpine container worked. The dump held two users, the seeded demo account and one throwaway test signup, so the new database started empty. I kept the dump anyway.

Then I destroyed the three Railyard apps on Fly: the control plane, its Postgres and the marketing site. Scaling to zero does not stop every charge. Destroying the app does, and it released the dedicated IP with it.

Every request returned 500

The first request to the new box failed with PG::UndefinedTable: relation "solid_cache_entries" does not exist, and so did every request after it.

Railyard uses Rails 8’s Solid Cache, Solid Queue and Solid Cable, all backed by Postgres. In database.yml they are configured as separate databases, but all four configurations point at the same physical database through one DATABASE_URL. Rack::Attack, which throttles requests per IP and on logins, keeps its counters in the cache. That puts the cache on every request, so a missing cache table is a full outage.

On RDS, the solid_queue_* tables existed, solid_cache_entries did not, and the schema files for the two are the same shape. db:prepare had loaded the queue and cable schemas and skipped the cache schema without a word.

I had seen this exact outage on 11 September, on Fly, right after adding Rack::Attack. Because the four configurations share one physical database, they share one ar_internal_metadata table, and db:prepare trusted that shared metadata over the presence of the table. The fix then was a rake task that checks the table itself, chained after db:prepare in Fly’s release_command. That chain lived in fly.toml. The new deploy script ran plain db:prepare, because that is what a Rails deploy runs, and the outage came back on the first fresh database.

The fix the second time was to stop depending on any host’s config. The task now runs as part of db:prepare itself:

namespace :db do
  task ensure_cache_schema: :environment do
    config = ActiveRecord::Base.configurations.configs_for(env_name: Rails.env, name: "cache")
    next unless config

    SolidCache::Record.establish_connection(config)
    next if SolidCache::Record.connection.table_exists?(:solid_cache_entries)

    ENV["DISABLE_DATABASE_ENVIRONMENT_CHECK"] = "1"
    Rake::Task["db:schema:load:cache"].invoke
  end
end

Rake::Task["db:prepare"].enhance { Rake::Task["db:ensure_cache_schema"].invoke }

I verified it against a fresh, empty database on the same RDS instance: plain bin/rails db:prepare, nothing chained, and the task fired and loaded the cache schema by itself. Any future host gets it without knowing it exists. When you move hosts, you rewrite the host-specific files from memory, and memory keeps the main command and drops the extra one that fixed an outage twelve days earlier. A fix belongs in the thing every host already calls.

Session Manager instead of an IP allowlist

The new box allowed SSH only from my home IP. Today my home IP changed, and I could not reach production at all. The box now has an instance profile for AWS Systems Manager, and I connect through Session Manager, which authenticates with AWS credentials and does not care which address I am on. The deploy script finds the instance by its tag and tunnels ssh and rsync through it:

HOST="$(aws ec2 describe-instances \
  --filters Name=tag:Name,Values=control-plane Name=instance-state-name,Values=running \
  --query 'Reservations[0].Instances[0].InstanceId' --output text)"
SSH_CONFIG="$(mktemp)"
printf 'Host i-*\n  ProxyCommand aws ssm start-session --target %%h --document-name AWS-StartSSHSession --parameters portNumber=%%p\n' > "$SSH_CONFIG"
SSH="ssh -F $SSH_CONFIG -i $KEY_FILE"

rsync -az --delete -e "$SSH" ./ "ubuntu@$HOST:~/app/"
$SSH "ubuntu@$HOST" "cd ~/app && sudo docker build -q -t control-plane:latest ."

The trade-offs are plain. It is one box in one availability zone with a single-AZ database, so a hardware failure means downtime until I bring up a replacement. I now patch the OS and Caddy myself, which Fly did for me. The bill no longer grows with the number of processes, and a fifth process costs nothing until the box runs out of memory.

The Fly config files are still in the repository as a record of how it was built, and the control plane has answered from the new box since yesterday afternoon.

Al Mokhtar Bekkour

Senior Rails & Go engineer in Quebec. I'm building Railyard, deploy software that runs your whole app on servers you own, and writing here about how it works. Open to work.

← All writing