Skip to content
A Bekkour
Writing

What a platform does between git push and a working URL

11 min read

The first time I pushed a Rails app to Heroku, a wall of text scrolled past and a minute later I had a URL. I used that workflow for years without being able to say what happened in that minute. Building Railyard means each part of that minute is now code I write.

There are five. A platform builds your code into something runnable, runs a one-off release step, starts your processes with their configuration, routes traffic to them, and rolls back when a version turns out to be bad. Heroku, Render, Railway and Fly all do these five, under different names. Below is each job, then what Railyard does for it today, two weeks after its first commit.

A quick map of Railyard: a Rails control plane decides what to deploy where, and a Go agent on each server does the work, streaming every line of output back over gRPC. For now the “servers” are three agent containers in Docker Compose on my laptop.

Build: turn source into an image

A server can’t run source code directly. It needs the language runtime, your dependencies installed, assets compiled, and a command to start the app. The build step produces one artifact that contains all of that. Today that artifact is almost always a container image: a layered filesystem snapshot plus metadata such as the start command and the port the app listens on.

There are two ways to get from a repo to an image.

The first is a Dockerfile. The repo tells the platform exactly how to build itself, and the platform runs docker build without needing to understand the language at all. This is the Dockerfile of the Sinatra app I use as a test target:

FROM ruby:3.3.6-slim

WORKDIR /app

RUN apt-get update -qq && \
    apt-get install -y --no-install-recommends build-essential && \
    rm -rf /var/lib/apt/lists/*

COPY Gemfile ./
RUN bundle install

COPY . .

EXPOSE 4567
CMD ["bundle", "exec", "puma", "-b", "tcp://0.0.0.0:4567"]

The order matters for speed. Docker caches each instruction as a layer and reuses it if nothing it depends on changed. Copying the Gemfile alone and running bundle install before copying the rest of the code means a change to app.rb reuses the installed gems and rebuilds only the last two layers.

The second way is detection, which Heroku buildpacks made famous. The platform looks at the repo and guesses: a Gemfile means Ruby, a package.json means Node, a requirements.txt means Python. Each buildpack has a detect step (do I recognise this repo?) and a build step (install the runtime and the dependencies). Cloud Native Buildpacks and Nixpacks are the modern versions, and both produce an ordinary image at the end. When detection fails, you’re debugging someone else’s guess about your app.

One detail matters more than it looks: tag the image with the exact commit it came from. A branch name like main moves; a commit SHA doesn’t. If the image tag is the commit (myapp: plus the first twelve characters of the SHA), you always know which code is running, and you can find the previous image when you need to go back.

Railyard builds from Dockerfiles only. The agent runs git clone --depth 1 --branch <ref>, resolves the commit with git rev-parse HEAD, and tags the image with a slug from the repo URL plus the first twelve characters of the SHA. A repo without a Dockerfile gets a line in the log saying the build was skipped, and nothing else happens.

The builds go through BuildKit, Docker’s newer builder. I turned it on after the first public Dockerfile that used --platform=$BUILDPLATFORM failed under the legacy builder with failed to parse platform : "" is an invalid OS component. BuildKit also made the first test repo faster: a cold rebuild of docker/welcome-to-docker went from 39 seconds to 27.

Release: run once, before the swap

Some work has to happen exactly once per deploy, after the new image exists and before any new code serves a request. The classic example is database migrations.

Consider migrating when the app boots instead. Three web containers start at once and all three try to migrate. Rails takes an advisory lock so nothing is corrupted, but a slow migration can make a health check give up on containers that are fine, and new code that serves before the migration finishes queries a column that doesn’t exist yet.

So platforms add a release phase. On Heroku you put a command on the release: line of your Procfile, and the platform runs it once in a one-off container from the new image. The deploy continues only if it exits zero. If the migration fails, the deploy stops and the old version keeps serving.

That ordering has a consequence. During the release step, the old code is still serving traffic against the new schema, so a migration has to be compatible with the code already running. You add a column before the code that uses it ships, and you remove a column only after the code that reads it is gone. The platform guarantees the order of operations. It can’t make an incompatible migration safe.

Railyard has no release phase yet. The Rails 8 test app migrates on boot anyway, because the bin/docker-entrypoint that rails new generates runs ./bin/rails db:prepare whenever the container’s command is ./bin/rails server. With one container per app that works. It’s exactly the pattern above that breaks once there are several.

Run: processes, config and addons

Most real apps are more than one process: a web server, a background worker, maybe a scheduler. Heroku’s answer is the Procfile, a small file that names each process type and its command, such as web: bundle exec puma and worker: bundle exec sidekiq. All of them run from the same image. The worker is the same code as the web process, started with a different command, which is why one build is enough no matter how many process types you have.

Configuration comes in through environment variables. The image is identical in staging and production; only the environment it starts with changes (RAILS_MASTER_KEY, an API token). That’s the twelve-factor rule.

Addons use the same mechanism. When you attach Postgres on Heroku or Render, the platform creates the database, a user and a password, then injects one variable, DATABASE_URL, into every process. Rails reads it with no configuration, and Redis works the same way with REDIS_URL. That indirection lets a platform move the database or rotate its password without a code change.

Railyard runs one container per app. Environment variables are a KEY=value textarea on the app’s form, sent to the agent as a map and passed to docker run as -e flags. There are no addons: the Rails test app gets a DATABASE_URL I typed by hand, pointing at the Postgres container in the same Compose network, and the agent attaches the app’s container to that network so the hostname resolves. Its home page renders the Postgres server version from a live query.

Route: proxy, TLS and the swap

A running container listens on a port on some server. Users type a hostname. Routing connects the two.

In front of the containers sits a reverse proxy (Nginx, Caddy, Traefik, or the platform’s own router). It accepts each request, reads the Host header, and forwards the request to the container that serves that hostname. It’s also where TLS ends: the proxy holds the certificate, decrypts HTTPS, and talks plain HTTP to the container over the local network.

Certificates come from Let’s Encrypt through a protocol called ACME. The proxy asks for a certificate for myapp.example.com, Let’s Encrypt asks it to prove control of that name (usually by serving a token over plain HTTP on port 80), and if the proof works it issues a short-lived certificate. Caddy runs the whole exchange and the renewals by itself, given only a hostname. For it to work, DNS has to point at the server and port 80 has to be reachable from the internet.

The proxy is also what makes a zero-downtime deploy possible. The sequence is to start the new container next to the old one, health-check the new one until it answers correctly, point the proxy at it, and then let the old container finish its in-flight requests before stopping it. If the health check fails, you stop the new container and nothing else has changed. In Caddy, a route with an active health check looks roughly like this:

myapp.example.com {
	reverse_proxy app-42-v7:3000 {
		health_uri /up
		health_interval 2s
		health_timeout 1s
	}
}

Swapping versions means pointing the upstream at the new container and reloading, which Caddy does without dropping open connections.

Railyard has no proxy yet, no hostnames and no TLS. Each app gets host port 4000 + application_id, and the dashboard prints http://localhost:4001 and calls it a URL. And the order of operations is the reverse of the safe one. This is the run step, trimmed:

_ = sendSystem(stream, "Replacing container "+name)
_, _ = s.runStep(stream, "docker", "rm", "-f", name)

args := []string{"run", "-d",
	"--name", name,
	"--restart", "unless-stopped",
	"-p", fmt.Sprintf("%d:%d", req.HostPort, req.ContainerPort),
}
if req.Network != "" {
	args = append(args, "--network", req.Network)
}
for k, v := range req.EnvVars {
	args = append(args, "-e", k+"="+v)
}
args = append(args, tag)

if code, err := s.runStep(stream, "docker", args...); err != nil {
	return err
} else if code != 0 {
	return s.sendCompleted(stream, code)
}

The old container is removed before the new one starts. Every deploy has a window where nothing is running, and if the new version is broken, the old one is already gone. It works this way because without a proxy there’s one host port per app, and two containers can’t bind the same port at once.

What protects the app instead is a health gate. docker run -d returns as soon as the container is created, and the first version of the pipeline reported success at that moment. I clicked the link and got connection refused: the app had already exited on a missing environment variable. Now the agent waits two seconds and asks Docker whether the container is still up:

time.Sleep(2 * time.Second)

out, err := exec.CommandContext(stream.Context(),
	"docker", "ps", "-a",
	"--filter", "name=^"+name+"$",
	"--format", "{{.Status}}").Output()
if err != nil {
	_ = sendSystem(stream, fmt.Sprintf("docker ps failed: %v", err))
	return false, nil
}

status := strings.TrimSpace(string(out))
if !strings.HasPrefix(status, "Up") {
	_ = sendSystem(stream, fmt.Sprintf("Container is %q, capturing logs:", status))
	_, _ = s.runStep(stream, "docker", "logs", "--tail", "200", name)
	return false, nil
}
return true, nil

When the container has died, the last 200 lines of its output go into the deploy log, so a failed deploy shows the crash reason instead of a red badge. That catches the app that dies at once. It doesn’t catch a Rails app that takes eight seconds to boot and then crashes, and it doesn’t check that the app answers HTTP at all. A container can be “Up” and return 500 on every request.

The fix for both problems is the same piece, which is why it’s next. With Caddy on each server, the new container can start next to the old one on its own port, the agent can probe it over HTTP until it answers, and only then move the route and stop the old container. A failed deploy then leaves the previous version serving, and the health gate becomes a readiness check that waits for a real response.

Rollback: deploy the previous image

Rollback falls out of the build and route jobs, provided you kept the old image. Rolling back is a deploy of the previous image with the previous config: start it, health-check it, swap. No rebuild, so it takes seconds, and tagging by commit gives the platform a name for “the version that worked”.

What rollback can’t do is undo a migration. If the bad release dropped a column, the previous image will query a column that’s gone. That’s the other reason migrations have to be compatible with both versions: the old code may need to run again.

Railyard has no rollback yet. The old images are still on disk, tagged by SHA, but nothing knows how to start one, and the agent clones with --branch, which accepts a branch or a tag but not a bare commit. Redeploying “the commit from yesterday” isn’t possible yet.

The order to build them in

Writing the five jobs out made the order obvious. The proxy comes first, because the safe swap, the HTTP readiness check and rollback all depend on starting a new container next to the old one. The release phase comes next, so migrations run once and before the swap. Then multiple processes from a Procfile, then Postgres and Redis as addons that inject their own URLs.

Under all of it, the deploy runs in a Ruby thread inside Puma, so restarting Rails mid-deploy loses it. A job queue has to fix that before any of this leaves my laptop.

Today, a good deploy of the first test repo takes 27 seconds and ends with a container on localhost:4001, and a deploy whose new container crashes takes the app down with it.

Al Mokhtar Bekkour

Senior Rails & Go engineer in Quebec. I'm building Railyard, deploy software that runs your whole app on servers you own, and writing here about how it works. Open to work.

← All writing