Skip to content
A Bekkour
Writing

Hardening a deploy platform against apps I didn't write

11 min read

Railyard’s deploy pipeline has a green test suite, and for most of this year that suite caught almost none of the bugs that matter most. Those bugs live between the pipeline and a repository someone else wrote, and every repository in the suite was one I wrote.

So in September I started deploying well-known open-source Rails apps through Railyard exactly the way a customer would: Mastodon, Discourse, Chatwoot, Forem, alonetone, Bike Index, diaspora and a couple of dozen more from the awesome-rails list. Paste the GitHub URL, add a database, press deploy, read what happens. This is the method I settled on.

The finish line

A platform test needs a definition of done that the platform cannot fake. Mine is an HTTP request to the deployed app that returns a rendered page: a 200 with <title>Mastodon</title> in the body. A green deploy badge does not count, because early on the platform marked deploys succeeded whose new container never became reachable.

The second half of the rule is to drive one app all the way to that state before starting the next. My first attempt ran the whole list as a batch, and it produced a table of first failures: two dozen shallow errors with no view of what sat behind each one. Working one app to completion is slower per day and much faster per bug, because each rerun gets one step further and exposes the next thing in line.

alonetone, a music hosting app, shows why. Driving that single app to a rendered page on 15 September surfaced six platform bugs in a row. A cleanup step deleted the app’s own lib/ directory when the source was a plain directory copy instead of a git clone. The synthetic commit ID for a directory-copy deploy was a timestamp, and the image tag truncated it to its first twelve characters, which kept only the year and month, so every rebuild that month reused the first image and every fix after it was invisible.

The app arrives as-is

A customer’s repository reaches Railyard exactly as it is in git. Nobody on my side gets to edit it first. So every fix I make has to live on Railyard’s side: a build-time fixup the agent applies to any repo with the same shape, a default it injects, routing behaviour, or a setting the operator changes in the dashboard. A change to the app’s tracked files is never allowed.

The rule changes what you look for. Mastodon, once it booted, rejected every request with Rails host authorization. Editing config.hosts would have made it work and proved nothing. Reading its config/initializers/1_hosts.rb showed that it builds the allowed hosts from LOCAL_DOMAIN and WEB_DOMAIN, which is Mastodon’s documented self-hosting convention. Setting LOCAL_DOMAIN through Railyard’s environment variable API is ordinary operator configuration, the same thing every Mastodon admin does. The next deploy returned a 200 with Mastodon’s title, from the untouched repository.

When an app cannot be deployed without editing it, that is a finding too, and it goes in the notes as “needs the operator’s own action”. alonetone imports a licensed animation plugin that is not in its public repository. No platform can supply that file.

Fixing the class

A failure in one app is a sample. What I fix is the class it belongs to: the set of repositories that share the shape that caused it.

Mastodon, after a fix that got it through a clean build, failed routing. The new container never became healthy and Railyard rolled it back. The clue sat a few lines earlier in the deploy log: a note that no release command was configured, so the release step was skipped and “any pending migrations will NOT run”.

A release command is the one-off step that runs between building an image and switching traffic to it. For Rails it is where migrations run. When Railyard detects Rails in a repository, it proposes db:prepare as the default:

def rails(files)
  if files["Procfile"].present?
    processes, release = split_procfile(files["Procfile"])
    release_command = release || "bundle exec rails db:prepare"
  else
    release_command = "bundle exec rails db:prepare"
    processes = [{ name: "web", command: "bundle exec rails server -p ${PORT:-3000} -b 0.0.0.0" }]
  end

  Detection.new(framework: "rails", release_command:, processes:)
end

That covers apps created from the dashboard form. For an app that reaches a build with no release command at all, the agent on the server applies the same default as one of a handful of Rails safety nets. Those safety nets had been wired into the Nixpacks and buildpack build paths and never into the path that builds a repository’s own Dockerfile. Mastodon ships a Dockerfile. So it got no migrations, booted against an empty schema and crash-looped. Bike Index, also Dockerfile-based, failed the same way in the same run.

The Mastodon-shaped fix was to set a release command on the Mastodon app and move on. The class was every Rails app with its own Dockerfile and no explicit release command, and every one of those had been getting zero migrations on every deploy. The deploy itself reported success. That is a much worse bug than “Mastodon doesn’t deploy”, and it only became visible because a famous Dockerfile-based app went through the pipeline.

The class also had a history. The day before, I had fixed the same shape for a different helper, a Gemfile.lock platform fixup that ran on two of the three build paths, and had not checked its siblings. This time I moved all of them above the point where the build paths split, so no build route can skip them, including whichever one gets added next.

A regression test for a class has a different shape from a test for one bug. It runs the whole class, so the next build path someone adds is covered the day it appears:

func TestRailsSafetyNetsRunForEveryBuildMethod(t *testing.T) {
	for _, method := range []string{"dockerfile", "nixpacks", "cnb"} {
		t.Run(method, func(t *testing.T) {
			dir := railsAppFixture(t) // config/environment.rb, no release command
			req := &pb.DeployRequest{BuildMethod: method}

			prepareSource(&recordingStream{}, dir, req)

			if req.ReleaseCommand != "bundle exec rails db:prepare" {
				t.Fatalf("%s: release command = %q", method, req.ReleaseCommand)
			}
		})
	}
}

Some bugs only exist at this level. When I pointed a Discourse app at a fork of Discourse, the next deploy reported success and built the original repository. The agent keeps a cached working directory per app and ran git fetch in it without updating the remote, so changing an app’s git URL did nothing. Anyone who renamed a repo, transferred it to an organisation or switched to a fork would have kept shipping the old code with nothing in the UI to say so. The fix is one git remote set-url origin before every fetch.

The second occurrence

The counterweight to fixing classes is not inventing them. One failure shows that one app broke. It does not show the shape of the class, and a generalisation from a single point usually picks the wrong axis.

So a general fix waits for the second occurrence. retrospring failed on a gem whose native extension needed a system library the build image lacked. I logged it and left it. When diaspora failed the same way, with a different gem and a different library, there were two points to draw a line through, and the fix became a general mechanism for native-extension system dependencies instead of a patch for one gem. alonetone’s Puma config bound to a unix socket, which cannot work when the proxy runs in a different container. Bike Index had done the same through its Procfile the day before. The second occurrence is what turned it into platform work.

Waiting has a cost: some single failures sit in the notes for days. A logged failure costs nothing to keep, though, and a speculative fix is code to maintain for a case nobody may hit again.

A flat count with real progress

In the early hours of 15 September my local run had 2 of 27 apps deployed. Over the next few hours I fixed eight bugs, each with a regression test and each confirmed against the app that found it. When I counted again, it was still 2 of 27.

The number had not moved, and the pipeline was in much better shape. Mastodon had been failing at the build. After its fix it cleared the build and failed at release, on the missing migrations above. A fix like that moves the failure later in the pipeline, where the next one is waiting. A pass/fail count says something stopped each app. It does not say where. An app that fails later than it did yesterday is progress that scores the same.

So the record I keep per app is the furthest step reached and the exact log line that stopped it.

The fix that wasn’t running

The count can also lie in the other direction, and the worst instance of that was mine.

Railyard has two programs. The Rails control plane serves the dashboard and decides what runs where. A Go agent on each customer server does the building and running. Agents update themselves: the control plane advertises the agent build it ships with, and each agent downloads and installs that build when it differs from its own. So the agent version in production is whatever binary was staged into the control plane’s image at deploy time.

On 11 September I found a server whose agent binary was dated the day the server was attached. Its journal had never once logged a self-update. Deploys had been going out without the step that stages a fresh agent build, so self-update had nothing new to offer. I fixed the staging and watched that server update itself for the first time.

Four days later a production deploy log for alonetone contradicted my own notes. I had recorded alonetone as fixed apart from the licensed plugin. The log showed it failing on a different error: the build ran Node 18 and one of its dependencies required Node 20 or newer. The agent’s Node version defaults had been merged and tested the day before. They were not on the servers.

I traced it to the release process. A control plane deploy rebuilt the Rails image from current code, but the agent binaries inside it came from a local staging directory that was ignored by git and refreshed only when someone ran the agent build first. A Rails-only release shipped whatever agent build was sitting there. Every agent fix merged since that build was absent from production, and nothing anywhere said so.

The fix is a check that runs before every control plane deploy. The build script stamps each agent build with the date and the repository commit it was built from. The check compares that stamp with the latest commit that touched the agent’s source:

#!/usr/bin/env bash
# bin/verify-agent-build [VERSION_FILE]
set -euo pipefail
ROOT="$(cd "$(dirname "$0")/.." && pwd)"
MARKER="${1:-$ROOT/dist/agent/VERSION}"   # "<date>-<short commit>"

latest="$(git -C "$ROOT" log -1 --format=%H -- agent/)"
[ -n "$latest" ] || exit 0
[ -f "$MARKER" ] || { echo "agent never built: run bin/build-agent" >&2; exit 1; }

stamp="$(tr -d '[:space:]' < "$MARKER")"
built="$(git -C "$ROOT" rev-parse --verify "${stamp##*-}" 2>/dev/null)" \
  || { echo "stamp $stamp is not a commit here: rebuild" >&2; exit 1; }

git -C "$ROOT" merge-base --is-ancestor "$latest" "$built" && exit 0
echo "agent build $stamp predates the latest agent change: rebuild" >&2
exit 1

The comparison has to be an ancestry check. The stamp records the whole repository’s HEAD at build time, which is usually a later commit than the last one that touched the agent. String equality would call almost every good build stale. merge-base --is-ancestor asks the right question: does the build include the newest agent change?

The deploy script runs bin/verify-agent-build before anything else and refuses to continue when it fails. The local test harness got the same check at startup, rebuilding its agent containers when the stamp is behind. When I first ran the check, it reported the staged production build as stale, behind the latest agent commit. After rebuilding and deploying, I confirmed on the control plane that the advertised agent version matched the version all three attached servers reported.

The incident also changed what “fixed” means in my notes. A fix counts when a fresh build of current code reproduces it against the app that found it. A merged commit, a passing test and a finished deploy are each necessary and none is sufficient.

Failures on the other side of the line

A good share of failures were never Railyard’s to fix, and recording those carefully matters as much as the bugs. splits-io and rubygems.org break on their own Dockerfiles, which I covered in building repos you didn’t write. retrospring and diaspora both failed inside rbenv install’s OpenSSL step with no output at all. I reproduced the same install alone in a fresh Ubuntu container, where it succeeded in about 15 seconds. The difference was my development machine, which at one point had about 64 MB of free memory while the batch ran.

Discourse is next on the list: its production bundle still has no application server, and it is the most famous app that has not reached a 200 yet.

Al Mokhtar Bekkour

Senior Rails & Go engineer in Quebec. I'm building Railyard, deploy software that runs your whole app on servers you own, and writing here about how it works. Open to work.

← All writing