Four ways to build someone else's repo
10 min read
Railyard takes a git URL and has to produce a container image from it that boots, without asking the person who owns the repo to change anything. On 12 September one line in Discourse’s Gemfile cost me five commits in about two hours, and the day before, a buildpack image ran bundle install when I had asked it to start Puma.
Both bugs live in the tools a platform uses to build code it didn’t write. So before the stories, the tools.
The four options
A container image is a filesystem plus a command to run. Something has to decide what goes into that filesystem: which base OS, which language runtime and version, which system packages, which build steps, and what to start at the end. When you deploy your own app you make those decisions once. A platform has to make them for every repo it sees, and there are four ways to do it.
The most direct is the repo’s own Dockerfile. If the author wrote one, the platform runs docker build and uses what comes out. This is the most faithful option, because the author already chose the base image, the packages, the build steps and the start command.
You also inherit every assumption the author made. rubygems.org has a well-made multi-stage Dockerfile; the image built cleanly in about 16 minutes and then db:prepare failed, because an initializer connects to OpenSearch. splits-io’s Dockerfile pipes a NodeSource setup script for Node 12 into bash, and that script no longer exists. Neither is a build problem a platform can fix. A Dockerfile also says nothing about what the app needs around it: which databases, which environment variables, what to run before the new version takes traffic.
The second option is a builder that infers a build plan. Nixpacks, from Railway, looks at the repo (a Gemfile means Ruby, a package.json means Node), writes a Dockerfile, and builds it. Railway is moving to a successor called Railpack that takes the same approach. The useful property is that the output is a plain Dockerfile. When something breaks you can open .nixpacks/Dockerfile and read the line that failed, and a platform can patch that file before it builds.
The inference is only as good as the language provider, and the Ruby provider installs Ruby with rbenv and ruby-build, which compiles the interpreter from source. On a 1 vCPU test server the build log said make -j 1, and compiling Ruby used up the 20-minute build budget before bundle install started.
The third option is Cloud Native Buildpacks, the descendant of Heroku’s git push heroku main. You run pack build with a builder image (I use heroku/builder:24), and a series of buildpacks detect the app and assemble the image from prebuilt layers. Ruby arrives as a prebuilt binary, so nothing compiles.
The cost is that there is no Dockerfile to read or patch. The image is assembled by the buildpack lifecycle and runs under a launcher process with its own rules. The run image is also smaller than the build image: git exists while building and not at runtime, which breaks any Rails app with a vendored gem whose gemspec calls git ls-files, because Bundler evaluates gemspecs on every boot. I also lost days to pack being unable to export into Docker’s containerd image store, a known pack issue. I worked around it by running a registry on loopback on the server, publishing the image there, and pulling it back.
The fourth option is to skip building. The customer builds an image in their own CI, pushes it to a registry, and the platform runs docker pull. Nothing compiles on the server that serves traffic, rollback is pulling an older tag, and the bytes that passed CI are the bytes that run. The customer now owns a build pipeline and a registry, and “push to deploy” has their CI in the middle of it.
Railyard can build all four ways. Most of my bugs come from the seams between them, and from tools behaving differently from their documentation’s happy path.
A version range in a PATH
Discourse is one of the largest open-source Rails apps. It ships no Dockerfile, no .ruby-version and no Procfile. Through the buildpack route it had already gotten further than I expected on a 2 GB server: Ruby, all the gems, the whole Ember asset build, no out-of-memory kill. It stopped compressing assets, because its lib/tasks/assets.rake shells out to a brotli binary the builder image doesn’t carry. On 12 September I tried the Nixpacks route, and spent the afternoon on a chain where each fix uncovered the next bug.
The Gemfile line at the start of it, and what I ended up doing about it:
# Discourse's Gemfile
ruby "~> 3.4"
# What the agent does before Nixpacks runs (simplified)
def resolved_ruby_version(gemfile, has_ruby_version_file)
return if has_ruby_version_file
raw = gemfile[/^\s*ruby\s+["']([^"']+)["']/, 1] or return
numeric = raw.sub(/\A[~><=!\s]+/, "").strip
return unless numeric.match?(/\A\d/)
return if numeric == raw && numeric.count(".") >= 2 # already exact, e.g. "3.3.6"
numeric # "~> 3.4" becomes "3.4", written to .ruby-version
end
~> 3.4 is ordinary Bundler. It means any 3.x from 3.4 up. The first failure was a Docker syntax error with no visible connection to Ruby: it couldn’t find = in a string ending in ruby-~>. Nixpacks’ Ruby provider takes the version string from the Gemfile and drops it, unresolved, into a PATH segment of the form ruby-<version>. With 3.4.1 that’s fine. With ~> 3.4 the PATH entry contains a space and a >, and the generated Dockerfile stops parsing.
So I resolve the range before Nixpacks sees it. If the repo has no .ruby-version and the Gemfile’s ruby line isn’t an exact version, the agent writes one into the checkout. My first version padded the result, so ~> 3.4 became 3.4.0. That got past the syntax error and into the next failure. Ruby 3.4.0 compiled, and then bundle install failed with Errno::ENOENT under lib/ruby/3.4.0+1/. ruby-build’s 3.4.0 build has a broken gem install step that later 3.4 patches don’t have. By padding with .0 I had picked the oldest and worst release in the range.
The second fix was to stop padding. ruby-build ships X.Y definitions that resolve to the newest patch it knows for that minor line, and Nixpacks installs a fresh ruby-build on every build. A bare 3.4 tracks new patches without anyone hardcoding a number. On that build it resolved to 3.4.10.
That exposed the third bug, and the fourth was waiting in the same generated file:
# Generated by Nixpacks (trimmed)
ARG GEM_HOME GEM_PATH
ENV GEM_HOME=$GEM_HOME GEM_PATH=$GEM_PATH
RUN rbenv install 3.4 && rbenv global 3.4
# After the agent patches it
RUN rbenv install 3.4 && \
rbenv global "$(rbenv versions --bare | grep -E '^3\.4(\.|$)' | tail -1)"
COPY . /app/.
RUN rbenv versions --bare | grep -E '^3\.4(\.|$)' | tail -1 > /app/.ruby-version
rbenv install 3.4 understands the alias, because ruby-build resolves it. rbenv global 3.4 doesn’t: rbenv’s version commands want the exact directory name of an installed version, so it failed with rbenv: version '3.4' not installed immediately after installing it. The patch makes global use whatever rbenv versions --bare reports.
The fourth bug was the two lines at the top. Nixpacks declares GEM_HOME and GEM_PATH as build arguments for every Ruby app whether or not anyone supplies them. Railyard never passed them, so Docker defaulted them to empty strings and baked GEM_HOME="" into the image. RubyGems treats an explicitly empty GEM_HOME as a real install location, so the first gem install failed trying to create a directory with an empty name. The agent now removes those two tokens from the generated ARG and ENV lines and leaves everything else on them alone, so RubyGems falls back to its own default. That one had nothing to do with Discourse. It would have hit every Ruby app built through Nixpacks on Railyard.
The fifth was the .ruby-version file I had written myself. It was copied into the image with the bare 3.4 in it, every later RUN step started a fresh shell, rbenv read the file, and bundle install failed with the alias error again, this time reported as set by /app/.ruby-version. The generated Dockerfile now rewrites the file to the concrete installed version right after the source is copied in.
Discourse still doesn’t run on Railyard as I write this. Its Gemfile declares puma in the test group, so a production bundle has no application server, and its supported deployment is its own discourse_docker image. What the afternoon produced was five fixes that apply to Ruby apps built through Nixpacks in general.
The command that ran bundle install
The second story is about the buildpack launcher. It took two attempts.
A buildpack image has ENTRYPOINT ["/cnb/process/web"]. Started with no arguments, it runs the app’s declared web process. Railyard lets you set your own command per process (bundle exec sidekiq for a worker, a different Puma bind for web), and the agent passed it the way you pass a command to any image, by appending it after the image name. Running the image by hand on a test server showed what that did:
$ docker run --env-file app.env <image> bundle exec puma -b tcp://0.0.0.0:3000
Bundle complete! 169 Gemfile dependencies, 221 gems now installed.
$ docker run --env-file app.env --entrypoint launcher <image> bundle exec puma -b tcp://0.0.0.0:3000
bundler: command not found: puma
With that entrypoint, trailing arguments to docker run become extra arguments to the declared process, and the command I asked for never runs. In this image what ran was bundle install. It succeeded, exited 0, and the container stopped. From Railyard’s side the new container “did not stay up” with a clean exit code and nothing in its log, because the agent read the log two seconds after start and the container hadn’t written anything yet. That was a separate bug: the agent now retries the log read a few times a second apart before giving up. A command that silently becomes a different command, succeeds, and exits is about as hard to diagnose as a deploy failure gets.
The first fix, on 11 September, was to start these containers with --entrypoint launcher, so the arguments are treated as a command. The second line of the repro shows the result: the command was finally running, and Puma’s own error came back.
Two days later apps with custom commands were still crash-looping on buildpack images, this time with bundle itself missing from the PATH. The launcher behaves differently depending on how many arguments it gets. Given a single string, it runs it through a shell that first sources the buildpacks’ profile.d scripts, which is where the Ruby buildpack sets PATH and the gem variables. Given more than one argument, it execs the first one directly and skips that setup. sh, -c, bundle exec puma is three arguments, so sh ran in an environment where the buildpack’s Ruby didn’t exist.
// Before: every buildpack command was wrapped, three argv elements.
func commandArgs(cmd string, isBuildpack bool) []string {
if isBuildpack || strings.ContainsAny(cmd, "$`|&;<>(){}*?~\n") {
return []string{"sh", "-c", cmd}
}
return strings.Fields(cmd)
}
// After: the launcher gets the whole command as one element,
// and sources the buildpack environment before running it.
func commandArgs(cmd string, isBuildpack bool) []string {
if isBuildpack {
return []string{cmd}
}
if strings.ContainsAny(cmd, "$`|&;<>(){}*?~\n") {
return []string{"sh", "-c", cmd}
}
return strings.Fields(cmd)
}
sh -c is what you pass to every other image, which is why I had written it into the buildpack path without thinking. The launcher’s one-argument path does the shell work itself, and does it with the environment the buildpacks built.
After the change I deployed a Rails app, a Laravel app and a Django app on buildpack images, each with a custom process command, and all three booted with their full buildpack environment.