Skip to content
A Bekkour
Writing

Why I'm building deploy software instead of a host

10 min read

I have deployed Rails apps two ways for most of my career. A platform did it for me (Heroku, then Render and Fly), or I did it by hand on a VPS with shell scripts, an Nginx config nobody wanted to touch, and a cron job that was supposed to back up Postgres. Neither one is the setup I want, so I’m building the one I want. It’s called Railyard.

Two ways to deploy, two bills

A platform as a service (PaaS) takes your code and gives you back a URL. You push, it builds an image, starts your processes, puts a proxy and a certificate in front of them, and attaches a database. For the first month it’s the best developer experience in the industry.

Then the app grows. You add a background worker, a second database, a bigger plan for the web process, and the monthly bill starts to look like the price of a much larger machine than the one you’re using. That gap is the platform’s margin, and it pays for something real: someone else is on call for the hardware. What you give up is visibility. You can’t see the machines, you choose from the regions they offer, and leaving means rebuilding your whole deploy story somewhere else.

A VPS is the other extreme. A server from Hetzner or DigitalOcean is cheap and fast, and you pay the provider and nobody else. You also inherit everything around the server: building images, swapping versions without dropping requests, TLS, logs, backups, the Postgres upgrade you keep putting off. Every team I’ve seen take this route ended up with one person who understood the scripts, and that person never had time to explain them.

I want the cost and control of the second option with the workflow of the first.

What Railyard is

Railyard is deploy software. You connect a cloud account at DigitalOcean, Hetzner or Vultr. Railyard creates servers in that account (or attaches ones you already have), installs a small agent on each, and from then on you get the platform workflow: point it at a git repository, and it builds, runs, routes traffic, streams logs and rolls back. The whole app runs there, including the frontend, the backend, Postgres, Redis, workers and scheduled jobs, managed by the same agent.

The money follows from the shape. Your provider bills you for the servers, directly, at their price. Railyard never resells compute, so there’s no hosting margin hidden in anything. And if you stop using Railyard, your servers and your data are still sitting in your account, running.

Where it sits

None of this is a new idea, and the neighbours are good.

Heroku, Render, Fly and Railway rent you the compute. They’re the experience I’m aiming for and the business model I’m avoiding.

Kamal, from Basecamp, deploys containers to your own servers and does it well. It’s a command-line tool you run from your laptop or CI: there’s no dashboard, databases are accessories you configure yourself, and you still need to understand Docker and its proxy to fix things when they break.

Cloud 66 and Hatchbox are the closest in shape. Both deploy onto servers you own and charge for the deploying, and both have done it for years. They prove the model works.

Coolify and Dokploy give you a web panel for your own servers, but you install, host and patch that panel yourself. When the panel’s server has a bad day, so does your ability to deploy.

What I’m building takes Heroku’s deploy engineering (build, release, run, rollback, a DATABASE_URL that just appears) and puts it in a product with as few screens and settings as I can manage, on servers the customer owns, with a control plane I run so they don’t have to.

The control plane is never in the request path

One constraint shapes most of the design, so I’m writing it down before the code grows around it.

The request path is every hop a user’s HTTP request travels before it reaches your app: DNS, a load balancer or proxy, maybe a CDN, then your process. Anything in that path can take your site down by being slow or broken. The control plane is the part that makes decisions: which app runs where, which version is current, what the environment variables are. Kubernetes draws the same line between its API server and the kubelet on each node; the half that serves traffic is usually called the data plane.

In Railyard, the control plane is a Rails app that I host, and it stays out of the request path entirely. A request for your app goes to your server, through a proxy on that server, to a container on that server. Railyard’s machines never see it. The containers run under Docker’s restart policy, so a reboot brings them back without asking anyone. If my control plane has an outage, the worst case is that you can’t deploy for a while. Your app keeps serving.

That rules out some convenient designs. I can’t route traffic through a central edge I control, I can’t have the agent ask Rails for permission before it restarts a crashed process, and anything an app needs at runtime (its config, its database, its certificates) has to live on the server it runs on. I’m fine with all of that. A deploy tool that can take down the apps it deployed is a worse product than one that occasionally can’t deploy.

The first two days

I started on a Sunday afternoon with a Docker Compose file and three services. Postgres is the only datastore. The Rails 8.1 app is the control plane: it owns the database, the dashboard, and every decision about what runs where. The Go program is the agent, the thing that will live on each server and run git, docker build and docker run. On day zero it did nothing but log a heartbeat and prove it could run docker ps against the host’s Docker daemon through a mounted socket.

On Monday the agent learned to talk. On boot it calls Register, gets a UUID back, and saves it to disk so it keeps the same identity across restarts. Every ten seconds after that it sends a Heartbeat with CPU, memory and disk usage, and the dashboard shows the server as online with live numbers.

The contract between the two programs is one protobuf file. Both sides generate their code from it, so if I rename a field, the build breaks on both sides instead of the two programs drifting apart. The deploy part looked like this by Monday evening:

service Agent {
  rpc Deploy(DeployRequest) returns (stream DeployEvent);
}

message DeployRequest {
  int64 application_id = 1;
  string git_url = 2;
  string git_ref = 3;
  int32 host_port = 4;
  int32 container_port = 5;
  map<string, string> env_vars = 6;
}

message DeployEvent {
  enum Kind {
    KIND_UNSPECIFIED = 0;
    KIND_STDOUT = 1;
    KIND_STDERR = 2;
    KIND_COMPLETED = 3;
    KIND_SYSTEM = 4;
  }
  Kind kind = 1;
  string line = 2;
  int32 exit_code = 3;
  google.protobuf.Timestamp emitted_at = 4;
}

Deploy is server-streaming: one request goes in, many events come out. Every line the build prints becomes an event, and the stream ends with exactly one KIND_COMPLETED carrying the exit code. KIND_SYSTEM came a few hours after the other kinds, for the agent’s own messages (“Cloning…”, “Building image…”) so the dashboard can colour them apart from the build output.

The part of the agent I’m most pleased with is the loop that streams a subprocess. A child process has two output pipes, stdout and stderr, and I want both forwarded as they’re written. The catch is that a gRPC stream’s Send is not safe to call from two goroutines at once; two concurrent writes can interleave on the wire and corrupt the framing. So each pipe gets a goroutine that scans lines into one channel, and a single loop drains that channel onto the stream:

events := make(chan *pb.DeployEvent, 64)
var wg sync.WaitGroup
wg.Add(2)
go scanLines(stdout, pb.DeployEvent_KIND_STDOUT, events, &wg)
go scanLines(stderr, pb.DeployEvent_KIND_STDERR, events, &wg)
go func() { wg.Wait(); close(events) }()

for ev := range events {
	if err := stream.Send(ev); err != nil {
		_ = cmd.Process.Kill()
		return 0, fmt.Errorf("stream send: %w", err)
	}
}

The command is created with exec.CommandContext(stream.Context(), ...), which ties the child process to the gRPC call. If Rails hangs up or the network drops, the context is cancelled and Go kills the process, so a stuck git clone doesn’t outlive the deploy that started it. Every step (clone, resolve the commit, build) runs through that one helper, which returns the exit code, and the first non-zero code ends the pipeline.

The first build of a public repo was docker/welcome-to-docker from GitHub. It took 39.2 seconds and streamed 115 lines back to Rails, which saved each one as a row. The 449 MB image landed on the host tagged docker-welcome-to-docker: plus the first twelve characters of the commit it was built from. The deploy page refreshed every two seconds to show the new lines; Turbo Streams can come later.

By the evening the agent also ran the container, passed environment variables, mapped a port, enabled BuildKit so modern Dockerfiles build, and checked the container was still up two seconds after starting it. I deployed a small Sinatra app and a fresh Rails 8 app with a Posts scaffold, so I could create and edit records through a deployed app.

The last change of the day turned one agent into three. A platform has many servers, and I wanted Rails to route each deploy to the right one from the start, so each agent reports its own gRPC address when it registers and Rails dials that address for apps assigned to that server. In Compose, the three agents share one definition through a YAML anchor:

agent: &agent_base
  build: ./agent
  environment:
    CONTROL_PLANE_URL: ${CONTROL_PLANE_URL}
    AGENT_SECRET: ${AGENT_SECRET}
    LISTEN_ADDR: agent:50052
  volumes:
    - agent-identity:/var/lib/agent
    - /var/run/docker.sock:/var/run/docker.sock

agent-2:
  <<: *agent_base
  environment:
    CONTROL_PLANE_URL: ${CONTROL_PLANE_URL}
    AGENT_SECRET: ${AGENT_SECRET}
    LISTEN_ADDR: agent-2:50052
  volumes:
    - agent-2-identity:/var/lib/agent
    - /var/run/docker.sock:/var/run/docker.sock

Each agent gets its own identity volume, because without one they’d all read the same UUID and overwrite each other’s server row. I deployed the Sinatra app to the second agent and the Rails app to the third at the same time, and both came up on their own ports.

What the toy cheats on

It’s a toy, and I know where. All three agents share the host’s Docker daemon through the mounted socket, which is root-equivalent access and means they aren’t separate machines at all. The deploy runs in a Ruby thread inside the Puma process, so restarting Rails mid-deploy loses it. The gRPC channels have no TLS. There’s no proxy, no hostname, no certificate, no database for the apps, and a repo without a Dockerfile gets skipped with a note in the log.

Each of those has a known fix, and none of them change the shape: a control plane that decides, an agent that does, and a typed contract between them.

What will be hard

The happy path took me a day. Getting a container from a repo onto a server is the easy part. A platform earns its keep on the repos it didn’t write: the app with no Dockerfile, the Ruby version written as a range, the migration that has to finish before new code takes traffic, the build that needs more memory than a small server has. Heroku has been collecting those cases for over fifteen years, and I’m starting from zero.

The other hard part is restraint. Every PaaS I’ve used has hundreds of settings, and I want Railyard to have as few as I can get away with, with defaults that are right for a Rails or Django app with a database, because that’s what most people I know run.

I’ll write here as I build it. The next post covers what happens between git push and a working URL, the five jobs every platform does, because I had to understand each one before I could build it into the agent.

Al Mokhtar Bekkour

Senior Rails & Go engineer in Quebec. I'm building Railyard, deploy software that runs your whole app on servers you own, and writing here about how it works. Open to work.

← All writing