Skip to content
A Bekkour
Writing

Why Railyard is a Rails app and a Go agent joined by one proto file

10 min read

Railyard has to run commands on servers it doesn’t live on, while the product (users, apps, deploy history) lives in one database somewhere else. I solved that by building it as two programs from the first commit: a Rails app that decides, and a Go agent on each server that does. They share exactly one file, a protobuf definition, and they talk gRPC.

The first post described the split in one line. Five weeks in, this is how it works, why I drew the line where I did, and what it already costs.

Deciding and doing

The line between the programs follows one question: does this code need to know about the product, or does it need to touch the machine?

Users, applications, deploy history, which app belongs on which server, whether a deploy succeeded: that’s product state. It lives in Postgres, it changes with every feature, and all of it belongs in the control plane.

Running git clone, running docker build, reading CPU and disk usage, killing a stuck build: that has to happen on the server, because that’s where Docker and the disk are. The agent does those things and keeps almost no state. It stores one thing on disk, a UUID the control plane hands it on first registration, so it keeps the same identity across restarts.

Each side fails in a contained way. If the control plane goes down, the agent stops getting orders, but every container it started keeps running. If an agent dies, one server stops heartbeating and the dashboard shows it going stale. Neither failure takes a customer’s app down.

What gRPC gives you

gRPC is a remote procedure call framework. You describe services and messages in a .proto file, a compiler generates client and server code for each language, and calls travel as compact binary protobuf messages over HTTP/2. A method can take one of four shapes: unary (one request, one response), server-streaming (one request, many responses), client-streaming (many requests, one response), and bidirectional (both sides stream at once).

The point for Railyard is the generated code. protoc runs twice from one script, once with the Go plugins into the agent’s tree and once with the Ruby plugin into the Rails app. If I rename a field, both sides break at build time or boot instead of drifting apart silently.

This is the whole service section today:

// Hosted by Rails on :50051. Agents dial it.
service ControlPlane {
  rpc Register(RegisterRequest) returns (RegisterResponse);
  rpc Heartbeat(HeartbeatRequest) returns (HeartbeatResponse);
}

// Hosted by the Go agent on :50052. Rails dials it.
service Agent {
  rpc Deploy(DeployRequest) returns (stream DeployEvent);
}

message DeployEvent {
  enum Kind {
    KIND_UNSPECIFIED = 0;
    KIND_STDOUT      = 1;
    KIND_STDERR      = 2;
    KIND_COMPLETED   = 3;
    KIND_SYSTEM      = 4;
  }
  Kind kind = 1;
  string line = 2;
  int32 exit_code = 3;
  google.protobuf.Timestamp emitted_at = 4;
}

A service belongs to whoever hosts it

The first version of the file had all three RPCs on one service called ControlPlane. It compiled and it looked tidy. When I sat down to implement Deploy, the generated code wanted Rails to host it, which made no sense: Rails doesn’t run builds, the agent does. I split it onto its own Agent service the same afternoon.

The rule I took from that: in gRPC a service is something a process serves, so it belongs to whoever listens. Register and Heartbeat are the agent telling Rails something, so Rails hosts them on port 50051. Deploy is Rails telling the agent to act, so the agent hosts it on 50052. Traffic goes both ways, so there are two services, two listeners and two ports.

That makes the agent a client and a server inside one process. main dials the control plane, registers, and runs a heartbeat every ten seconds with CPU, memory and disk percentages. Just before that loop, a single go runAgentServer(secret) starts the other role in a goroutine that listens on 50052 and serves Deploy. The main goroutine keeps heartbeating until SIGTERM. Both directions check the same bearer token, carried as gRPC metadata.

Streaming a build log

Deploy is server-streaming: one DeployRequest goes in, many DeployEvents come back. Each event has a kind, a line, a timestamp, and on the last one an exit code. That shape fits a build exactly. Rails asks once, then watches.

The agent’s job inside that call is to run a command and turn its output into events while it runs. A process has two output pipes, stdout and stderr, and both produce lines at the same time. So I read each pipe in its own goroutine. The catch is that gRPC-Go does not allow two goroutines to call Send on the same stream at once; its documentation says so directly, and violating it corrupts the stream. So the readers never touch the stream. They push events into a channel, and one loop drains the channel and is the only thing that sends:

cmd := exec.CommandContext(stream.Context(), name, args...)
stdout, _ := cmd.StdoutPipe()
stderr, _ := cmd.StderrPipe()
cmd.Start()

events := make(chan *pb.DeployEvent, 64)
var wg sync.WaitGroup
wg.Add(2)
go scanLines(stdout, pb.DeployEvent_KIND_STDOUT, events, &wg)
go scanLines(stderr, pb.DeployEvent_KIND_STDERR, events, &wg)
go func() { wg.Wait(); close(events) }()

for ev := range events {
	if err := stream.Send(ev); err != nil {
		cmd.Process.Kill()
		return 0, fmt.Errorf("stream send: %w", err)
	}
}
return exitCode(cmd.Wait())

Three things fall out of this shape that I didn’t have to write separately.

Ordering: the channel puts interleaved stdout and stderr lines into one sequence, so the log Rails stores reads the way the build printed it.

Backpressure: if Rails is slow to read, Send blocks, the channel fills to 64 events, the scanner goroutines block on the channel, the OS pipe buffer fills, and the build process itself blocks on its next write. A slow consumer slows the build down instead of growing memory on the server.

Cancellation: the command is tied to the call’s context with exec.CommandContext. If Rails gives up on the call or the network drops, the context is cancelled and Go kills the process. A git clone hanging on a dead remote doesn’t outlive the deploy that started it.

The system messages (“Cloning…”, “Building docker image…”, “Container is up”) are sent from the same handler goroutine, between steps, never while a drain loop is running. So there is still one sender at any moment. The final KIND_COMPLETED event carries the exit code of the step that ended the pipeline.

Consuming the stream in Ruby

On the Rails side, the generated stub returns an Enumerator for a server-streaming call, so reading the stream is a plain each. A deploy job dials the agent of the server that app lives on and writes each event as a log row:

stub.deploy(request, metadata: auth).each do |event|
  case event.kind
  when :KIND_STDOUT, :KIND_STDERR, :KIND_SYSTEM
    @deploy.deploy_logs.create!(
      stream:     event.kind.to_s.delete_prefix("KIND_").downcase,
      line:       event.line,
      emitted_at: timestamp_to_time(event.emitted_at),
    )
  when :KIND_COMPLETED
    exit_code = event.exit_code
    succeeded = exit_code.zero?
  end
end

@deploy.update!(status: succeeded ? :succeeded : :failed, finished_at: Time.current)

Each iteration blocks until the next event arrives. If the agent crashes mid-build, the iteration raises a gRPC error, the job’s rescue writes the error into the log, and the deploy is marked failed.

KIND_SYSTEM wasn’t in the first version. I added it about twenty-five minutes after the other three, as enum value 4, because the log had no way to say what the agent itself was doing between commands. That’s the general protobuf rule for changing a contract that’s already in use: add fields and enum values with new numbers, never renumber or reuse old ones. An old reader that sees value 4 falls into its default case instead of crashing.

Why Go for the agent

The agent has to run on a server I don’t control, next to someone’s production app. I want installing it to mean copying one file. Go compiles to a single binary, and with CGO_ENABLED=0 it’s fully static: no Ruby, no gems, no shared libraries to match against the distro. Today that binary runs inside three Compose containers that stand in for servers. Making it static now means moving it onto a real Ubuntu box later is a file copy and a systemd unit.

The agent’s work is also mostly concurrency: pipes, timeouts, cancellation, a heartbeat running beside a server. Goroutines, channels and context cover all of it with the standard library, as the loop above shows.

Why Rails for the control plane

The control plane is a product: forms, a dashboard, a deploy log that updates while you watch, and later users, teams and permissions. That’s what Rails is for, and I move faster in it than in anything else. One monolith means a new feature is a migration, a model, a controller and a view in one repo with one test suite.

Rails never touches a server directly. It doesn’t SSH anywhere and doesn’t run Docker. Every physical action goes through the agent, which keeps the code that can damage a server in one place.

What the split costs

Generated code has a cost. The Ruby and Go stubs are checked in, and regenerating them needs protoc plus a Go plugin and a Ruby plugin installed locally. Editing the proto and forgetting to regenerate leaves the two sides disagreeing until the next run of the script.

gRPC in Ruby costs more. The grpc gem is a large native extension. Zeitwerk won’t autoload the generated files because their _pb.rb file names don’t match the constant names inside them, so an initializer adds the directory to the load path and requires them by hand. The gRPC server runs in a background thread inside the Puma process, which is fine on a laptop and not how I’ll run it in production. And binding the port from short-lived commands like rails db:migrate segfaulted on shutdown, so the listener only starts when Rails is running as a server.

REST would have done half of this job. Register and Heartbeat are request and response; a JSON endpoint in Rails and an HTTP client in Go would have worked, and I could debug them with curl. Deploy decided it. I wanted a live stream of typed events, an explicit end carrying an exit code, and cancellation that reaches a subprocess on another machine. Over plain HTTP I’d have built most of that myself out of chunked responses or server-sent events and a home-made framing convention. gRPC puts it in the contract, and once one RPC needs gRPC it’s simpler to run every call on the same channel than to keep two protocols.

The direction problem, and the fix I’ve designed

There’s a weakness in this layout that I can see now. Deploy works because Rails opens a connection to the agent on port 50052. In Compose that’s free: every container can reach every other one. A real server may sit behind NAT, or a cloud firewall that drops inbound traffic, and asking customers to open a port to the internet for my agent is a bad first step.

The design I’ve settled on is to stop dialing the agent. The agent already dials out to the control plane for heartbeats, and outbound connections get through almost any network. So the agent will open one long-lived bidirectional stream to Rails at boot and hold it. Rails sends commands down that stream; the agent sends events back up it. The connection direction flips, and the server never needs an inbound port. This is the shape I’ve sketched:

service ControlPlane {
  rpc Register(RegisterRequest) returns (RegisterResponse);
  rpc Heartbeat(HeartbeatRequest) returns (HeartbeatResponse);
  rpc Connect(stream AgentEnvelope) returns (stream ControlEnvelope);
}

message AgentEnvelope {
  oneof msg {
    ConnectHello hello = 1;  // server UUID, sent first
    Ping ping = 2;
  }
}

message ControlEnvelope {
  oneof msg {
    ConnectAck ack = 1;
    Pong pong = 2;
  }
}

The oneof envelopes are the important part. One stream has to carry many kinds of messages, so each message on it is a wrapper with exactly one payload set, and new commands become new fields in the envelope with new numbers. The first step only proves the stream can open, identify the server and stay alive with pings. Once that holds on real networks, Deploy and everything after it can move onto the stream one RPC at a time while the current dial-in path keeps working.

The service-ownership rule still holds under this design, because Connect is served by Rails and dialed by the agent. What changes is that every byte of control traffic will ride a connection the server opened itself.

Al Mokhtar Bekkour

Senior Rails & Go engineer in Quebec. I'm building Railyard, deploy software that runs your whole app on servers you own, and writing here about how it works. Open to work.

← All writing