Why parallel deploys waited on each other
11 min read
One deploy at a time, Railyard was fine. With several at once, deploys sat in “Waiting for a build slot” while another app built, the dashboard stopped responding for seconds at a time, and once in a while a deploy that had done nothing wrong was marked failed.
Every one of those traced back to the same shape of problem: a shared resource that one kind of work could fill, leaving nothing for everything else.
One build slot per host
The Go agent on each server runs the builds. In front of the build step sits a gate, a buffered channel used as a semaphore: a build sends into the channel to take a slot and receives from it to give the slot back. The channel’s capacity is the number of builds that can run at once.
The capacity was 1. That was a safe default when the agent was new and servers were small, but it meant every other deploy on the same server, or on a shared build server, queued behind whichever one arrived first. A large server built one app at a time, exactly like the smallest one.
I changed the agent to size its slots from the host it runs on, from memory and core count with a ceiling, and kept a setting that overrides it. A small server still builds one thing at a time on purpose. Two image builds on a box with little memory both run out of it, and two failed builds are worse than one that waited.
The gate had a second problem that only showed up when the queue got long. A separate sweep job fails any running deploy with no log output for 30 minutes, because silence for that long usually means the worker running it died. A build waiting for a slot produces no output. With enough deploys queued, the reaper killed deploys that were waiting exactly as designed. The wait loop now has a five-minute ticker next to the channel send, and every tick emits “Still waiting for a build slot on this host…” as a log line. That line counts as output for the reaper, and the person watching the deploy sees why nothing is happening.
One job pool for everything
On the Rails side, every background job ran in one Solid Queue worker with eight threads and queues: "*", so any thread took any job. The deploy job holds its thread for the whole build, often several minutes. Database backups, server moves, app commands and scheduled tasks do the same. Run a few of those together and all eight threads are busy with long work. The next deploy waits. So does the cancel button, which is also a job, and so do the per-minute sweeps that check server health and reconcile deploys.
The pattern for this is called a bulkhead, after the walls that divide a ship’s hull so one flooded compartment can’t sink the ship. Each kind of work gets its own pool, so filling one can’t starve another. The worker config became three pools:
# config/queue.yml
default: &default
dispatchers:
- polling_interval: 1
batch_size: 500
workers:
- queues: deploys # one thread per running deploy
threads: 16
polling_interval: 0.1
- queues: operations # backups, restores, moves, imports, commands, restarts
threads: 8
polling_interval: 0.2
- queues: default # sweeps, notifications, housekeeping
threads: 8
polling_interval: 0.1
No pool takes * anymore. A test fails if any queue has no pool, if a * pool comes back, or if a job class doesn’t declare its queue explicitly, so a new job has to choose its pool on purpose. That last check came a day later, after I noticed how easy it would be for a long-running job to slip onto the housekeeping queue by default.
The worker runs Solid Queue in async mode, so all three pools are threads inside one process instead of three forked processes. On a small 2-vCPU box that matters: after the deploy the whole worker sat at 206 MB.
The per-minute sweeps had a smaller version of the same problem. A recurring job that sometimes runs longer than a minute gets enqueued again while the previous run is still going, and the copies pile up and take threads from everything else. Solid Queue’s concurrency controls handle it: every sweep now declares limits_concurrency with on_conflict: :discard, so a new tick that finds the previous one still running is dropped instead of queued. Dropping a tick is safe for a sweep, because the next one runs a minute later and sees the same state.
A database write per log line
Build output streams from the agent over gRPC, one event per line. For each line the Rails side did two writes: an INSERT into the deploy logs table, and an update to the deploy’s step_progress JSON column to save the stream’s resume cursor. The cursor is the sequence number of the last line received, and it lets a dropped stream reattach without losing or repeating output; I wrote about that design in deploys that survive a restart.
Updating a JSON column in Postgres rewrites the whole value. A noisy npm install produces thousands of lines. Several builds at once meant a constant stream of tiny transactions against a micro RDS instance, two per line.
Batching is the usual answer when the producer is fast and each write has a fixed cost. Lines go into an array, and the array is written with one insert_all when it reaches 200 lines or 0.3 seconds have passed, whichever comes first:
LOG_BATCH_LINES = 200
LOG_BATCH_SECONDS = 0.3
def flush_due?(event, state)
!event.output_line? ||
state[:lines].size >= LOG_BATCH_LINES ||
monotonic_now - state[:flushed_at] >= LOG_BATCH_SECONDS
end
def flush_logs(step, request, state)
state[:flushed_at] = monotonic_now
DeployLog.insert_all(state[:lines]) if state[:lines].any?
state[:lines] = []
return if state[:seq] == state[:saved_seq]
save_step_progress(step, "session" => request.session_id, "seq" => state[:seq])
state[:saved_seq] = state[:seq]
end
def monotonic_now = Process.clock_gettime(Process::CLOCK_MONOTONIC)
The order inside flush_logs is the part that matters. The cursor is saved after the lines it covers are stored, so the database never claims a line was received that it doesn’t hold. If the process crashes between flushes, the cursor points at the last stored line and the agent replays from there. The stream’s error handlers and its ensure block flush too, so a broken connection keeps the partial batch. Any non-output event (a system step, the exit code) flushes first, so ordering in the log view stays correct. The test sends 300 lines and expects fewer than 20 writes.
insert_all skips Active Record callbacks and validations. That was safe here only because the one callback on the log model fires for system lines, and system lines still go through create!.
A canvas refreshing ten times a second
That callback was the next problem. The Canvas is Railyard’s diagram view of a team’s apps, servers and databases, and it updates live through Turbo’s page refresh: the server broadcasts “refresh”, every open Canvas refetches the page, and Turbo morphs in the difference. Every system line of a deploy (“Cloning”, “Building”, “Starting web”) triggered one of those broadcasts for the whole team.
With five deploys running, that was about ten full page renders per second for each person with the Canvas open, each one a real request on a web process with five threads. A couple of viewers could keep Puma busy redrawing a diagram that had barely changed, and every other page waited behind those renders. That was the freeze.
The fix is a throttle that uses the cache as the clock:
class DeployLog < ApplicationRecord
after_create_commit -> { deploy.application.team.refresh_canvas_soon },
if: -> { stream == "system" }
end
class Team < ApplicationRecord
def refresh_canvas = broadcast_refresh_later_to(self, :canvas)
def refresh_canvas_soon
refresh_canvas if Rails.cache.write([ self, :canvas_refresh ], true,
unless_exist: true, expires_in: 2.seconds)
end
end
unless_exist: true makes the write succeed only when the key is absent, so the first step in any two-second window refreshes and the rest are skipped. The cache is shared across processes, so the throttle holds no matter which worker thread logs the step. Status changes, a deploy succeeding or failing, still refresh immediately, because those are what people are waiting to see. The steps in between can arrive two seconds late without anyone noticing.
The same family of problem, real-time updates fanning out far more than anyone needed, also showed up as Solid Cable traffic on the app page. That one is its own story, in why my dashboard took two seconds to load.
Bugs I added
All of that shipped on the morning of September 30 with the suite green. That evening I made two follow-up fixes.
The first was a connection budget. Each Rails process keeps a database connection pool, and the pool size is the most connections that process will ever open. Postgres has a hard limit, max_connections, for the whole server. The budget is multiplication: the sum over every process of its pool size has to stay under the server’s limit, with room left for migrations and a console. This box runs several processes against one micro RDS instance (web, gRPC, the log forwarder, the job worker, and during a release a second web container), and that instance had already run out of connection slots once, at 78.
When I added the job pools, I sized the pool in database.yml to cover every job thread. database.yml is read by every process, so web, gRPC and the log forwarder each got the job worker’s pool too:
# pool ceiling per process, before and after the follow-up fix
job_threads = 16 + 8 + 8 # deploys + operations + default
jobs_pool = job_threads + 5 # => 37
other_pool = [ 5, 8 ].max + 3 # puma threads vs old job threads => 11
before = { web: 37, grpc: 37, log_forwarder: 37, jobs: 37, web_next: 37 }
after = { web: 11, grpc: 11, log_forwarder: 11, jobs: jobs_pool, web_next: 11 }
before.values.sum # => 185, against a limit already hit at 78
after.values.sum # => 81 during a release
after.except(:web_next).values.sum # => 70 the rest of the time
Nothing had failed yet, because connections open lazily and idle ones are reaped after 30 seconds, but 185 on paper against a limit near 80 only needed the right burst of load. The fix moves the number to the process that needs it: the job worker’s boot script sets its own pool to its thread total plus a margin, and every other process falls back to the old pool of 11.
The arithmetic above still has one multiplier missing. Rails keeps a separate pool per database config, and this app has four of them (primary, queue, cache and cable), all pointing at the same Postgres in production. A process that touches all four can hold connections in four pools. In practice most of those stay idle and get reaped, but a complete budget includes them, and mine doesn’t yet. I wrote about what multiple pools in one process can do to a migration in a deadlock Postgres can’t see.
The second follow-up was in the script that deploys Railyard itself. I started two deploys from two terminals at nearly the same time, and mine failed twice, first with gzip: invalid compressed data and then with archive/tar: invalid tar header. Both runs had uploaded the image to the same fixed path in /tmp on the box, and each overwrote the other’s file mid-transfer. Production stayed up, because loading the image failed before any restart. There was a quieter race too: both runs tagged their local build latest, so a deploy could have shipped the other run’s code with no error at all.
A fixed path in /tmp is shared state, and a script that can run twice is a concurrent program. I fixed it the way I’d fix any other one, with a lock and names that can’t collide:
DEPLOY_ID="$(git rev-parse --short HEAD)-$$"
LOCK=/tmp/deploy.lock
until ssh "$HOST" "find $LOCK -maxdepth 0 -mmin +45 -exec rm -rf {} +;
mkdir $LOCK 2>/dev/null && echo '$(whoami) $DEPLOY_ID' > $LOCK/owner"; do
[ "$waited" -ge 1800 ] && die "another deploy still holds the lock"
sleep 15; waited=$((waited + 15))
done
trap 'ssh "$HOST" "grep -q \" $DEPLOY_ID\" $LOCK/owner && rm -rf $LOCK"' EXIT
REMOTE_TGZ="/tmp/control-plane-$DEPLOY_ID.tar.gz"
ssh "$HOST" "gunzip -t $REMOTE_TGZ && gunzip -c $REMOTE_TGZ | docker load -q \
&& docker tag control-plane:$DEPLOY_ID control-plane:latest"
mkdir is atomic, so exactly one run can create the lock directory, and the owner file records who holds it. A second deploy prints the holder and waits up to 30 minutes. A lock older than 45 minutes belongs to a crashed deploy and is taken over. The exit trap releases the lock only if this run owns it. Each run builds and uploads its own tag and file, named after the commit and the process ID, checks the archive with gunzip -t before docker load, and only then retags latest. I tested it with a fake ssh: two runs at once, a timeout, a stale lock, and release on exit.
What comes next
The fixes are in production. The worker came up with its three pools in one process and no errors after the recurring ticks, and the agents picked up the new build-slot logic on their next self-update. Every cause above came from reading code and tracing where work queued, because production rarely ran enough deploys at once to show them under load. The next step is the live test: three or four real deploys at once on a throwaway server, watching them build in parallel to find what breaks after this.
Until that test runs, parallel deploys are fixed in code and verified one deploy at a time.