/ tag
Railyard
14 posts on this topic.
-
A deadlock Postgres can't see
A Rails migration waited forever on a lock held by its own transaction, through a second connection pool. Why Postgres' deadlock detector can't catch that, how to spot it in pg_stat_activity, and why the logs were empty.
-
Why parallel deploys waited on each other
Several Railyard deploys at once queued behind each other and froze the dashboard. The shared resources one kind of work could fill, how I split them, and the two bugs the fixes created.
-
One door for every deploy
Railyard had nineteen places that could start a deploy, each with its own copy of the rules. How I moved every one of them behind a single model method, and the table-driven test that keeps them there.
-
Why my Rails dashboard took two seconds to load
A performance postmortem on Railyard's Rails and Hotwire control plane: an image build on the live box, a hidden tab that streamed logs into Solid Cable, and a page that rendered fourteen hidden tabs to show one.
-
Deploys that survive a dropped stream or a control-plane restart
How I moved each Railyard build into an agent-side session that outlives the gRPC stream, so a ping timeout or a control-plane release no longer kills a customer's deploy, and the Go context bug that came with it.
-
Point-in-time recovery for Postgres on a server you own
How I added point-in-time recovery to Railyard's managed Postgres: WAL archiving, a daily base backup, and a restore that rebuilds the past in a throwaway Postgres so the live database is never touched until you say so.
-
Moving Railyard's control plane from Fly.io to one EC2 box
Why the Railyard control plane needed four Fly machines and a dedicated IPv4, how it runs now on one Graviton EC2 instance with RDS and Caddy, and the cache-table bug that returned 500 on every request after the cutover.
-
The deploy that hung for an hour
Two bugs in Railyard's Go agent about the end of a process's life: a build timeout that killed one process out of a tree, and a stale Puma pidfile that survived every restart.
-
Hardening a deploy platform against apps I didn't write
How I harden Railyard against open-source Rails apps: deploy each one unmodified, group the failures into classes, fix the class on its second occurrence, and only call something fixed when a fresh build reproduces it.
-
Four ways to build someone else's repo
The options a deploy platform has for turning a repository it didn't write into a running container, what each costs, and two bugs in Nixpacks and the buildpack launcher I hit deploying Discourse.
-
A cloud firewall broke every deploy, and the fix was to stop dialing Postgres
Railyard created each app's Postgres database over the server's public port 5432, so a cloud firewall made every deploy time out. I moved the SQL into the agent's existing gRPC channel and put app-to-database traffic on the Docker network.
-
Why Railyard is a Rails app and a Go agent joined by one proto file
Railyard splits into a Rails control plane that decides and a Go agent that does the work on each server, joined by a gRPC contract. How the contract is laid out, how build logs stream, and what the split costs.
-
What a platform does between git push and a working URL
The five jobs every PaaS does after you push: build an image, run a release step, start your processes, route traffic to them, and roll back when it goes wrong. Explained from first principles, with Railyard's two-week-old pipeline as the running example.
-
Why I'm building deploy software instead of a host
Railyard deploys your whole app onto servers in your own DigitalOcean, Hetzner or Vultr account. Why I chose that shape, where it sits next to Heroku, Kamal and Coolify, and what the first two days of code looked like.