whoami

Lucas Jenß

cat /etc/motd

The Coding Journal ツ — Notes taken on an epic coding journey. Technical solutions, debugging notes, and practical guides from the trenches of software development.

Close-up of white dominoes with black dots standing on green felt surface, shallow depth of field, focused mood

ls -la ~/languages/

total 8
drwxr-xr-x
▶ PHP
▶ Ruby
▶ Scala
▶ C#
▶ JavaScript
▶ Objective-C
▶ Shell Scripting

ls -la ~/toolchain/

total 7
drwxr-xr-x
▶ Typo3
▶ Akka
▶ Capistrano
▶ Git
▶ MAMP
▶ Adobe Illustrator
▶ NSTrackingArea (Cocoa)

uname -a

platforms
drwxr-xr-x
▶ Mac OS X
▶ Unix

istrano Recipes for Syncing Uploaded Files Across Servers

When you run a Ruby web application across multiple boxes, keeping user uploads consistent between hosts becomes an operational headache. Capistrano handles code release beautifully, but anything written to disk at runtime — profile photos, generated PDFs, CSV exports — lives outside the repository and can drift between servers. A stray file on one node that is missing on another often surfaces as a confusing 404 for users in Melbourne, Sydney, or Brisbane.

This walkthrough shows how to assemble a Capistrano task that propagates uploaded files across every server in your fleet. The approach leans on standard Unix utilities already present on most Linux distributions, so it does not require extra daemons. We will walk through the recipe structure, contrast a few transport options, wire the sync into the deploy flow, and round out with a checklist you can apply on your own cluster.

Understanding the shared folder model in Capistrano

Capistrano's standard layout places deploys under releases/ and a persistent shared/ directory. Anything that must survive between releases — log files, configuration, and user uploads — belongs in shared/. The convention is to symlink from current/ into the matching shared/ path so application code reads from a stable location regardless of which release is live.

For uploads specifically, a common pattern looks like shared/public/uploads. When a user submits an avatar through a Rails controller, Carrierwave or Paperclip writes the file under that path, and the running app serves it directly. On a single server this is trivial, but on a multi-host setup you suddenly need the same file to exist on every web node, otherwise load balancers send a request to a machine that does not have it.

The recipe we will build simply enumerates the contents of shared/public/uploads on a designated primary server and replicates them to the rest. The primary role remains a logical choice for many Australian SaaS shops running their stack in the AWS Sydney region, since latency between an admin uploading a marketing asset and that asset being visible to readers elsewhere stays well under fifty milliseconds across the ap-southeast-2 availability zone.

Building a basic file sync recipe from scratch

A Capistrano recipe is a plain Ruby file that runs inside the same DSL the deployment tasks use. You can drop it under config/deploy/ and load it through Capfile. The skeleton below captures the goal: locate files on one server, push them to the rest, and skip anything that has not changed since the previous run.

The core task uses on(roles(:web)) to scope commands to web nodes and either upload! or execute with streaming to move the bytes. Rsync remains the most predictable tool for this job because it handles incremental transfers, partial files, and SSH authentication in one binary that is already installed by default on Ubuntu and Amazon Linux instances.

namespace :uploads do
  desc "Sync uploaded files from the primary to all web servers"
  task :sync do
    primary = roles(:web).select { |h| h.properties.primary }[0]
    receivers = ->(cmd) { on(roles(:web).reject { |h| h.properties.primary }) { |h| execute cmd } }

    files = `ssh #{primary.hostname} "find #{shared_path}/public/uploads -type f"`.split("\n")
    files.each do |relpath|
      cmd = "rsync -az #{primary.hostname}:#{relpath} #{File.dirname(relpath)}/"
      receivers.call(cmd)
    end
  end
end

The primary flag on a host tells Capistrano which node owns the canonical copy. In practice you might mark your first Sydney host as primary and let the others in ap-southeast-2 receive the update. The task above is intentionally simple — production deployments usually need timestamps, excludes for cache files, and a lock file to prevent concurrent runs that could collide on a busy Friday afternoon.

Comparing sync tools for production reliability

The recipe can use any number of transports under the hood. Choosing one comes down to file count, average file size, and how often uploads occur. A small marketing site that gets a few hero images per week has very different needs from a media platform processing hundreds of gigabytes nightly from contributors across Perth, Adelaide, and Hobart.

Tool Best for Direction Requires daemon Handles deletes
rsync over ssh One-off batch sync, small to medium sets One to many No Optional with --delete
unison Two-way mirror, occasional conflicts Bi-directional No Yes
lsyncd Continuous, low-latency mirroring One to many Yes (inotify) Yes
S3 with lifecycle Stateless storage, large objects Indirect via bucket No Native versioning
scp loop Trivial transfers, debugging One to one No No

For most Australian teams running a handful of EC2 instances out of ap-southeast-2, plain rsync over SSH strikes a healthy balance. There is nothing extra to install beyond the OpenSSH client, the failure modes are well understood, and bandwidth between hosts in the same region is essentially free. Teams running at greater scale often move uploads to S3 entirely, which sidesteps the sync problem altogether. While you are deciding, you might also be interested in solving SVG import issues, since uploaded vector assets frequently surface similar edge cases when published through the application.

Integrating the sync into the deploy lifecycle

Recipes are only useful if they run automatically. The right hook depends on what you are trying to achieve. before 'deploy:published' fires once the new release has been symlinked, which is a sensible moment if the sync needs to happen alongside the rollout. after 'deploy:finishing' runs later and is safer when uploads are written by the running app rather than by the deployment process itself.

Many teams prefer to keep the sync as a standalone task that operators can invoke manually. The pattern cap production uploads:sync reads naturally from memory during a Friday afternoon incident in Adelaide when someone has just pushed a config file but forgot the matching assets. Manual invocation also makes it easier to dry-run with -n so you can preview the rsync invocations before they hit production.

A subtle point: the primary host itself should be excluded from the receive side, otherwise rsync will copy files onto the same machine and waste a round trip. The receivers lambda in the earlier example already filters out hosts with properties.primary = true, but if you copy the snippet into another project, double-check that filter before the first deploy.

Practical recommendations for steady-state operations

Once the recipe is live, day-to-day reliability depends on a handful of habits. The list below covers the points I wish I had applied from day one rather than after a support ticket arrived at half past seven on a Tuesday morning, AEST.

  • Mark exactly one host with primary = true in config/deploy/production.rb and keep a comment explaining why that node was chosen.
  • Wrap the sync in a File.open('/tmp/uploads_sync.lock', File::LOCK_EX) block so two operators cannot run it simultaneously and trample each other.
  • Exclude generated thumbnails or cache files that can be reconstructed from the originals on the receiving side.
  • Add set :uploads_keep_versions, 5 if you decide to retain historical copies in shared/backups/uploads.
  • Log every rsync invocation through info so the output lands in the deploy log alongside the rest of the rollout.
  • Schedule a nightly cron run of the sync as a safety net, even though deploys normally trigger it manually.
  • Monitor disk usage on the primary so a runaway upload does not silently fill the root volume before anyone notices.

If your team eventually migrates to a containerised setup with ECS or EKS in the Sydney region, much of this complexity disappears in favour of an EFS mount or an S3 prefix. Until then, a tidy Capistrano recipe is often the simplest path forward and keeps the operational surface small.

Rolling this out on a test stack first catches the silly mistakes. Use a disposable environment, push a few sample files through your Rails console, then run cap staging uploads:sync and verify the byte counts match on every host. Once you are confident, promote the same task to production and keep the manual cap production uploads:sync command in your team's runbook for those inevitable after-hours fixes.

If this walkthrough helped tighten up your deploy pipeline, subscribe to the RSS feed or drop a comment below with the variant you ended up shipping on your own fleet. Happy shipping from a warm Hobart office, and may your uploads always arrive where they are needed.


cat ~/interests.json

KeyValue
editorTerminal-first workflow
osMac OS X / Unix
vcsGit, distributed version control
deployCapistrano, cron automation
graphicsSVG, Adobe Illustrator troubleshooting
networkingIP validation, SSH, VPN

git log --oneline --reverse

2013-10-30

Solving SVG import issues in Adobe Illustrator CS6 and CC

When importing an SVG into Illustrator, the operation fails with an unknown error [CANT]. A workaround for this Adobe-side bug.

2013

Solving NDK build issues on OS X

Troubleshooting native development kit compilation problems on Mac OS X.

2013

Programmatically adding PHP generated TypoScript to the backend configuration

Integrating dynamically generated TypoScript into Typo3 backend setups using PHP.

2013

ArgumentError: Could not parse PKey: no start line

Debugging an SSH key parsing error encountered during deployment.

2011-08-04

Validating IP-Addresses in PHP

Using PHP filter functions with flags like FILTER_FLAG_IPV4 and FILTER_FLAG_IPV6, and understanding how filter_var handles reserved IP addresses.

2011-07-09

Cocoa: Using NSTrackingArea

A short tutorial on using Cocoa's NSTrackingArea to capture mouseEntered and mouseExited events.


cat ~/contact.txt

github: github.com/x3ro
stackoverflow: x3ro
coderwall: coderwall.com/x3ro
twitter: @x3rames