whoami

Lucas Jenß

cat /etc/motd

The Coding Journal ツ — Notes taken on an epic coding journey. Technical solutions, debugging notes, and practical guides from the trenches of software development.

Close-up of white dominoes with black dots standing on green felt surface, shallow depth of field, focused mood

ls -la ~/languages/

total 8
drwxr-xr-x
▶ PHP
▶ Ruby
▶ Scala
▶ C#
▶ JavaScript
▶ Objective-C
▶ Shell Scripting

ls -la ~/toolchain/

total 7
drwxr-xr-x
▶ Typo3
▶ Akka
▶ Capistrano
▶ Git
▶ MAMP
▶ Adobe Illustrator
▶ NSTrackingArea (Cocoa)

uname -a

platforms
drwxr-xr-x
▶ Mac OS X
▶ Unix

Tracking Down a PHP Session Corruption Bug in a Load-Balanced TYPO3 Setup

The call came in on a Tuesday morning, Sydney time, while I was still working through my first coffee. A client in Melbourne running a TYPO3 site on a pair of load-balanced application servers was losing user sessions at random. Logged-in editors were getting logged out mid-edit, and the shopping cart on the public-facing part of the site was emptying itself between clicks. The traffic was split between two nodes sitting behind a HAProxy instance in a Sydney data centre, and the more we looked, the more it felt like a classic case of session state drifting between the two.

TYPO3 has long been a workhorse for Australian government departments, universities, and mid-sized publishers, and many of those deployments sit behind load balancers to weather traffic spikes during events like the Sydney New Year's Eve broadcast or a university open day in Brisbane. When sessions start corrupting, the impact is rarely limited to a single page. It cascades through the CMS, the frontend, and any custom extensions the agency has bolted on over the years.

I have written about my share of head-scratchers on The Coding Journal, but this one earned its own post because the root cause was not what we initially suspected. It looked like a frontend caching issue, then a cookie domain problem, then a database bottleneck. The actual culprit sat a level lower, in how PHP writes session files when the load balancer hands a user back to a different application node than the one that started the session.

This post walks through the investigation, the fix, and the configuration changes that keep the bug from coming back when a third or fourth node is added to the pool. If you run TYPO3 on anything more elaborate than a single VM, the pattern here is worth recognising before it bites you.

The Setup and the Symptoms

The site in question runs TYPO3 11 on two Debian 12 application servers behind a HAProxy load balancer in a Sydney-based facility. Sessions were stored using PHP's default file handler, with each node writing to its own local /var/lib/php/sessions directory. Static assets and uploads lived on a shared NFS mount, but the session directory was deliberately kept local to avoid latency and lock contention.

The first symptom was reported by an editor in Perth who noticed that saving a page in the TYPO3 backend would occasionally fail with a "Your session has expired" message. A refresh fixed it, but only briefly. The editor assumed it was a cookie clearing on his browser, a common enough complaint on Windows machines with aggressive privacy settings. Within a few days, the public-facing site started showing the same behaviour for logged-in customers in Adelaide, and the support inbox filled up.

The smoking gun appeared in the PHP error log on one of the nodes. Lines like "ps_files_cleanup_dir: opendir(/var/lib/php/sessions) failed" and "session_decode(): Session is corrupted" were appearing on both nodes, but with different request IDs in the same second. That told us the issue was not confined to a single machine. It was a side effect of how the load balancer distributed requests and how the application servers each treated the session cookie as belonging to them.

Why Load Balancers and PHP Sessions Clash

PHP's default file-based session handler is fast and simple, but it assumes a one-to-one relationship between a session ID and a directory on disk. When you put two or more application servers behind a load balancer, that assumption breaks unless you do one of three things: share the session storage between nodes, route the same session ID to the same node every time using sticky sessions, or switch to a session backend that is itself distributed.

Most Australian managed hosting providers, including the boutique outfits that specialise in TYPO3 hosting, recommend sticky sessions as the path of least resistance. HAProxy can be configured with a balance source directive that hashes the client IP and sticks a visitor to a particular backend for the life of the session. It works most of the time, but it falls over when editors log in from corporate networks where dozens of staff share a single egress IP through a proxy in Sydney or Melbourne. All of those users end up on the same node, and the load balancer quietly defeats its own purpose.

The other common shortcut is to point the session save path at a shared NFS mount or a GlusterFS volume. The temptation is real, especially when the storage layer sits in a data centre with a 10-gigabit backbone. The problem is that file locking on network filesystems is not as reliable as the PHP session handler assumes, and a slow NFS response during a busy checkout on a Saturday afternoon can produce exactly the kind of partially written session file that triggers the "session is corrupted" error.

Reproducing the Corruption Locally

Before changing anything in production, I wanted a way to reproduce the bug on my laptop. I set up a minimal TYPO3 11 instance with the introduction package, put it behind a local HAProxy listening on port 8080, and spun up two PHP-FPM pools on different Unix sockets. I pointed both pools at the same Redis instance running in a Docker container, which is the storage backend the production fix would eventually use.

I wrote a small script that hammered the site with concurrent requests, half of them routed through HAProxy and half hitting the backends directly. Within a few hundred requests, the local setup started producing the same session_decode() errors I had seen in the Sydney logs. The pattern was clear: when a request created a session on node A and the next request for the same user landed on node B before node A had finished flushing the session file, the second node read a half-written file and treated it as corrupt.

The reproduction mattered because it gave the team a controlled environment to test fixes without poking at production. It also made it easier to explain to the client why the answer was not a HAProxy tweak but a change in how TYPO3 talks to its session storage.

Switching to a Distributed Session Backend

The fix that worked was to move the session backend off the local filesystem and onto Redis. TYPO3 has a configuration option called SYS/session/Redis in the LocalConfiguration.php file, and it accepts a DSN that points at one or more Redis nodes. We ran Redis as a primary-replica pair on dedicated small instances in the same Sydney availability zone, with a Sentinel setup to handle failover.

The relevant configuration looks roughly like this:

'SYS' => [
    'session' => [
        'BE' => [
            'backend' => 'Redis',
            'Redis' => [
                'host' => 'tcp://10.0.0.10:6379',
                'port' => 6379,
            ],
        ],
        'FE' => [
            'backend' => 'Redis',
            'Redis' => [
                'host' => 'tcp://10.0.0.10:6379',
                'port' => 6379,
            ],
        ],
    ],
],

For sites that need even tighter control over session data, the same approach works with Memcached or a relational database. The point is that the session store needs to be a single source of truth that every application node can read from, regardless of which one handled the previous request. ACSC publishes general guidance on this kind of distributed-state design under the Essential Eight maturity model, and it is worth reading if you are about to roll the same pattern out across a larger estate.

Locking, TTLs, and the Bits You Forget

Moving to Redis solves the corruption problem, but it introduces a new set of concerns that are easy to overlook during a late-night deployment. The first is session locking. PHP's file handler takes an exclusive lock on the session file while a request is being processed, which prevents another request from writing to the same session at the same time. The Redis handler does the same, but the lock is held by a key in Redis, and if the request crashes without releasing it, the next visitor can be stuck waiting for the lock to expire.

Three settings that make a measurable difference in this setup:

  • A short lock acquisition timeout, around two seconds, so a crashed worker does not block the next visitor indefinitely
  • A reasonable session lifetime, set to match the application's business needs rather than PHP's default of 24 minutes
  • A separate Redis database index for sessions, so a FLUSHDB on the cache database does not nuke every logged-in user

The second concern is garbage collection. PHP's session garbage collector looks at the session files on disk and removes any older than the configured lifetime. When sessions live in Redis, the equivalent is to set an explicit TTL on each session key and let Redis handle the cleanup. The default TTL when you configure TYPO3 to use Redis is conservative, and on a busy site that means a slow build-up of stale keys. Bumping the GC probability and tuning the TTL to the actual session lifetime keeps the Redis instance lean.

A handful of operational checks that are worth running during the first week after the change:

  • Watch the evicted_keys and expired_keys counters in Redis to confirm garbage collection is firing
  • Tail the TYPO3 log for any new session_decode warnings, which would indicate a misconfiguration
  • Confirm that both application nodes are reading and writing sessions by checking the cmdstat counters in INFO

Hardening the Configuration for Production

Once the session corruption was gone, the client asked what else could be done to keep the stack healthy as it grew. The first recommendation was to enable HAProxy's stick-table-based session affinity, but as a fallback rather than a primary mechanism. Sticky sessions reduce the round trips to Redis for the common case, but the application has to assume at any moment that the next request will land on a different node.

The second was to switch the application servers to a shared-nothing image that does not store any state on local disk, including PHP's upload tmp directory. The site already used S3-compatible object storage in a Melbourne region for media, and pushing session storage and temporary uploads out of the local filesystem meant every node was interchangeable. Adding a third node in a Brisbane availability zone became a configuration change rather than a project.

The third was to introduce a health check on the Redis primary that the load balancer could use to drain traffic from an application server when its session backend became unreachable. This kind of graceful degradation is what makes the difference between a brief hiccup during a data centre upgrade in Sydney and a full outage that ends up on the front page of an industry publication.

If you maintain a TYPO3 site on a load-balanced infrastructure, take an hour this week to audit your session storage. The bug we chased here had been latent for months and only surfaced when traffic patterns changed. Documenting the setup on the Suema R engineering notes made the eventual fix easier to apply consistently across client sites, and the same notes are mirrored on my own blog for anyone who wants to dig deeper. Reach out if you would like a second pair of eyes on a similar setup.


cat ~/interests.json

KeyValue
editorTerminal-first workflow
osMac OS X / Unix
vcsGit, distributed version control
deployCapistrano, cron automation
graphicsSVG, Adobe Illustrator troubleshooting
networkingIP validation, SSH, VPN

git log --oneline --reverse

2013-10-30

Solving SVG import issues in Adobe Illustrator CS6 and CC

When importing an SVG into Illustrator, the operation fails with an unknown error [CANT]. A workaround for this Adobe-side bug.

2013

Solving NDK build issues on OS X

Troubleshooting native development kit compilation problems on Mac OS X.

2013

Programmatically adding PHP generated TypoScript to the backend configuration

Integrating dynamically generated TypoScript into Typo3 backend setups using PHP.

2013

ArgumentError: Could not parse PKey: no start line

Debugging an SSH key parsing error encountered during deployment.

2011-08-04

Validating IP-Addresses in PHP

Using PHP filter functions with flags like FILTER_FLAG_IPV4 and FILTER_FLAG_IPV6, and understanding how filter_var handles reserved IP addresses.

2011-07-09

Cocoa: Using NSTrackingArea

A short tutorial on using Cocoa's NSTrackingArea to capture mouseEntered and mouseExited events.


cat ~/contact.txt

github: github.com/x3ro
stackoverflow: x3ro
coderwall: coderwall.com/x3ro
twitter: @x3rames