whoami
Lucas Jenß
cat /etc/motd
The Coding Journal ツ — Notes taken on an epic coding journey. Technical solutions, debugging notes, and practical guides from the trenches of software development.
ls -la ~/languages/
- drwxr-xr-x
- ▶ PHP
- ▶ Ruby
- ▶ Scala
- ▶ C#
- ▶ JavaScript
- ▶ Objective-C
- ▶ Shell Scripting
ls -la ~/toolchain/
- drwxr-xr-x
- ▶ Typo3
- ▶ Akka
- ▶ Capistrano
- ▶ Git
- ▶ MAMP
- ▶ Adobe Illustrator
- ▶ NSTrackingArea (Cocoa)
uname -a
- drwxr-xr-x
- ▶ Mac OS X
- ▶ Unix
Akka Failure Detection And Supervision For Resilient Actors
Akka applications are built around the assumption that work will fail. An actor can throw an exception, a node can disappear during a deployment, or a network route can become unreliable without warning. Resilience comes from treating these events as part of the system’s normal operating conditions rather than as exceptional surprises.
Failure detection and supervision solve different problems. A failure detector estimates whether another process is reachable, while supervision decides how an actor should respond when its child throws an exception. Confusing those responsibilities often leads to systems that restart healthy actors, hide serious defects, or continue processing with invalid state.
This distinction matters in production environments where connectivity is variable. A service running in Sydney may briefly lose contact with a node in Melbourne because of a cloud networking issue, an overloaded NBN connection, or a rolling deployment. The right response is usually measured and observable, rather than an immediate shutdown of the whole actor hierarchy.
The practical notes collected on Akka development are useful alongside the official documentation because small configuration details often determine whether recovery works as expected. Actor paths, dispatcher capacity, mailbox behaviour and cluster roles all influence the result.
Separate Reachability From Application Failure
Akka’s cluster failure detector uses heartbeats and an accrual algorithm to calculate how suspicious a missing response appears. The phi value is a suspicion level, not a simple Boolean. Short pauses may be tolerated, while a sustained lack of heartbeats eventually causes the node to be considered unreachable.
This approach is useful because garbage collection, CPU starvation and temporary network congestion can delay a heartbeat. A detector that marks a node dead after one missed response would create false positives. At the same time, excessively generous thresholds make the cluster slow to react to a genuinely failed machine.
The detector does not prove that an application is healthy. A node may answer heartbeats while its business actors are blocked, its database pool is exhausted, or its mailbox is growing without limit. Health checks, request timeouts, circuit breakers and business-level metrics should complement cluster membership events.
Treat an unreachable event as information that requires a policy. Some applications should stop sending work to the node immediately. Others may preserve queued commands and retry them after a backoff. Repeatedly retrying without limits can turn a brief outage into a traffic storm, especially during the busy morning period when Australian services receive concentrated mobile and web traffic.
Design Supervision Around Actor Responsibilities
A supervisor monitors child actors and handles failures that occur while processing messages. In classic Akka, common directives are resume, restart, stop and escalate. Akka Typed expresses similar decisions through a SupervisorStrategy, often combined with a backoff strategy for repeated crashes.
Resuming keeps the actor’s current state and skips the failed message. It is appropriate only when the exception is understood and the state remains trustworthy. Restarting reconstructs the actor’s behaviour and usually clears in-memory state, making it suitable for recoverable resources such as a connection wrapper. Stopping is safer when the actor cannot continue without risking corruption.
Escalation passes the failure to a parent, which can apply a broader policy. This creates a useful hierarchy: a small worker may restart, a group manager may stop the failed worker, and a service guardian may terminate the component if its invariant has been broken. Supervision should reflect ownership of state and responsibility, rather than being applied uniformly across the tree.
A restart is not a substitute for handling the message that caused the exception. Unless the message is safely retried, persisted, or sent to a dead-letter workflow, it may be lost. Idempotent commands, durable event streams and explicit acknowledgement protocols make recovery much more predictable.
Select Recovery Policies Deliberately
Transient infrastructure faults deserve a different response from programming defects. A database timeout may justify a bounded retry with exponential backoff. A NullPointerException caused by an invalid assumption should generally be logged with context and allowed to reach a policy that prevents endless restart loops.
Akka’s restart limits are important safeguards. A supervisor can permit a small number of restarts within a time window and then stop or escalate the actor. Without a limit, an actor that fails instantly can consume CPU, fill logs and repeatedly reconnect to a broken dependency. Backoff supervisors reduce this pressure while making the recovery attempt visible in metrics.
State placement also determines the correct strategy. Ephemeral state can often be discarded during a restart, but financial transactions, inventory reservations and customer notifications require durable coordination. Australian businesses handling personal information should also consider the Privacy Act and the Notifiable Data Breaches scheme: logs and dead-letter messages must not casually expose sensitive customer data during diagnosis.
A useful pattern is to keep side effects behind dedicated actors or adapters. A command-processing actor can validate input and coordinate work, while a database actor owns connection recovery. This narrows the blast radius and allows each component to have a clear restart, stop or retry policy.
Make Cluster Recovery Observable
Cluster membership changes should be treated as operational events. Subscribe to the relevant cluster or typed system notifications, record the affected address and role, and expose counters for reachable, unreachable and removed members. Logs should include correlation identifiers so an operator can connect a failed request with the actor and node that handled it.
Quarantine deserves special attention in systems using remote messaging. When Akka detects a serious association problem, it may prevent further communication with a remote system to avoid inconsistent behaviour. Restarting an application process may not clear every situation; the deployment and remoting configuration must support a clean recovery path.
Avoid using cluster membership as a substitute for application-level ownership. When a node disappears, another actor may need to take over a shard, lease or partition. That handover should be coordinated using explicit state and fencing, so an old process cannot continue writing after a new owner has been selected.
Operational assumptions should match the local deployment environment. A service spread across Sydney and Brisbane may experience different latency and maintenance windows, while daylight-saving changes between states can complicate scheduled messages. Store timestamps in UTC, use monotonic timing for intervals and test failover across the actual network topology rather than relying solely on a laptop development cluster.
Test Failure Paths Before Production
A resilient actor system needs tests for delayed heartbeats, dropped messages, slow persistence, repeated exceptions and partial recovery. TestKit can drive messages through an actor hierarchy, while cluster test tools can simulate member changes and network conditions. The aim is to verify observable outcomes: no duplicate side effect, bounded retries, correct handover and useful diagnostics.
Fault injection should include failures that look similar but require different responses. A stopped node, a saturated node and a partitioned node may all appear unreachable from one perspective. A test that covers only process termination will miss mailbox overload, dispatcher starvation and a database that accepts connections but never completes queries.
The following comparison helps keep the policies distinct:
| Situation | Primary signal | Suitable response | Main risk |
|---|---|---|---|
| Child throws a known, recoverable exception | Supervision event | Restart or bounded retry | Repeating a bad message |
| Child loses trusted in-memory state | Supervision event | Stop and recreate from durable state | Continuing with corrupted state |
| Remote node misses heartbeats | Failure detector | Quarantine, reroute or wait for membership change | False positive during a pause |
| Database or HTTP dependency times out | Request timeout | Backoff, circuit breaker and retry limit | Cascading overload |
| Cluster member disappears during ownership change | Membership and lease events | Fence old owner and transfer responsibility | Duplicate side effects |
Review telemetry after each injected failure. A dashboard should show actor restarts, escalation counts, dead letters, mailbox depth, heartbeat suspicion and recovery duration. Alert on patterns, such as a rising restart rate or repeated unreachable transitions, instead of paging someone for every isolated transient event.
Clear runbooks are valuable for teams working across Australian time zones and on-call rotations. They should state which node roles can be restarted, which messages are safe to replay, how personal data is protected in logs and when a cluster should be deliberately shut down. Treating incident procedures as part of the design makes supervision decisions easier to trust.
A useful way to explain these ideas to a mixed engineering team is to compare actor roles with coordinated positions in a local team sport; a readable football team story can provide that human reference point. Each role has a defined responsibility, and replacing one player should not require abandoning the entire team’s plan.
Build the failure paths into your Akka design from the first actor diagram: define ownership, choose restart limits, protect side effects, instrument membership changes and test realistic partitions. Then review the policies against production data and refine them before a real outage turns an untested assumption into customer impact.
cat ~/interests.json
| Key | Value |
|---|---|
| editor | Terminal-first workflow |
| os | Mac OS X / Unix |
| vcs | Git, distributed version control |
| deploy | Capistrano, cron automation |
| graphics | SVG, Adobe Illustrator troubleshooting |
| networking | IP validation, SSH, VPN |
git log --oneline --reverse
Solving SVG import issues in Adobe Illustrator CS6 and CC
When importing an SVG into Illustrator, the operation fails with an unknown error [CANT]. A workaround for this Adobe-side bug.
Solving NDK build issues on OS X
Troubleshooting native development kit compilation problems on Mac OS X.
Programmatically adding PHP generated TypoScript to the backend configuration
Integrating dynamically generated TypoScript into Typo3 backend setups using PHP.
ArgumentError: Could not parse PKey: no start line
Debugging an SSH key parsing error encountered during deployment.
Validating IP-Addresses in PHP
Using PHP filter functions with flags like FILTER_FLAG_IPV4 and FILTER_FLAG_IPV6, and understanding how filter_var handles reserved IP addresses.
Cocoa: Using NSTrackingArea
A short tutorial on using Cocoa's NSTrackingArea to capture mouseEntered and mouseExited events.
cat ~/contact.txt