Hands-on Linux infrastructure & operational security.
I get brought in when important systems have become fragile, opaque, risky to change, or difficult for a small team to own. The goal is usually the same: make them safer, more repeatable, and easier for humans to understand.
- Production needs stabilising without a risky rewrite
- Knowledge lives in people's heads instead of code and runbooks
- Security, privacy, or compliance requirements have increased
- You need senior infrastructure depth, but not another full-time role
What could work better?
Sometimes the right first step is architecture work; sometimes it is a broken deployment, a noisy monitoring system, an undocumented host, an access-control problem, or a restore test.
I prefer incremental improvements that remove real risk and leave a clearer system behind them.
Inherited infrastructure & stabilisation
Understand what is really running, identify the dangerous dependencies, and create a safe path out of snowflakes.
- Infrastructure inventory, dependency mapping, and risk discovery
- Incremental remediation instead of unnecessary rebuilds
- Turning undocumented production knowledge into something maintainable
Configuration & change management
Infrastructure as code and automated configuration for repeatability, reviewability, and faster recovery.
- Declarative server and application configuration
- Safe change workflows and reviewable history
- Runbooks that match what the systems actually do
Continuous integration & deployment
Build pipelines that test, ship, migrate, and roll back safely, with less manual button pushing.
- Build and test automation with clear release steps
- Database migrations and reliable rollback paths
- Secrets handling and least privilege in CI/CD
Security hardening & compliance
Practical controls around Linux, networks, identity, encryption, secrets, patching, and logging.
- Baseline hardening and patching strategy
- Identity, access, MFA/OAuth/OIDC, and secrets management
- Audit-friendly change and evidence for frameworks such as ISO 27001
Monitoring & incident response
Know when something is wrong, get useful context quickly, fix it, and use the incident to make the system better.
- Metrics, logs, alerting, and noise reduction
- Production triage, mitigation, and root-cause analysis
- UTC+10/11 coverage that can complement teams elsewhere
High availability & disaster recovery
Design for failure without pretending every system needs hyperscale architecture.
- Redundancy, replication, load balancing, and sensible failure domains
- Backups that are tested by restoring them
- Recovery runbooks and exercises people can actually execute
Read how I approach the work.
The News archive includes practical write-ups and older field notes that show the reasoning behind the services above.
Join the team, improve the system, leave useful artefacts.
Some engagements are a focused piece of work; others become ongoing support. Either way, I try to avoid creating dependency on me as the person who alone understands what changed.
- Understand: inventory, risks, constraints, and what is hurting now
- Prioritise: quick wins alongside the deeper work that matters
- Implement: hands-on changes, automation, hardening, monitoring, and testing
- Explain: documentation, runbooks, knowledge transfer, and what should happen next
- Infrastructure and configuration as code where it makes sense
- Monitoring and alerts that match your operational reality
- Runbooks for deployment, rollback, restore, and common failure modes
- Security notes that explain what changed and why
- A clearer view of remaining risks and priorities
Security, privacy, open source & public-interest work
I particularly enjoy working with organisations where operational choices have consequences beyond keeping a website online: privacy and security work, civil society, human rights, public-interest technology, and open-source projects. I also work with commercial organisations when the infrastructure problem is a good fit.