status: available_for_hire

Production never sleeps. Neither do my scripts.

Freelance system administrator & infrastructure engineer. Twelve+ years keeping Linux fleets, cloud platforms, and CI pipelines humming through outages, audits, and migrations no one wants to do at 2am.

12+
years
0 sev1
incidents
$1.4M
saved_$/yr
marcus@prod-edge-01: ~/manifesto
zsh
root@sysadmin:~$ whoami
marcus_holloway — senior system administrator
root@sysadmin:~$ uptime --since
since 2013-04-12  •  servers managed: 1,240  •  uptime: 99.997%
root@sysadmin:~$ cat manifesto.txt
I keep production calm at 3am.
Linux, cloud, security & ruthless automation —
so your team can ship without paging me.
region:eu-west-3 / us-east-1
stack:linux/aws/k8s/terraform
sla:99.99% p95
response:< 15 min, 24/7
current_load:2 active engagements
next_slot:open from Q1
region:eu-west-3 / us-east-1
stack:linux/aws/k8s/terraform
sla:99.99% p95
response:< 15 min, 24/7
current_load:2 active engagements
next_slot:open from Q1
~/identity.jpg
Portrait of the system administrator
node: edge-01kernel: 6.6.21uptime: 11y
$ cat about.md

I run the boring, critical stuff so your product feels effortless.

I'm a senior system administrator and SRE who has spent the last twelve years inside terminals, runbooks, and on-call rotations — for fintechs, healthtech platforms, ecommerce, and a handful of unsexy B2B companies you've never heard of but that move serious money.

My work usually starts the moment your team says "we just need to keep it stable" and ends when nobody talks about infrastructure in standups anymore. That's the goal: invisible, observable, recoverable.

// principles
  • Boring infrastructure is good infrastructure.
  • If it isn't in git, it doesn't exist.
  • Observability before optimization.
  • Security is a default, not a feature.
  • Document like the next responder is you at 3am.
// certifications
  • AWS Solutions Architect — Pro
  • CKA — Certified Kubernetes Admin
  • RHCE — Red Hat Engineer
  • HashiCorp Terraform Associate
  • CompTIA Security+
$ ls ./services

What you can offload to me, in plain English.

Engagements typically run 2–12 weeks. Retainers available for ongoing on-call & infra ownership.

// 01

Linux & Server Operations

Hardened Debian/Ubuntu/RHEL fleets, kernel tuning, patching cadence, configuration management via Ansible & Salt.

scope_it
// 02

Cloud Infrastructure

AWS, GCP, Hetzner. Terraform-first. VPC design, autoscaling, cost reviews that actually move the bill.

scope_it
// 03

Security & Compliance

CIS hardening, SOC2 / ISO27001 prep, secrets management (Vault), IAM cleanup, incident response runbooks.

scope_it
// 04

Observability & SRE

Prometheus, Grafana, Loki, OpenTelemetry. SLOs, error budgets, on-call rotations that don't burn the team out.

scope_it
// 05

CI/CD & Automation

GitHub Actions, GitLab, Jenkins migrations. Bash & Python tooling. Deploys that take minutes, not Mondays.

scope_it
// 06

Migrations & Disaster Recovery

Datacenter → cloud, monolith → containers, MySQL/Postgres major versions. Tested DR with real RPO/RTO numbers.

scope_it

root@sysadmin:~$ stats --since=career_start

Numbers I can defend in a postmortem — not vanity metrics.

$ echo $years_in_ops
12+
$ echo $servers_managed
1,240
$ echo $uptime_average
99.997%
$ echo $client_renewals
96%
$ uname -a

The toolchain I reach for first.

Boring, battle-tested, and chosen for the team that has to operate it after I'm gone.

Linux Linux os
Windows Server os
AWS AWS cloud
GCP GCP cloud
Cloudflare Cloudflare cloud
Docker Docker platform
Kubernetes Kubernetes platform
Terraform Terraform iac
Ansible Ansible iac
Vault Vault iac
Nginx Nginx edge
Redis Redis data
Postgres Postgres data
MongoDB MongoDB data
Elastic Elastic obs
Prometheus Prometheus obs
Grafana Grafana obs
Bash Bash lang
Python Python lang
Git Git lang
OpenSSL/TLS OpenSSL/TLS sec
os = operating systems cloud = cloud / edge platform = containers / orchestration iac = infrastructure as code edge = web / proxies data = databases & caches obs = observability lang = scripting & vcs sec = security
$ tail -n 3 ./case_studies.log

Recent work, with receipts.

Names anonymized. Postmortems, runbooks & metrics available on request after NDA.

Cut p99 deploy time from 47min → 4min across 23 microservices.
fintech / payments
−91%
deploy time
Ledger-grade ledger co.

Cut p99 deploy time from 47min → 4min across 23 microservices.

  • +Replaced bespoke Jenkins with reusable GitHub Actions matrix
  • +Introduced ephemeral preview envs on Kubernetes
  • +Auto-rollback on SLO breach via Prometheus alert routing
request_full_brief
Hardened 180-host fleet & passed SOC2 Type II in 9 weeks.
healthtech / soc2
0
audit findings
Series B telehealth

Hardened 180-host fleet & passed SOC2 Type II in 9 weeks.

  • +Full CIS Level 1 baseline via Ansible across Ubuntu 22.04
  • +Centralized secrets in Vault, killed 412 hard-coded creds
  • +Wrote IR runbooks & ran the first 3 game-days with the team
request_full_brief
Migrated 14TB Postgres to AWS RDS with 22s of write downtime.
ecommerce / migration
22s
of write downtime
DTC retailer, EU

Migrated 14TB Postgres to AWS RDS with 22s of write downtime.

  • +Logical replication with pglogical, cut over inside a maintenance window
  • +Terraform-native VPC + IAM rewrite, no console clicks
  • +AWS spend dropped 38% post-rightsizing & RI plan
request_full_brief
$ cat references.txt

What the people who paid the invoice said.

"
Marcus showed up the week our payment processor caught fire. Two months later we hit our first 30-day zero-incident streak. He left behind better docs than I write for my own team.
Priya N.
VP Engineering, Series C fintech
"
We hired him for a 'small Terraform cleanup'. He quietly rewrote our entire AWS footprint, dropped the bill 38%, and trained two of our engineers to own it after he left.
Ben K.
CTO, EU ecommerce
"
The only sysadmin I've worked with who treats observability as a product. SLOs, dashboards, alerting that actually wakes the right person. Cannot recommend strongly enough.
Dr. Helena R.
Head of Platform, telehealth
$ ssh marcus@inbox

Send a real brief.
Get a real human back.

Tell me what's broken, what's coming, or what you'd like to stop losing sleep over. I read every message, and I respond within one business day.

./email
marcus@root-sysadmin.dev
./based
Lisbon, PT — remote · EU/US hours
./response_time
< 24h business · < 15min on-call
marcus@inbox: ~/new_message
vim
encrypted in transit · stored eu-west