Kubernetes Platform Audit & Linux Infrastructure Hardening for SaaS Scale-Up

Posted 2 days ago

Worldwide

Summary

We're a bootstrapped-to-profitable B2B SaaS company (analytics platform, ~200 paying customers, ~15M requests/day). Stack is Go and Python microservices, PostgreSQL and Redis, all running on self-managed Kubernetes (kubeadm) across a fleet of Ubuntu VMs split between AWS EC2 and Hetzner bare metal. We have 12 engineers, one of whom "does DevOps when he has time." The cluster was set up two years ago by a contractor who's no longer around. It works - until it doesn't. We've had three incidents this quarter (a full-disk etcd node, a cert expiry that took down the API server, and a noisy-neighbor pod that OOM-killed production). We need a senior Kubernetes / Linux platform engineer who has run clusters in anger to audit what we have, fix the scary parts, and leave us with something maintainable. What we're looking for Full Kubernetes cluster health and architecture review: control plane HA, etcd backup/restore (we currently have neither tested), upgrade path (we're 3 minor versions behind), findings prioritized by blast radius. Decision support and migration plan: stay on self-managed kubeadm vs. move to EKS - with a real cost and operational-overhead comparison, not a sales pitch. Linux hardening across the node fleet: CIS benchmarks, kernel parameters, SSH lockdown, automated patching strategy (unattended-upgrades vs. immutable node images), auditd. Kubernetes security hardening: RBAC least-privilege review (everything currently runs as cluster-admin), Pod Security Standards, NetworkPolicies (we have zero), secrets encryption at rest, image scanning and admission control (Kyverno or Gatekeeper). Resource management: requests/limits across all workloads, HPA where it makes sense, PodDisruptionBudgets, priority classes - so one bad deploy can't take down the cluster again. Observability: Prometheus / Grafana / Alertmanager stack with alerts that actually page on the right things, plus centralized logging (Loki or ELK). We currently SSH into nodes and grep. GitOps: move deployments from kubectl apply off laptops to ArgoCD or Flux with proper environments and rollback. Ingress and TLS: sane ingress-nginx (or alternative) setup with cert-manager so certificates never silently expire again. Disaster recovery runbook: tested etcd restore, node replacement procedure, and a written incident playbook. Documentation and knowledge transfer sessions so our team can operate this without you. Requirements Deep hands-on Kubernetes experience: you've administered self-managed clusters (kubeadm/kubespray), not just consumed EKS/GKE. Strong Linux systems engineering: systemd, networking, storage, performance debugging, security hardening on Ubuntu/Debian. Production experience with Prometheus, Grafana, and at least one GitOps tool (ArgoCD / Flux). IaC fluency: Terraform for the cloud layer, Ansible (or similar) for node configuration. You've handled real incidents - etcd recovery, control plane failures, cascading OOMs - and can talk about them specifically. Available for meaningful overlap with Central European business hours. Nice to have CKA / CKS certification. Experience running hybrid cloud + bare-metal clusters (Hetzner, OVH, Equinix). Cilium / eBPF networking experience. Cost optimization experience - our AWS bill and Hetzner fleet have both grown without anyone checking. Willing to stay on retainer for ongoing platform support and on-call escalation. Please include specific examples of Kubernetes clusters you've audited or rescued, what you found, and what you fixed - plus your estimated timeline. Long-term potential for the right person.

  • $500.00

    Fixed-price
  • Expert
    Experience Level
  • Remote Job
  • One-time project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Terraform
Docker
Linux System Administration
Activity on this job
  • Proposals:20 to 50
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Aug 16, 2026
  • Slovakia
    4:34 PM

Explore similar jobs on Upwork

Apache and PHP-FPM Configuration ExpertHourly‐ Posted 2 months ago
Ubuntu
Apache HTTP Server
Senior DevOps & SRE / Infrastructure EngineerFixed-price‐ Posted 3 weeks ago
DevOps

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo