MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps

Posted 17 hours ago

Worldwide

Summary

MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps Summary We are working on a proprietary AI/ML infrastructure engagement and are looking for an experienced SRE Architect to support our small technical team and share the hands-on engineering workload. The environment includes 50+ Kubernetes clusters across AWS, GCP, on-premises infrastructure, and other cloud providers, supporting production GPU and AI/ML inference workloads. The work will involve troubleshooting multi-cluster Kubernetes environments, managing node lifecycle activities, developing StackStorm auto-remediation workflows, maintaining Terraform and Flux CD configurations, debugging container runtimes, improving observability, and supporting structured incident response. This is a hands-on role for someone who can investigate complex infrastructure problems, write automation, implement safe solutions, and clearly document technical findings. Our Tech Stack Kubernetes & Fleet Management: Kubernetes, Rancher, multi-cluster operations, node cordon/drain/reconfiguration Automation: StackStorm/ST2, Python, Bash, event-driven auto-remediation Infrastructure & GitOps: Terraform, Flux CD, Helm, Git-based infrastructure workflows Cloud: AWS, GCP, on-premises Kubernetes; neocloud experience is a plus Container Runtime: containerd, stargz, image caching, snapshotters, cgroups GPU & AI Infrastructure: NVIDIA GPU Operator, DCGM, GPU workloads, AI/ML inference and model-serving platforms Observability: Grafana, VictoriaMetrics, Prometheus-compatible alerting, observability-as-code, runbooks Reliability: Service catalogs, SLI/SLO implementation, incident.io or similar incident-management platforms Requirements ● Proven experience as a senior SRE, SRE Architect, Platform Engineer, or MLOps Infrastructure Engineer in large-scale production environments. ● Deep Kubernetes troubleshooting experience across multiple clusters, cloud providers, and on-premises environments. ● Strong experience managing Kubernetes node lifecycle activities, including cordoning, draining, reconfiguration, recovery, and safe workload rescheduling. ● Hands-on StackStorm/ST2 experience for operational automation and auto-remediation. ● Strong Terraform, Flux CD, Rancher, Python, and Bash experience. ● Experience debugging containerd, image-cache, stargz, snapshotter, and cgroup-related issues. ● Experience operating GPU workloads using NVIDIA GPU Operator and DCGM. ● Experience supporting AI/ML inference infrastructure or model-serving platforms. ● Ability to create Grafana dashboards, alerting rules, runbooks, and observability configurations, not only monitor existing dashboards. ● Structured on-call and incident-response experience, including root-cause analysis, remediation tracking, and postmortems. ● This role is not suitable for general cloud engineers without deep Kubernetes experience or engineers who rely primarily on manual operations without scripting and automation. ● Please include brief examples of your experience with large Kubernetes fleets, StackStorm, GPU infrastructure, containerd debugging, and Terraform/Flux CD when applying

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Expert
    Experience Level
  • $10.00

    -

    $25.00

    Hourly
  • Remote Job
  • Complex project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Terraform
Bash
Kubernetes
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:7 hours ago
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Mar 1, 2017
  • United States
    Irving5:12 PM
  • $1.8K total spent
    21 hires, 3 active
  • 143 hours

Explore similar jobs on Upwork

Docker
CI/CD
Git
GitLab
Linux System Administration
DevOps
Performance Testing
Docker
DevOps
Linux System Administration

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo