MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps
Worldwide
MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps Summary We are working on a proprietary AI/ML infrastructure engagement and are looking for an experienced SRE Architect to support our small technical team and share the hands-on engineering workload. The environment includes 50+ Kubernetes clusters across AWS, GCP, on-premises infrastructure, and other cloud providers, supporting production GPU and AI/ML inference workloads. The work will involve troubleshooting multi-cluster Kubernetes environments, managing node lifecycle activities, developing StackStorm auto-remediation workflows, maintaining Terraform and Flux CD configurations, debugging container runtimes, improving observability, and supporting structured incident response. This is a hands-on role for someone who can investigate complex infrastructure problems, write automation, implement safe solutions, and clearly document technical findings. Our Tech Stack Kubernetes & Fleet Management: Kubernetes, Rancher, multi-cluster operations, node cordon/drain/reconfiguration Automation: StackStorm/ST2, Python, Bash, event-driven auto-remediation Infrastructure & GitOps: Terraform, Flux CD, Helm, Git-based infrastructure workflows Cloud: AWS, GCP, on-premises Kubernetes; neocloud experience is a plus Container Runtime: containerd, stargz, image caching, snapshotters, cgroups GPU & AI Infrastructure: NVIDIA GPU Operator, DCGM, GPU workloads, AI/ML inference and model-serving platforms Observability: Grafana, VictoriaMetrics, Prometheus-compatible alerting, observability-as-code, runbooks Reliability: Service catalogs, SLI/SLO implementation, incident.io or similar incident-management platforms Requirements ● Proven experience as a senior SRE, SRE Architect, Platform Engineer, or MLOps Infrastructure Engineer in large-scale production environments. ● Deep Kubernetes troubleshooting experience across multiple clusters, cloud providers, and on-premises environments. ● Strong experience managing Kubernetes node lifecycle activities, including cordoning, draining, reconfiguration, recovery, and safe workload rescheduling. ● Hands-on StackStorm/ST2 experience for operational automation and auto-remediation. ● Strong Terraform, Flux CD, Rancher, Python, and Bash experience. ● Experience debugging containerd, image-cache, stargz, snapshotter, and cgroup-related issues. ● Experience operating GPU workloads using NVIDIA GPU Operator and DCGM. ● Experience supporting AI/ML inference infrastructure or model-serving platforms. ● Ability to create Grafana dashboards, alerting rules, runbooks, and observability configurations, not only monitor existing dashboards. ● Structured on-call and incident-response experience, including root-cause analysis, remediation tracking, and postmortems. ● This role is not suitable for general cloud engineers without deep Kubernetes experience or engineers who rely primarily on manual operations without scripting and automation. ● Please include brief examples of your experience with large Kubernetes fleets, StackStorm, GPU infrastructure, containerd debugging, and Terraform/Flux CD when applying
- More than 30 hrs/weekHourly
- 6+ monthsDuration
- ExpertExperience Level
$10.00
-
$25.00
Hourly- Remote Job
- Complex projectProject Type
Skills and Expertise
Activity on this job
- Proposals:20 to 50
- Last viewed by client:7 hours ago
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- United StatesIrving5:12 PM
- $1.8K total spent21 hires, 3 active
- 143 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by