Hire the Best Distributed Systems Engineers

More than 3,000 reviews on G2
Rating is 4.5 out of 5.
4.5/5
of Upwork by G2 peer reviewers
Anatoliy K.

Bratislava, Slovakia

$55/hr
5.0
6 jobs

Software engineer, Linux enthusiast - Leading tech decisions, architecture design; - Multi-model AI orchestration: OpenAI (GPT), Perplexity (Sonar-Pro), Google Gemini; function-calling agents with massive context - Real-time WebSocket infrastructure; Socket-io with Redis adapter for horizontal scaling - AWS-native architecture: ECS Fargate, Lambda, S3, SQS, EventBridge, CloudFront, Cognito - Infrastructure as Code; Terraform, GitHub Actions CI/CD, zero-downtime deployments - High-load API development; NestJS, Node.js, TypeScript - PostgreSQL (Prisma ORM), Redis/ElastiCache

  • Python
  • React
  • Node.js
  • WebRTC
  • TypeScript
  • CI/CD
  • Terraform
  • AWS Lambda
  • Kubernetes
  • Amazon ECS
  • Apache Kafka
  • Hyperledger Fabric
  • PostgreSQL
  • NestJS
  • NestJS Development
  • Websockets
  • Socket.io
  • Product Concept
Choudhry S.

Haripur, Pakistan

$65/hr
5.0
23 jobs

๐Ÿ”น 5 Years Experience| ๐Ÿ”น 100% JSS | ๐Ÿ”นDeployed Enterprise Infrastructure | ๐Ÿ”น 20+ Nvidia GPUs cluster Deployed | ๐Ÿ”น 30+ Kubernetes GPU Clusters Deployed | ๐Ÿ”น 50+ LLMs deployments | ๐Ÿ”น 20+ AI Servers Designed & Deployed | I wrote the IaC for GPU cluster orchestration on KAUST Shaheen III, the Middle East's #1 supercomputer (ranked 18th globally). That's the level of GPU and LLM infrastructure I work on, on-prem and in the cloud. ๐Ÿ“Œ What I'm an expert in: ๐Ÿ”นSelf-hosted & local LLM serving (vLLM, NVIDIA Triton, TensorRT, SGLang) ๐Ÿ”นGPU inference optimization (FP8/INT8 quantization, FlashAttention, tensor parallelism, MIG on H100/A100/GH200/Blackwell) ๐Ÿ”นLLM fine-tuning & customization (LoRA, PEFT, Unsloth, domain adaptation) ๐Ÿ”นBare-metal GPU Kubernetes (Proxmox, Rancher, K3s, OpenShift, Calico, MetalLB, HAProxy) ๐Ÿ”นManaged cloud Kubernetes with GPU (GKE, EKS, AKS, Vertex AI, SageMaker, Cloud Run) ๐Ÿ”นNVIDIA GPU Operator, Kubeflow, Ray, Triton Inference Server, custom K8s operators (I built VRAM GPU Operator + Aliyun GPU Resource Manager for Kubeflow) ๐Ÿ”นPrivate & on-prem LLM deployments (HIPAA, GDPR, air-gapped, regulated teams) ๐Ÿ”นHeterogeneous HPC & AI infrastructure (GPU, FPGA, ASIC, XPU, multi-accelerator scheduling) ๐Ÿ”นAI agents & orchestration (LangGraph, LangChain, CrewAI, Claude Agent SDK, MCP servers) ๐Ÿ”นVoice agent infrastructure (Retell AI, Vapi, Twilio, ElevenLabs, n8n, WhatsApp integrations) ๐Ÿ”นRAG & document intelligence (Docling, Microsoft GraphRAG, pgvector, FAISS, Pinecone, Qdrant, ChromaDB) ๐Ÿ”นInfrastructure-as-code (Terraform, Ansible, Pulumi, Crossplane, Helm) ๐Ÿ”นCI/CD for ML/AI (GitHub Actions, GitLab CI, Jenkins, CircleCI, ArgoCD, Flux) ๐Ÿ”นObservability for AI workloads (Prometheus, Grafana, Loki, Datadog, OpenTelemetry) ๐Ÿ”นCloud spend optimization for AI infrastructure ๐Ÿ“Œ Real results, defensible on a call: ๐Ÿ”นCut vLLM inference latency from 3.2s to 800ms at a US startup ๐Ÿ”นDropped their monthly cloud bill from $10K to $4K ๐Ÿ”นServed LLMs to 1,000+ concurrent users on bare-metal Kubernetes, zero unplanned downtime for 12 months ๐Ÿ”นShipped the Docling + GraphRAG pipeline for commercial real estate lease analysis at Nestbox AI, deployed on GCP with a human review queue ๐Ÿ“Œ What clients say: "single-handedly built our Kubernetes infrastructure for our GPU setup" "a full stack AI guy" "energetic, hungry, and sharp" "gets concepts quickly and is an incredibly hard worker" "when he says he'll do it by a deadline, he does. Owns accountability" "approaches the project as if it were his own" "worked late into the night to get it done" "delivered what he promised and provided clear documentation" "incredibly efficient, addressed every concern in 30 minutes" "a safe bet" A bit about me: I'm Acting CTO at Peregrine Ventures AI, currently building a product called Deplexify. I also run a dual RTX Pro 6000 Blackwell server at home (EPYC, Proxmox, GPU passthrough), so on-prem GPU work isn't a side skill. It's my daily reality.

  • Kubernetes
  • DevOps Engineering
  • Linux System Administration
  • MLOps
  • Terraform
  • Docker
  • Machine Learning
  • Computer Vision
  • Artificial Intelligence
  • Natural Language Processing
  • LLM Prompt Engineering
  • AI App Development
  • Kubeflow
  • Python
  • Bash
Francisco S.

Valparaiso, Chile

$96/hr
5.0
4 jobs

Hi, I'm Fran ๐Ÿ‘‹ I architect and ship production AI systems and cloud infrastructure that actually scale. โ†’ Architected & shipped a production AI copilot (agentic, RAG-grounded, human-in-the-loop) now serving customers โ†’ Sr DevOps running cloud infra for a NASDAQ-listed biotech, supporting Twist Bioscience (NASDAQ: TWST, $2B+) โ†’ Cut report generation time 50% at IBM ($150B+ market cap) with Python microservices โ†’ Cut cloud infrastructure costs 30% for a US biotech SaaS company using GCP rightsizing + autoscaling โ†’ 2ร— release velocity at a US biotech SaaS by streamlining CI/CD โ†’ Built recurring AI consulting from $0 to $1,500+/client serving LATAM tech professionals 8+ years building production systems that move real money 24/7. CKA + CKAD certified (Linux Foundation / Cloud Native Computing Foundation). ๐Ÿ’ผ What I do: โ†’ Cloud architecture (AWS, GCP, Azure) โ€” Kubernetes, Terraform, ArgoCD โ†’ AI Agents, RAG & Automation โ€” Claude / OpenAI, Python, production-grade โ†’ Infrastructure cost optimization โ€” typical 25-40% savings โ†’ CI/CD pipeline acceleration โ€” typical 3-5ร— speedup โ†’ Production systems on your existing stack (no rip-and-replace) ๐ŸŽฏ Best fit for: โ†’ B2B SaaS with infrastructure scaling challenges โ†’ Legal/professional firms needing AI document automation โ†’ Marketing agencies needing content automation systems โ†’ Teams needing senior engineering on fractional/project basis Stack: Kubernetes ยท Terraform ยท ArgoCD ยท AWS ยท GCP ยท Azure ยท Python ยท Go ยท Claude Code ยท GitHub Actions Let's chat ๐Ÿ‘‡

  • Python
  • DevOps
  • Kubernetes
  • Docker
  • Terraform
  • AI Agent Development
  • CI/CD
  • Cloud Architecture
  • Google Cloud Platform
  • Infrastructure as Code
  • Amazon Web Services
  • Microsoft Azure
  • Prometheus
  • Grafana
  • Bash
  • HighLevel
  • n8n
  • Make.com
Mikhail S.

Les Echelles, France

$45/hr
4.6
248 jobs

๐Ÿ‘‹ Iโ€™m a Solution Architect and Cloud Engineer with 20+ years of hands-on experience building, fixing, and scaling production systems on AWS, Kubernetes, and Terraform. I help companies solve the most common infrastructure headaches: lowering runaway AWS costs, fixing unstable deployments, and cleaning up manual "click-ops" messes. As a Solution Architect, I transform fragile setups into resilient, automated environments. Whether you need to optimize a slow Kubernetes cluster, refactor unmaintainable Terraform code, or secure a leaking AWS environment, I provide the architectural oversight and hands-on execution to ensure your systems are production-ready and cost-effective for the long haul. I help teams turn messy infrastructure into clear, automated DevOps environments using AWS, Kubernetes, and Terraform as the core stack. Most projects involve fixing unstable AWS setups, rebuilding pipelines with Terraform, and moving workloads into production-ready Kubernetes clusters. I also work with GCP and Azure when projects require multi-cloud or provider-specific services. ๐Ÿš€ Core DevOps Stack DevOps with AWS, Kubernetes, and Terraform is the foundation of my work. As a Cloud Engineer, I design full DevOps systems where AWS infrastructure is managed by Terraform and applications run on Kubernetes with automated pipelines. My daily work includes: - Solution Architect oversight for AWS/Kubernetes migrations - Cloud Engineer automation on AWS using Terraform - Kubernetes cluster design and operations - Terraform modules for repeatable AWS environments - Production DevOps pipelines for Kubernetes workloads - Multi-cloud setups across AWS, GCP, and Azure โš™๏ธ What I Deliver โ˜๏ธ DevOps & Cloud Infrastructure - Full DevOps setups on AWS using Terraform - High-availability AWS environments managed through Terraform - Cost-optimized AWS architectures built with Solution Architect best practices - Secure AWS networking and access policies via Terraform - Production setups on GCP and Azure when required โ˜ธ๏ธ Kubernetes Platforms - Production Kubernetes clusters on AWS - Scalable Kubernetes environments managed by Terraform - DevOps pipelines for Kubernetes deployments - Monitoring and autoscaling for Kubernetes workloads - Kubernetes clusters on GCP (GKE) and Azure (AKS) ๐Ÿ”ง Terraform Automation - End-to-end Terraform infrastructure on AWS - Modular Terraform code for repeatable DevOps setups - Terraform pipelines for Kubernetes clusters - Automated AWS provisioning with Terraform - Multi-cloud Terraform environments across AWS, GCP, and Azure ๐Ÿ”„ DevOps CI/CD - DevOps pipelines deploying to AWS and Kubernetes - Container builds and rollouts into Kubernetes - Terraform driven release environments on AWS - Zero-downtime DevOps deployments - CI/CD for systems running on GCP and Azure ๐Ÿ“Š Monitoring, Security, and Databases - DevOps monitoring for AWS and Kubernetes - Metrics and alerts for Kubernetes clusters - Log aggregation for AWS and Kubernetes - Secure Terraform secrets and access management - Database tuning in AWS, GCP, and Azure environments ๐Ÿ’ก Why Clients Choose My DevOps Work - Long-term infrastructure support and maintenance (SLA-focused) - Real production experience as a Cloud Engineer on AWS, GCP, and Azure - Stable Kubernetes clusters designed by a professional Solution Architect - Clear Terraform code your team can maintain - Predictable AWS infrastructure costs Most of my projects involve: - Fixing broken DevOps pipelines - Rebuilding AWS environments with Terraform - Moving apps into stable Kubernetes clusters - Improving DevOps reliability and monitoring - Supporting multi-cloud setups on AWS, GCP, and Azure ๐Ÿ› ๏ธ Typical DevOps Projects - Migrate legacy servers into AWS with Terraform - Build new Kubernetes clusters for SaaS platforms - Set up full DevOps pipelines on AWS - Refactor manual infrastructure into Terraform - Stabilize production Kubernetes environments - Deploy or migrate workloads to GCP or Azure If you need a Cloud Engineer and Solution Architect who works daily with AWS, Kubernetes, and Terraform, and can provide long-term infrastructure support, I can help you build an automated, stable, and scalable system. Letโ€™s discuss your setup and improve your infrastructure.

  • DevOps
  • Amazon Web Services
  • Docker
  • CI/CD
  • Kubernetes
  • Linux
  • Python
  • Deployment Automation
  • Terraform
  • Linux System Administration
  • Google Cloud Platform
  • Amazon EC2
  • Git
  • NGINX
  • CI/CD Platform
  • Configuration Management
  • DevOps Engineering
  • Ansible
  • Jenkins
Fendri F.

Sfax, Tunisia

$59/hr
5.0
65 jobs

๐Ÿ† ๐•‹๐• ๐•ก 1% - ๐•‹๐• ๐•ก โ„๐•’๐•ฅ๐•–๐•• โ„™๐•๐•ฆ๐•ค ๐Ÿง‘๐Ÿปโ€๐Ÿคโ€๐Ÿง‘๐Ÿป 60+ โ„๐•’๐•ก๐•ก๐•ช ๐•”๐•๐•š๐•–๐•Ÿ๐•ฅ๐•ค โฑ๏ธ 6000+ ๐•‹๐• ๐•ฅ๐•’๐• โ„๐• ๐•ฆ๐•ฃ๐•ค ๐ŸŽญ 100% ๐•๐•Š๐•Š) Hi ๐Ÿ‘‹, I specialize in designing and deploying ๐™ž๐™ฃ๐™ฃ๐™ค๐™ซ๐™–๐™ฉ๐™ž๐™ซ๐™š, ๐™จ๐™˜๐™–๐™ก๐™–๐™—๐™ก๐™š, ๐™–๐™ฃ๐™™ ๐™จ๐™š๐™˜๐™ช๐™ง๐™š cloud-native solutions๐Ÿ’ก. With more than ๐Ÿ– years of experience as ๐˜พ๐™ก๐™ค๐™ช๐™™ ๐˜ผ๐™ง๐™˜๐™๐™ž๐™ฉ๐™š๐™˜๐™ฉ, ๐™†๐™ช๐™—๐™š๐™ง๐™ฃ๐™š๐™ฉ๐™š๐™จ ๐˜ผ๐™™๐™ข๐™ž๐™ฃ๐™จ๐™ฉ๐™ง๐™–๐™ฉ๐™ค๐™ง ๐™–๐™ฃ๐™™ ๐˜ฟ๐™š๐™ซ๐™Ž๐™š๐™˜๐™Š๐™ฅ๐™จ ๐™€๐™ฃ๐™œ๐™ž๐™ฃ๐™š๐™š๐™ง. Tฬฒhฬฒiฬฒsฬฒ ฬฒiฬฒsฬฒ ฬฒmฬฒyฬฒ ฬฒPฬฒrฬฒoฬฒfฬฒeฬฒsฬฒsฬฒiฬฒoฬฒnฬฒaฬฒlฬฒ ฬฒCฬฒeฬฒrฬฒtฬฒiฬฒfฬฒiฬฒcฬฒaฬฒtฬฒiฬฒoฬฒnฬฒ ฬฒ:ฬฒ ๐Ÿฅ‡ ๐‘ฒ๐’–๐’ƒ๐’†๐’“๐’๐’†๐’•๐’†๐’” Administrator CKA ๐Ÿฅ‡ ๐‘ฒ๐’–๐’ƒ๐’†๐’“๐’๐’†๐’•๐’†๐’” Security specialist CKS ๐Ÿฅ‡ ๐‘ฒ๐’–๐’ƒ๐’†๐’“๐’๐’†๐’•๐’†๐’” Application Developer CKAD ๐Ÿฅ‡ ๐‘ฒ๐’–๐’ƒ๐’†๐’“๐’๐’†๐’•๐’†๐’” ๐’‚๐’๐’… ๐‘ช๐’๐’๐’–๐’… ๐‘ต๐’‚๐’•๐’Š๐’—๐’† Associate (KCNA) ๐Ÿฅ‡ ๐‘ฒ๐’–๐’ƒ๐’†๐’“๐’๐’†๐’•๐’†๐’” ๐’‚๐’๐’… ๐‘ช๐’๐’๐’–๐’… ๐‘ต๐’‚๐’•๐’Š๐’—๐’† ๐‘บ๐’†๐’„๐’–๐’“๐’Š๐’•๐’š Associate (KCSA) ๐Ÿฅ‡ ๐‘จ๐‘พ๐‘บ Solution Architect ๐Ÿฅ‡ ๐‘จ๐’›๐’–๐’“๐’† DevOps Expert๐Ÿฅ‡ ๐‘ป๐’†๐’“๐’“๐’‚๐’‡๐’๐’“๐’Ž Associate๐Ÿฅ‡ ๐‘จ๐’“๐’ˆ๐’ ๐‘ช๐‘ซ, GitOps At Scale ๐Ÿฅ‡ Certified ๐‘ฐ๐’”๐’•๐’Š๐’ and ๐‘ฌ๐’๐’—๐’๐’š Service Mesh With my expertise in cloud-native design, I can help you build scalable, high-performance applications using microservices, deployed seamlessly in containers on private, public, or hybrid cloud platforms. This will be realized using this ๐Ÿ†ƒ๐Ÿ…พ๐Ÿ…พ๐Ÿ…ป๐Ÿ†‚ ๐Ÿ› ๏ธ following the ๐™Ž๐™š๐™ซ๐™š๐™ฃ ๐™ˆ๐™ค๐™™๐™š๐™ก๐™จ ๐™ค๐™› ๐˜พ๐™ก๐™ค๐™ช๐™™ ๐™‰๐™–๐™ฉ๐™ž๐™ซ๐™š ๐˜ฟ๐™š๐™จ๐™ž๐™œ๐™ฃ ๐Ÿ”ฅ : ๐Ÿ‘จโ€๐Ÿ’ป ๐Ÿ. ๐Œ๐จ๐๐ž๐ซ๐ง ๐ƒ๐ž๐ฌ๐ข๐ ๐ง & ๐ƒ๐ž๐ฏ๐ž๐ฅ๐จ๐ฉ๐ฆ๐ž๐ง๐ญ ๐Œ๐จ๐๐ž - Programming Languages: ๐™‚๐™ค, ๐™‹๐™ฎ๐™ฉ๐™๐™ค๐™ฃ, ๐™‰๐™ค๐™™๐™š ๐™Ÿ๐™จ, ๐˜ฝ๐™–๐™จ๐™ - Containerization Tools: ๐˜ฟ๐™ค๐™˜๐™ ๐™š๐™ง, ๐˜ฟ๐™ค๐™˜๐™ ๐™š๐™ง-๐˜พ๐™ค๐™ข๐™ฅ๐™ค๐™จ๐™š, ๐™†๐™–๐™ฃ๐™ž๐™ ๐™ค, ๐™‹๐™ค๐™™๐™ข๐™–๐™ฃ - Container Image Registry: ๐™ƒ๐™–๐™ง๐™—๐™ค๐™ง, ๐˜ฟ๐™ค๐™˜๐™ ๐™š๐™ง๐™๐™ช๐™— ๐™‚๐™ž๐™ฉ๐™ก๐™–๐™— ๐™๐™š๐™œ๐™ž๐™จ๐™ฉ๐™ง๐™ฎ, ๐™‚๐™ž๐™ฉ๐™๐™ช๐™— ๐™๐™š๐™œ๐™ž๐™จ๐™ฉ๐™ง๐™ฎ - AฬฒPฬฒIฬฒ ฬฒDฬฒrฬฒiฬฒvฬฒeฬฒnฬฒ: ๐˜ผ๐™ฅ๐™ž๐™จ๐™ž๐™ญ, ๐™„๐™จ๐™ฉ๐™ž๐™ค ๐™‚๐™–๐™ฉ๐™š๐™ฌ๐™–๐™ฎ, ๐™†๐™ค๐™ฃ๐™œ, ๐™๐™ง๐™–๐™š๐™›๐™ž๐™ , ๐˜ผ๐™ข๐™—๐™–๐™จ๐™จ๐™–๐™™๐™ค๐™ง, ๐™๐™ฎ๐™  - Mฬฒoฬฒdฬฒeฬฒrฬฒnฬฒ ฬฒDฬฒaฬฒtฬฒaฬฒbฬฒaฬฒsฬฒeฬฒsฬฒ: ๐˜พ๐™–๐™จ๐™จ๐™–๐™ฃ๐™™๐™ง๐™–, ๐™€๐™ก๐™–๐™จ๐™ฉ๐™ž๐™˜๐™จ๐™š๐™–๐™ง๐™˜๐™, ๐˜พ๐™ค๐™ช๐™˜๐™๐™™๐™—, ๐™๐™š๐™™๐™ž๐™จ, ๐™‹๐™ค๐™จ๐™ฉ๐™œ๐™ง๐™š๐™Ž๐™Œ๐™‡, ๐™‘๐™ž๐™ฉ๐™š๐™จ๐™จ - Eฬฒvฬฒeฬฒnฬฒtฬฒ-ฬฒDฬฒrฬฒiฬฒvฬฒeฬฒnฬฒ ฬฒDฬฒeฬฒsฬฒiฬฒgฬฒnฬฒ:ฬฒ ๐™‰๐˜ผ๐™๐™Ž, ๐™๐™–๐™—๐™—๐™ž๐™ฉ๐™ˆ๐™Œ, ๐™†๐™–๐™›๐™ ๐™–, ๐˜ผ๐™˜๐™ฉ๐™ž๐™ซ๐™š๐™ˆ๐™Œ, ๐˜ฝ๐™š๐™–๐™ฃ๐™จ๐™ฉ๐™–๐™ก๐™ ๐™™ ๐Ÿ—๏ธ ๐Ÿ. ๐Œ๐จ๐๐ž๐ซ๐ง ๐ˆ๐ง๐Ÿ๐ซ๐š๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ฎ๐ซ๐ž/๐ƒ๐ž๐ฏ๐Ž๐ฉ๐ฌ - ๐‚๐ˆ/๐‚๐ƒ - CI/CD Tools: ๐™‚๐™ž๐™ฉ๐™ก๐™–๐™— ๐˜พ๐™„, ๐™‚๐™ž๐™ฉ๐™๐™ช๐™— ๐˜ผ๐™˜๐™ฉ๐™ž๐™ค๐™ฃ๐™จ, ๐™…๐™š๐™ฃ๐™ ๐™ž๐™จ, ๐˜ฟ๐™–๐™œ๐™œ๐™š๐™ง, ๐™๐™ง๐™–๐™ซ๐™ž๐™จ ๐˜พ๐™„, ๐˜พ๐™ž๐™ง๐™˜๐™ก๐™š ๐˜พ๐™„, ๐™๐™š๐™ ๐™ฉ๐™ค๐™ฃ. - Security Scanning Tools: ๐™๐™ง๐™ž๐™ซ๐™ฎ, ๐˜ฟ๐™ค๐™˜๐™ ๐™š๐™ง ๐™Ž๐™š๐™˜๐™ช๐™ง๐™ž๐™ฉ๐™ฎ ๐™Ž๐™˜๐™–๐™ฃ๐™ฃ๐™š๐™ง, ๐™Ž๐™ฃ๐™ฎ๐™  - Secret Management: ๐™Ž๐™ค๐™ฅ๐™จ, ๐™‘๐™–๐™ช๐™ก๐™ฉ, ๐™†๐™š๐™ฎ๐™—๐™–๐™จ๐™š, ๐™‘๐™–๐™ช๐™ก๐™ฉ ๐™Ž๐™š๐™˜๐™ง๐™š๐™ฉ๐™จ ๐™Š๐™ฅ๐™š๐™ง๐™–๐™ฉ๐™ค๐™ง, ๐™„๐™ฃ๐™›๐™ž๐™จ๐™ž๐™˜๐™–๐™ก - Dฬฒeฬฒcฬฒlฬฒaฬฒrฬฒaฬฒtฬฒiฬฒvฬฒeฬฒ ฬฒAฬฒPฬฒIฬฒ ฬฒ(ฬฒIฬฒaฬฒaฬฒCฬฒ)ฬฒ: ๐™๐™š๐™ง๐™ง๐™–๐™›๐™ค๐™ง๐™ข, ๐™‹๐™ช๐™ก๐™ช๐™ข๐™ž, ๐˜ผ๐™ฃ๐™จ๐™ž๐™—๐™ก๐™š, ๐™‹๐™ช๐™ฅ๐™ฅ๐™š๐™ฉ, ๐˜พ๐™๐™š๐™›, ๐™†๐™ช๐™—๐™š๐™ซ๐™š๐™ก๐™– - Service Discovery & Service Mesh: ๐™„๐™จ๐™ฉ๐™ž๐™ค, ๐™€๐™ฃ๐™ซ๐™ค๐™ฎ, ๐˜พ๐™ค๐™ฃ๐™จ๐™ช๐™ก, ๐™‡๐™ž๐™ฃ๐™ ๐™š๐™ง๐™™, ๐™•๐™ค๐™ค๐™ ๐™š๐™š๐™ฅ๐™š๐™ง, ๐™ˆ๐™š๐™จ๐™๐™š๐™ง๐™ฎ - SSL: ๐˜พ๐™š๐™ง๐™ฉ ๐™ˆ๐™–๐™ฃ๐™–๐™œ๐™š๐™ง, ๐˜พ๐™š๐™ง๐™ฉ๐™—๐™ค๐™ฉ, ๐™‡๐™š๐™ฉโ€™๐™จ ๐™€๐™ฃ๐™˜๐™ง๐™ฎ๐™ฅ๐™ฉ - Applications Platforms: ๐™†๐™ช๐™—๐™š๐™ง๐™ฃ๐™š๐™ฉ๐™š๐™จ, ๐˜ฟ๐™ค๐™˜๐™ ๐™š๐™ง ๐™Ž๐™ฌ๐™–๐™ง๐™ข, ๐™Š๐™ฅ๐™š๐™ฃ๐™จ๐™๐™ž๐™›๐™ฉ, ๐™๐™–๐™ฃ๐™˜๐™๐™š๐™ง, ๐™Š๐™ฅ๐™š๐™ฃ๐™จ๐™ฉ๐™–๐™˜๐™  - Internal Developer Platforms: ๐˜ฝ๐™–๐™˜๐™ ๐™จ๐™ฉ๐™–๐™œ๐™š, ๐™†๐™ง๐™–๐™ฉ๐™ž๐™ญ, ๐™‹๐™ค๐™ง๐™ฉ - Kubernetes Package Managers: ๐™ƒ๐™š๐™ก๐™ข , ๐™ƒ๐™š๐™ก๐™ข๐™›๐™ž๐™ก๐™š, ๐™ƒ๐™š๐™ก๐™ข๐™จ๐™ข๐™–๐™ฃ โš™๏ธ ๐Ÿ‘. ๐๐ฎ๐ข๐ฅ๐ & ๐ƒ๐ž๐ฉ๐ฅ๐จ๐ฒ๐ฆ๐ž๐ง๐ญ ๐Œ๐จ๐๐ž๐ฅ - Worker Nodes Scaling: ๐™†๐™–๐™ง๐™ฅ๐™š๐™ฃ๐™ฉ๐™š๐™ง, ๐˜พ๐™ก๐™ช๐™จ๐™ฉ๐™š๐™ง ๐˜ผ๐™ช๐™ฉ๐™ค๐™จ๐™˜๐™–๐™ก๐™š๐™ง - Pods Replication Scaling: ๐™ ๐™ฃ๐™–๐™ฉ๐™ž๐™ซ๐™š, ๐™ƒ๐™ค๐™ง๐™ž๐™ฏ๐™ค๐™ฃ๐™ฉ๐™–๐™ก ๐™‹๐™ค๐™™ ๐˜ผ๐™ช๐™ฉ๐™ค๐™จ๐™˜๐™–๐™ก๐™š๐™ง (๐™ƒ๐™‹๐˜ผ), ๐™‘๐™š๐™ง๐™ฉ๐™ž๐™˜๐™–๐™ก ๐™‹๐™ค๐™™ ๐˜ผ๐™ช๐™ฉ๐™ค๐™จ๐™˜๐™–๐™ก๐™š๐™ง (๐™‘๐™‹๐˜ผ) - GฬฒiฬฒtฬฒOฬฒpฬฒsฬฒ: ๐™๐™ก๐™ช๐™ญ๐˜พ๐˜ฟ, ๐˜ผ๐™ง๐™œ๐™ค๐˜พ๐˜ฟ, ๐™๐™–๐™ฃ๐™˜๐™๐™š๐™ง ๐™๐™ก๐™š๐™š๐™ฉ, Argo Rollout , Canary, Blue/Green ๐Ÿ”ญ ๐Ÿ’. ๐‚๐ฅ๐จ๐ฎ๐ ๐Ž๐›๐ฌ๐ž๐ซ๐ฏ๐š๐›๐ข๐ฅ๐ข๐ญ๐ฒ - Monitoring: ๐™‹๐™ง๐™ค๐™ข๐™š๐™ฉ๐™๐™ช๐™š๐™จ, ๐™‚๐™ง๐™–๐™›๐™–๐™ฃ๐™– - Tracing: ๐™๐™š๐™ข๐™ฅ๐™ค, ๐™š๐˜ฝ๐™‹๐™, ๐™Š๐™ฅ๐™š๐™ฃ๐™๐™š๐™ก๐™š๐™ข๐™š๐™ฉ๐™ง๐™ฎ,๐™…๐™–๐™š๐™œ๐™š๐™ง, ๐™•๐™ž๐™ฅ๐™ ๐™ž๐™ฃ - APM: ๐™‰๐™š๐™ฌ ๐™๐™š๐™ก๐™ž๐™˜, ๐™Ž๐™ ๐™ฎ๐™ฌ๐™–๐™ก๐™ ๐™ž๐™ฃ๐™œ, ๐˜ฟ๐™–๐™ฉ๐™–๐™™๐™ค๐™œ - Continuous Profiling & Analysis: ๐™‹๐™–๐™ง๐™ฆ๐™–, ๐™๐™š๐™ฉ๐™ง๐™–๐™œ๐™ค๐™ฃ, ๐™‹๐™ฎ๐™ง๐™ค๐™จ๐™˜๐™ค๐™ฅ๐™š - Logging: ๐™‡๐™ค๐™ ๐™ž, ๐™€๐™‡๐™†, ๐™๐™ก๐™ช๐™š๐™ฃ๐™ฉ๐˜ฝ๐™ž๐™ฉ - Observability Platforms: ๐™‹๐™ž๐™ญ๐™ž๐™š, ๐™๐™ฅ๐™ฉ๐™ง๐™–๐™˜๐™š, ๐™‹๐™ž๐™ฃ๐™ฅ๐™ค๐™ž๐™ฃ๐™ฉ, ๐˜พ๐™ค๐™ง๐™ค๐™ค๐™ฉ - Security Policies: ๐™†๐™š๐™ฎ๐™˜๐™ก๐™ค๐™–๐™˜๐™ , ๐™๐™–๐™ก๐™˜๐™ค, ๐™†๐™ช๐™—๐™š๐˜ผ๐™ง๐™ข๐™ค๐™ง ๐Ÿ” ๐Ÿ“. ๐Ÿ’๐‚'๐ฌ ๐จ๐Ÿ ๐‚๐ฅ๐จ๐ฎ๐ ๐’๐ž๐œ๐ฎ๐ซ๐ข๐ญ๐ฒ - Container Security | Cluster Security | Cloud Security | Container Image Security โ˜๏ธ ๐Ÿ”. ๐‚๐ฅ๐จ๐ฎ๐ ๐๐ฅ๐š๐ญ๐Ÿ๐จ๐ซ๐ฆ๐ฌ - Private: ๐˜ผ๐™’๐™Ž ๐™‚๐™Š๐™‘ ๐˜พ๐™ก๐™ค๐™ช๐™™ , ๐˜ผ๐™ฏ๐™ช๐™ง๐™š ๐™‚๐™Š๐™‘ ๐˜พ๐™ก๐™ค๐™ช๐™™ - Public: ๐˜ผ๐™’๐™Ž , ๐˜ผ๐™ฏ๐™ช๐™ง๐™š, ๐™‚๐˜พ๐™‹, ๐˜ฟ๐™ž๐™œ๐™ž๐™ฉ๐™–๐™ก๐™Š๐™˜๐™š๐™–๐™ฃ, ๐™‘๐™ˆ๐™ฌ๐™–๐™ง๐™š, ๐™‡๐™ž๐™ฃ๐™ค๐™™๐™š, Scaleway, Hertzner, OVH Cloud ๐Ÿ” ๐Ÿ•. ๐€๐ฎ๐ญ๐จ๐ฆ๐š๐ญ๐ข๐จ๐ง - Chaos Engineering: ๐™‡๐™ž๐™ข๐™ž๐™ฉ๐™ช๐™จ, ๐˜พ๐™๐™–๐™ค๐™จ ๐™ˆ๐™š๐™จ๐™ , ๐˜พ๐™๐™–๐™ค๐™จ ๐™†๐™ช๐™—๐™š - MLOps Tools: ๐™ˆ๐™‡๐™๐™ก๐™ค๐™ฌ, ๐™‹๐™š๐™ง๐™›๐™š๐™˜๐™ฉ, ๐™ˆ๐™š๐™ฉ๐™–๐™›๐™ก๐™ค๐™ฌ, ๐™†๐™ช๐™—๐™š๐™๐™ก๐™ค๐™ฌ, ๐™๐™–๐™ฎ, ๐™ˆ๐™š๐™ฉ๐™–๐™๐™ก๐™ค๐™ฌ ๐Ÿ’ก What Clients Say: โญโญโญโญโญ"Firas has become a reliable resource for our company whenever we face any issues with our devops needs from cloud architecture, or cloud native app" โญโญโญโญโญ"Firas went above and beyond in delivering the work! His dedication and skillset are commendable. Highly recommended!!" โญโญโญโญโญ"Firas was like a GINUIS that solved a long persisting issues"

  • DevOps
  • DevOps Engineering
  • Kubernetes
  • Docker
  • Terraform
  • Amazon Web Services
  • Google Cloud Platform
  • Microsoft Azure
  • CI/CD
  • Infrastructure as Code
  • Cloud Engineering
  • Jenkins
  • Ansible
  • Linux
  • Prometheus
  • Grafana
  • Amazon ECS for Kubernetes
  • Python
  • System Administration
  • Network Administration
Yoga W.

Jakarta, Indonesia

$20/hr
4.8
13 jobs

I build AI-powered systems that work while you sleep. Hi ๐Ÿ‘‹ I'm Yoga, an Full Stack Developer & Automation Engineer with 5+ years of experience building production system end-to-end, from infrastructure to frontend. I turn messy manual workflows into clean, automated pipelines: โ†’ AI Agents? Built and deployed. โ†’ LLM pipelines? Designed and optimized. โ†’ RAG systems? Vector DB, embeddings, retrieval. โ†’ Content automation? Scraping, transcription, analysis. โ†’ Workflow automation? n8n, Zapier, Make. โ†’ Infrastructure? Docker, CI/CD, VPS, Proxmox, Monitoring with Grafana. โ†’ Networking? VPN, Firewall, Routing & Switching, RADIUS. What makes me different: I don't just build the AI part. I build the whole system around it, the backend, the infra, the deployment, the monitoring. You get one engineer who handles architecture to production. I've operated and maintained 150+ services accessed globally. I know what it takes to keep systems running reliably at scale. Startups building 0โ†’1? I'm your guy. Existing systems that need AI integration? Let's talk. If it can be automated, I'll automate it. If it needs to scale, I'll make it scale. Let's build something that runs itself ๐Ÿ˜Š Keywords: AI Agents, LLM, RAG, Vector Databases, Claude, OpenAI, Langchain, Python, Rust, Node.js, Go, React, Next.js, FastAPI, n8n, Zapier, Make, Docker, Kubernetes, Terraform, CI/CD, PostgreSQL, WireGuard, MikroTik, Linux, DevOps, Infrastructure, Automation, Content Pipeline, Whisper, Web Scraping, API Integration, Routing & Switching.

  • Network Administration
  • Rust
  • DevOps Engineering
  • Infrastructure Management
  • Cloud Computing
  • Cloud Architecture
  • n8n
  • AI Development
  • AI Agent Development
  • AI Chatbot
  • AI Consulting
  • AI Platform
  • AI Implementation
  • AI App Development
  • AI Mobile App Development
  • AI Model Integration
  • AI Model Development
  • AI Bot
  • AI Builder

How it works

Post a job for freePost a job

Tell us what you need. Create your own job post or generate one with AI then filter talent matches.

Hire top talent fast

Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.

Collaborate easily

Use Upwork to chat or video call, share files, and track project progress right from the app.

Payment simplified

Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.

Don't just take our word for it

What does a Distributed systems engineer do?

A distributed systems engineer builds software that runs across many computers at once, treating them as a single coordinated unit. This work focuses on keeping data consistent and services available even when individual servers fail or network connections drop. You design architectures that scale horizontally by adding more machines rather than upgrading a single box. Your code handles partial failures gracefully so users experience no interruption during outages.

  • Design system architectures that distribute workloads across multiple nodes to prevent bottlenecks and single points of failure. You define how services communicate, manage state, and handle retries when remote calls time out. This includes selecting consensus algorithms and data partitioning strategies that match your consistency requirements.
  • Implement observability using tools like OpenTelemetry to collect traces, metrics, and logs from every service in the cluster. You configure the OpenTelemetry Collector to export this telemetry to backends for analysis. This visibility lets you spot latency spikes and error rates before they impact customers.
  • Troubleshoot complex issues that span multiple layers of the stack, from application logic to network configuration. You perform root cause analysis after incidents to identify why a failure occurred and how to prevent it next time. This work involves reading distributed traces to pinpoint where a request stalled or failed.
  • Write runbooks and automate operational tasks to reduce manual toil during routine maintenance or emergency responses. You build tooling that simplifies system adoption for other developers and reduces the cognitive load of managing distributed state. This includes creating self-service features that handle credential distribution or configuration updates safely.
  • Participate in on-call rotations to respond to production incidents and restore service availability quickly. You lead post-incident reviews to document lessons learned and assign corrective actions to specific team members. This process ensures that each outage results in concrete improvements to system resilience.

How to hire a Distributed systems engineer on Upwork

Step 1: Post a job

Define the specific distributed computing challenges your infrastructure faces. The Job Post Generator powered by Umaโ„ข, Upwork's Mindful AI drafts a complete post from a few sentences describing your needs. You can write a new post, update a saved draft, or reuse an existing post to start hiring.

  • Specify requirements for designing scalable architectures that coordinate across multiple hosts and handle failure conditions gracefully.
  • List necessary experience with observability stacks like OpenTelemetry for collecting traces, metrics, and logs to monitor system health.
  • Detail expectations for owning operational excellence through runbook development and proactive risk identification mechanisms.

Step 2: Evaluate candidates

Look for portfolios demonstrating work on tier-0 critical capabilities such as credential distribution platforms. Uma runs instant video interviews and builds shortlists with side-by-side comparisons to help you assess technical depth quickly.

  • Verify experience troubleshooting complex distributed-system issues across multiple layers and performing root cause analysis after incidents.
  • Check for contributions to system design decisions that prioritize long-term maintainability and secure operation in production environments.
  • Review examples of operational automation that reduce toil and improve mean-time-to-resolution for ongoing service reliability.

Step 3: Interview your top choices

Discuss how candidates approach simplifying system behavior and adoption via tooling or features. Schedule and conduct interviews within Upwork Messages to receive an immediate transcript and summary after each conversation.

  • Ask about their process for working backwards from customer needs to build productionized tooling that reduces operational burden.
  • Explore their experience with Kubernetes observability practices for managing cluster application metrics, logs, and traces effectively.
  • Evaluate their participation in on-call rotations and incident response workflows to gauge their readiness for real-time troubleshooting.

Step 4: Agree on scope and begin work

Set clear milestones for delivering operational automation and monitoring improvements. Use Upwork Messages and the contract workroom for communication and project management, plus identity verification, payment protection, hourly tracking, and project funds for security.

  • Define deliverables for system architecture decisions that address specific scalability requirements and secure data handling protocols.
  • Establish expectations for submitting root cause analyses and corrective action plans following any production incidents during the contract.
  • Agree on metrics for success such as reduced operational toil and improved system reliability scores through proactive maintenance.

Upwork is not affiliated with and does not sponsor or endorse any of the tools or services discussed in this article. These tools and services are provided only as potential options, and each reader and company should take the time needed to adequately analyze and determine the tools or services that would best fit their specific needs and situation.

The rates and information provided in this article are based on current data and industry sources available at the time of publication. Freelance rates can vary depending on factors such as experience, location, project scope, and market conditions. Readers are encouraged to conduct their own research to confirm current rates and trends, as this information may change over time.

How much does hiring a Distributed systems engineer cost?

$800-$2,500 per project is a typical range for focused Distributed systems engineer work. Final pricing depends on scope, technical complexity, required integrations, source-material quality, revision needs, and the freelancer's experience level.

System architecture review

$800-$1,500/project

Mid-level
  • Identified scalability bottlenecks and reliability risks
  • Prioritized list of architectural improvements
  • Phased implementation plan for system upgrades

Observability setup

$1,500-$3,000/project

Mid-level to senior-level
  • Integrated OpenTelemetry Collector for metrics and traces
  • Visualized key performance indicators and error rates
  • Defined thresholds for proactive incident detection

Incident response automation

$3,000-$5,500/project

Senior-level
  • Documented procedures for common failure scenarios
  • Automated remediation steps for known issues
  • Standardized format for root cause analysis reports

Scalability optimization

$5,500-$9,000/project

Senior-level
  • Optimized distributed components for higher throughput
  • Validated system behavior under peak traffic conditions
  • Measured improvements in latency and resource usage

Custom distributed platform build

$9,000-$15,000/project

Expert-level
  • Built fault-tolerant microservices with secure communication
  • Configured Kubernetes clusters for automated scaling
  • Comprehensive guide for system maintenance and expansion

Frequently asked questions

Is hiring a Distributed systems engineer worth it?

For most businesses, yes: hiring a Distributed systems engineer is worthwhile. These specialists build architectures that scale across multiple servers while maintaining reliability during outages. They reduce operational toil by automating monitoring and incident response workflows.

How do I evaluate Distributed systems engineer candidates?

Evaluate candidates by reviewing their approach to root cause analysis and system observability. Look for specific examples where they used tools like OpenTelemetry to trace errors across services and wrote runbooks that reduced mean-time-to-resolution.

What deliverables does a Distributed systems engineer produce?

A Distributed systems engineer produces system architecture designs, operational runbooks, and automated monitoring configurations. They also generate root cause analysis reports after incidents to prevent future failures.

Which tools do Distributed systems engineers use?

Distributed systems engineers use observability stacks such as OpenTelemetry Collector to gather metrics, logs, and traces. They also apply Kubernetes observability practices to monitor cluster and application performance.