Principal AI Infrastructure & HPC Engineer
Only freelancers located in the U.S. may apply.U.S. located freelancers only
Principal AI Infrastructure & HPC Engineer We are a PE-backed, AI-native technology services company building a new AI Infrastructure & HPC practice across AWS and Microsoft Azure. We're looking for a deeply technical engineer to help build the practice from the ground up. This is not a traditional cloud or DevOps role. The focus is large-scale GPU infrastructure, distributed AI workloads, high-performance networking, and getting expensive compute environments to perform at their potential. What You'll Work On GPU cluster benchmarking, performance tuning and optimization NCCL benchmarking and tuning AWS EFA and Azure GPU/HPC infrastructure InfiniBand, RDMA, RoCE and GPUDirect RDMA CUDA, NVLink/NVSwitch and NVIDIA GPU environments Distributed AI training and inference optimization Linux, kernel, driver and systems-level performance Kubernetes/EKS/AKS and Slurm-based GPU environments GPU cloud / neocloud infrastructure What We're Looking For We want someone with deep hands-on expertise in GPU/HPC systems who can benchmark an environment, identify where performance is being lost, and fix it. Experience with NCCL, CUDA, InfiniBand/RDMA, distributed training, Linux performance engineering, and large multi-node GPU clusters is particularly relevant. Deep AWS and/or Microsoft Azure experience is a major plus, particularly experience designing or optimizing GPU/HPC workloads using EFA, EC2 accelerated computing, Azure GPU infrastructure, EKS/AKS, and high-performance networking. You don't need to check every box. Depth in this domain matters more than breadth. More Than a Project We're building a practice around this capability. The right person can play an important role in defining our technical offerings, developing repeatable optimization methodologies, working directly with AWS and Microsoft, and helping us build the engineering team as the practice grows. If you've worked deep in GPU infrastructure, HPC, distributed systems, or high-performance networking, we'd like to talk.
- More than 30 hrs/weekHourly
- 6+ monthsDuration
- ExpertExperience Level
$60.00
-
$128.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:20 to 50
- Last viewed by client:11 hours ago
- Interviewing:21
- Invites sent:10
- Unanswered invites:2
About the client
- USACosta Mesa 4:35 AM
- $7.9K total spent9 hires, 5 active
- 78 hours
- Tech & ITSmall company (2-9 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by