You will get a Production ready GPU Kubernetes cluster for AI and LLMs
Top Rated

Top Rated

Project details
Get Your Kubernetes Cluster Running in 60 Minutes 🚀
Deploy production-ready Kubernetes clusters quickly using battle-tested automation scripts. Whether you need CPU or GPU workloads, I'll set up your cluster with essential management tools:
✓ CPU Clusters: Rancher dashboard for easy management
✓ GPU Clusters: Nvidia-device-plugin + Rancher + DCGM + Grafana for monitoring
✓ Basic monitoring
Three tiers available:
• Starter: 3-node CPU-K8s cluster
• Standard: 3-node GPU-K8s cluster
• Advanced: Multi-node cluster + Infrastructure as Code scripts Included
Expertise backed by successful deployments of Kubernetes on Pakistan's largest HPC at SPCAI Data-center .
Video demo available showing automated cluster creation.
Skip weeks of configuration headaches. Get your cluster running today.
Note: Requires your own VMs/hardware. GPU support needs compatible hardware.
Let's get your cluster up and running! 🔧
Deploy production-ready Kubernetes clusters quickly using battle-tested automation scripts. Whether you need CPU or GPU workloads, I'll set up your cluster with essential management tools:
✓ CPU Clusters: Rancher dashboard for easy management
✓ GPU Clusters: Nvidia-device-plugin + Rancher + DCGM + Grafana for monitoring
✓ Basic monitoring
Three tiers available:
• Starter: 3-node CPU-K8s cluster
• Standard: 3-node GPU-K8s cluster
• Advanced: Multi-node cluster + Infrastructure as Code scripts Included
Expertise backed by successful deployments of Kubernetes on Pakistan's largest HPC at SPCAI Data-center .
Video demo available showing automated cluster creation.
Skip weeks of configuration headaches. Get your cluster running today.
Note: Requires your own VMs/hardware. GPU support needs compatible hardware.
Let's get your cluster up and running! 🔧
Machine Learning Tools
Kubeflow, NVIDIA AI Platform, Python, PyTorch, TensorFlow, Vertex AIWhat's included
| Service Tiers |
Starter
$499
|
Standard
$1,199
|
Advanced
$2,499
|
|---|---|---|---|
| Delivery Time | 3 days | 7 days | 14 days |
Number of Revisions | 1 | 2 | 3 |
Model Validation/Testing | - | - | - |
Model Documentation | - | - | - |
Data Source Connectivity | - | - | - |
Source Code | - | - | - |
Optional add-ons
You can add these on the next page.
vLLM deployment with your model
+$400
Prometheus + Grafana monitoring stack
+$250
30 days post-delivery support
+$300Frequently asked questions
12 reviews
(12)
(0)
(0)
(0)
(0)
This project doesn't have any reviews.
AA
Ahmed A.
Aug 15, 2026
Senior Full-Stack Engineer (Python/FastAPI + React) — Enterprise GRC SaaS
Great quality work. Goes above and beyond and delivers in tight deadlines. Very detail oriented and amazing technical expertise!
AA
Ahmed A.
Jul 8, 2026
DevOps & Observability Engineer
Quality work. Very organized and great technical adjustments.
EG
Elena G.
Jul 1, 2026
Technical Audit for Luxury Marketplace Website (Code, Security, Ownership & Infrastructure Review)
Choudhry was excellent to work with. He conducted a thorough security audit, identified critical vulnerabilities, and clearly explained every issue along with practical solutions. He was highly responsive, detail-oriented, and proactive throughout the project, often going beyond the original scope to improve our platform’s security. I highly recommend him to anyone looking for a knowledgeable, reliable, and professional cybersecurity expert. I would definitely hire him again.
AS
Abhishek S.
Jun 24, 2026
Full-Stack AI Engineer Needed for Question Paper Generator
Choudhry built a backend document intelligence pipeline for us that takes educational PDFs (NCERT-style Maths and Physics, both typed and scanned) and produces a single source-of-truth JSON: page-wise, block-wise, with text, LaTeX-encoded equations, and cropped image assets. Same JSON drives a web-ready output and a basic printable PDF, so reviewers can verify extraction quality against the original.
Scope was tight (about three weeks, fixed price), and he delivered everything on the list: full source code, API docs, sample JSONs across multiple PDFs, extracted image assets with captions and references preserved, reconstructed PDF and HTML samples, and a README that actually explains the pipeline end to end. Multi-column reading order was a known risk area and he handled it well.
He is technically sharp, picked the right tooling combination for our case (Docling for general layout, equation-specific OCR for math) instead of forcing one tool to do everything. He is also a clear communicator. When my spec had ambiguity, he flagged it instead of silently making assumptions, which saved us rework.
Solid hire for anyone doing document AI, OCR, or structured extraction work.
Scope was tight (about three weeks, fixed price), and he delivered everything on the list: full source code, API docs, sample JSONs across multiple PDFs, extracted image assets with captions and references preserved, reconstructed PDF and HTML samples, and a README that actually explains the pipeline end to end. Multi-column reading order was a known risk area and he handled it well.
He is technically sharp, picked the right tooling combination for our case (Docling for general layout, equation-specific OCR for math) instead of forcing one tool to do everything. He is also a clear communicator. When my spec had ambiguity, he flagged it instead of silently making assumptions, which saved us rework.
Solid hire for anyone doing document AI, OCR, or structured extraction work.
RA
Richbourgs A.
Jun 23, 2026
N8n hooked to Mattermost for a tekmetric webhook
Absolutley wonderful experience very knowledgeable and helpful
About Choudhry
On-Prem GPU & AI Infrastructure Engineer | MLOps, LLMs & Kubernetes |
100%
Job Success
Haripur, Pakistan - 1:00 pm local time
🔹100% Job Success | Top Rated
🔹60+ LLM deployments
🔹30+ GPU Kubernetes clusters
🔹30+ AI servers designed and deployed
🔹20+ NVIDIA GPUs deployed in a single cluster.
I wrote the IaC for GPU cluster orchestration on KAUST Shaheen III, one of the world's leading supercomputers. That is the level of infrastructure I work with, both on prem and in the cloud.
Most of my work is around self hosted and private AI, GPU infrastructure, LLM deployment and making AI systems reliable enough for real production use.
🔹I work hands on with:
🔹Self hosted LLMs and model serving: vLLM, NVIDIA Triton, TensorRT, SGLang, Ollama, Qwen, Llama, DeepSeek and other open source models
🔹GPU inference optimization: FP8, INT8, FlashAttention, tensor parallelism, batching, quantization and GPU performance tuning
🔹GPU Kubernetes: NVIDIA GPU Operator, Proxmox, Rancher, K3s, OpenShift, Calico, MetalLB, HAProxy and multi GPU scheduling
🔹Cloud AI infrastructure: AWS, GCP, Azure, GKE, EKS, AKS, Vertex AI, SageMaker and GPU compute
🔹AI agents and applications: LangGraph, LangChain, CrewAI, Claude Agent SDK, MCP, RAG, document intelligence and voice agents
🔹MLOps and platform engineering: Terraform, Ansible, Pulumi, Helm, Kubeflow, Ray, GitHub Actions, GitLab CI, ArgoCD, Flux, Prometheus, Grafana and OpenTelemetry
🔹A few results from projects I've worked on:
Reduced LLM inference latency from 3.2 seconds to 800ms through vLLM and serving optimization
Reduced a client's monthly cloud bill from $10K to $4K
Served LLM workloads for 1,000+ concurrent users on bare metal Kubernetes with zero unplanned downtime for 12 months
Built and deployed a Docling + GraphRAG pipeline on GCP for commercial real estate lease analysis
🔹Clients have described my work as:
"single handedly built our Kubernetes infrastructure for our GPU setup"
"a full stack AI guy"
"gets concepts quickly and is an incredibly hard worker"
"when he says he'll do it by a deadline, he does"
"approaches the project as if it were his own"
"incredibly efficient"
I'm currently CTO / AI Infrastructure Lead at Peregrine Ventures AI, where I work hands on with production AI systems. I also run my own dual RTX Pro 6000 Blackwell server with EPYC, Proxmox and GPU passthrough, so on prem GPU infrastructure isn't something I just know from the cloud. It's something I work with every day.
If you're building, deploying, self hosting, scaling or optimizing an AI or LLM system, feel free to reach out. I'm happy to look at the architecture, figure out what is actually needed and help get it into production.
Steps for completing your project
After purchasing the project, send requirements so Choudhry can start the project.
Delivery time starts when Choudhry receives requirements from you.
Choudhry works on your project following the steps below.
Revisions may occur after the delivery date.
Review your hardware, workloads and access, confirm the final plan
Install and configure Kubernetes with the NVIDIA GPU Operator