LLM Infrastructure Specialist (Local / Self-Hosted Deployment)
Worldwide
About the role We're building an AI platform for commercial real estate workflows — document ingestion, AI-assisted extraction, and human-in-the-loop review — where security, data privacy, and auditability are first-class concerns. We're looking for a specialist who can deploy and operate large language models on our own infrastructure, so sensitive documents never have to leave our environment. You'll own the local LLM stack end to end: selecting models, standing up inference servers, optimizing them for our hardware, and making them reliable enough for production use. What you'll do Deploy open-weight LLMs (e.g. Llama, Mistral, Qwen, DeepSeek) on local / on-prem / private-cloud GPU infrastructure Set up and tune inference servers (vLLM, TGI, Ollama, llama.cpp, or similar) for throughput and latency Apply quantization and optimization techniques (GGUF, AWQ, GPTQ, etc.) to fit models to available hardware Expose models behind stable, OpenAI-compatible APIs for our application services Configure GPU environments (CUDA drivers, containerization, orchestration) Benchmark models for quality, speed, and cost; recommend the right model for each use case Establish monitoring, logging, and scaling for inference workloads Document the setup so the team can maintain and reproduce it What we're looking for Proven experience deploying LLMs locally or in a self-hosted/private environment (not just calling hosted APIs) Hands-on with at least one inference framework (vLLM, TGI, Ollama, llama.cpp, LM Studio, or equivalent) Solid understanding of GPU hardware, VRAM constraints, and model quantization trade-offs Comfortable with Linux, Docker/containers, and Python Able to reason about model selection: size vs. quality vs. speed vs. cost Nice to have Experience with fine-tuning, LoRA/PEFT, or RAG pipelines Kubernetes / GPU orchestration at scale Background in regulated or security-sensitive environments (data privacy, audit requirements) Familiarity with Azure / cloud GPU instances
- Less than 30 hrs/weekHourly
- 3-6 monthsDuration
- ExpertExperience Level
$35.00
-
$45.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:20 to 50
- Last viewed by client:yesterday
- Interviewing:7
- Invites sent:0
- Unanswered invites:0
About the client
- United StatesBrooklyn12:55 PM
- $280K total spent27 hires, 7 active
- 9,910 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by