LLM Infrastructure Specialist (Local / Self-Hosted Deployment)

Posted yesterday

Worldwide

Summary

About the role We're building an AI platform for commercial real estate workflows — document ingestion, AI-assisted extraction, and human-in-the-loop review — where security, data privacy, and auditability are first-class concerns. We're looking for a specialist who can deploy and operate large language models on our own infrastructure, so sensitive documents never have to leave our environment. You'll own the local LLM stack end to end: selecting models, standing up inference servers, optimizing them for our hardware, and making them reliable enough for production use. What you'll do Deploy open-weight LLMs (e.g. Llama, Mistral, Qwen, DeepSeek) on local / on-prem / private-cloud GPU infrastructure Set up and tune inference servers (vLLM, TGI, Ollama, llama.cpp, or similar) for throughput and latency Apply quantization and optimization techniques (GGUF, AWQ, GPTQ, etc.) to fit models to available hardware Expose models behind stable, OpenAI-compatible APIs for our application services Configure GPU environments (CUDA drivers, containerization, orchestration) Benchmark models for quality, speed, and cost; recommend the right model for each use case Establish monitoring, logging, and scaling for inference workloads Document the setup so the team can maintain and reproduce it What we're looking for Proven experience deploying LLMs locally or in a self-hosted/private environment (not just calling hosted APIs) Hands-on with at least one inference framework (vLLM, TGI, Ollama, llama.cpp, LM Studio, or equivalent) Solid understanding of GPU hardware, VRAM constraints, and model quantization trade-offs Comfortable with Linux, Docker/containers, and Python Able to reason about model selection: size vs. quality vs. speed vs. cost Nice to have Experience with fine-tuning, LoRA/PEFT, or RAG pipelines Kubernetes / GPU orchestration at scale Background in regulated or security-sensitive environments (data privacy, audit requirements) Familiarity with Azure / cloud GPU instances

  • Less than 30 hrs/week
    Hourly
  • 3-6 months
    Duration
  • Expert
    Experience Level
  • $35.00

    -

    $45.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
LLM Prompt Engineering
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:yesterday
  • Interviewing:
    7
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Sep 9, 2008
  • United States
    Brooklyn3:04 PM
  • $280K total spent
    27 hires, 7 active
  • 9,910 hours

Explore similar jobs on Upwork

Apache and PHP-FPM Configuration ExpertHourly‐ Posted 1 month ago
Ubuntu
Apache HTTP Server
Docker
CI/CD
Git
GitLab
Linux System Administration
DevOps
Performance Testing

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo