Senior AI/ML Infrastructure Engineer

Posted yesterday

Worldwide

Summary

We are looking for a Senior AI/ML Infrastructure Engineer to design, deploy, and integrate a custom, privacy-first AI assistant ("Resi") directly into our urban planning and climate tech platform, Planner360. Our infrastructure runs entirely on our own dedicated, bare-metal physical servers with zero cloud hosting dependencies. The goal is to build an air-gapped Retrieval-Augmented Generation (RAG) and local inference pipeline capable of processing complex, multi-format urban datasets—including unstructured files (PDFs, Word documents), structured data (Excel, SQL/NoSQL databases), and complex spatial/GIS layers (building footprints, infrastructure networks, map data). The assistant must integrate cleanly via a custom backend API with Planner360, supporting structured inputs and output formatting that aligns with user requirements. Infrastructure Profile Your solution must be engineered and optimized specifically for our physical hardware cluster: Server Count: 3x Physical Bare-Metal Servers Processors: AMD Ryzen 5900X (per server) Memory: 128GB RAM per server (Total system capacity: 384GB RAM) Storage: 2x NVMe drives per server (ultra-fast storage for vector DB and caching) Note: Because this is a CPU/RAM-optimized environment (rather than enterprise NVIDIA GPUs), the architecture must utilize efficient CPU runtimes (such as Ollama with GGUF/quantized weights) and intelligently distribute workloads across our 3 nodes. Key Responsibilities Local Inference Deployment: Configure, optimize, and run local model inference engines (Ollama / GGUF runtimes) across our 3-node physical server cluster using open-source weights (e.g., DeepSeek, Qwen). Advanced RAG & Data Ingestion Pipeline: Build robust ingestion parsers for multi-format urban datasets, including text documents, tabular databases, and geospatial/GIS formats (GeoJSON, shapefiles, building footprints). Vector Indexing & Storage: Set up and maintain a local vector database (Qdrant or Milvus) leveraging our fast NVMe drives for high-speed semantic search across city-scale data. Planner360 API Integration: Develop a secure, high-throughput backend API (using FastAPI) to seamlessly connect the local AI engine with our existing Planner360 application interface. Structured Output Mapping: Implement response-formatting logic so the LLM generates precise structured output blocks, code schemas, or JSON snippets that the Planner360 backend can programmatically compile back into user files (Excel, GIS layers, reports). Tech Stack Experience Required Inference Runtimes & Models: Ollama, GGUF/quantized model workflows, Hugging Face ecosystem. Vector DBs & Orchestration: Qdrant, Milvus, LangChain, or LlamaIndex. Backend & APIs: Python, FastAPI, Docker, RESTful APIs, WebSockets. Data & Spatial Formats: Experience with Pandas, Excel processing, and geospatial/GIS data handling (GeoPandas, GeoJSON). Infrastructure: Linux (Ubuntu Server), bare-metal CPU/RAM resource optimization, Docker Compose, containerized network isolation across multiple nodes.

  • $499.00

    Fixed-price
  • Expert
    Experience Level
  • Remote Job
  • Complex project
    Project Type
Skills and Expertise
Mandatory skills
Embedded System
Microcontroller Programming
Nice-to-have skills
Reverse Engineering
Embedded C
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:7 hours ago
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Dec 1, 2020
  • United States
    Basking Ridge11:12 AM
  • $4.5K total spent
    8 hires, 4 active
  • 20 hours
  • Individual client

Explore similar jobs on Upwork

Gen AI Developer (Contract)Fixed-price‐ Posted 2 months ago
AI Agent Development
Python
JavaScript
API
Node.js
Deep Learning
React
PostgreSQL
Football Prediction Tool DeveloperHourly‐ Posted 4 weeks ago
JavaScript
PHP
HTML5
jQuery

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo