Senior Python/AI Performance Engineer

Posted yesterday

Worldwide

Summary

We are looking for an experienced Python / AI / LLM performance engineer to investigate a production-like PDF processing platform that uses Ollama, Docker, NVIDIA GPU, Python, SQL Server, and a Medallion-style processing architecture. The system is functional and can successfully process 100+ PDF files, and the server remains stable during load testing. However, the overall processing time is significantly slower than expected. An important observation is that calling Ollama directly is fast and stable, while processing the same files through our complete application/Medallion pipeline introduces significant additional latency. We need a senior engineer who can go beyond surface-level debugging and determine why this performance difference exists, where the bottleneck is, and what should be changed to make the platform production-ready and scalable. Main Objective The objective of this assignment is to: Analyse the existing codebase and architecture. Analyse the server, Docker containers, GPU/CPU/RAM and Ollama configuration. Identify the actual performance bottleneck(s). Compare direct Ollama processing against the complete application pipeline. Determine why the Medallion pipeline introduces additional latency. Identify scalability and concurrency limitations. Provide concrete technical solutions and, where practical, implement the required fixes. We are not looking for someone to simply rewrite the application. We want someone who can first understand the existing system, measure it properly, identify the root cause and then make targeted improvements. Current Technology Stack The environment includes: Python Ollama Local LLM NVIDIA GPU Docker / Docker Compose SQL Server PDF processing / OCR Medallion-style file processing architecture REST/API services Linux server The exact implementation details and source code will be provided to the selected freelancer. Key Problem We currently observe behaviour similar to: Direct Ollama PDF + prompt → Ollama → result This is relatively fast and stable. Application pipeline PDF → file detection → Medallion processing → OCR/extraction → configuration/schema processing → Ollama → post-processing → database/filesystem → result This takes significantly longer. We need to understand exactly which part of the second flow is responsible for the additional processing time. Scope of Work: Performance Profiling Instrument/profile the application where necessary and establish measurable timings for each major stage. For example: File discovery File copying/movement PDF reading OCR/extraction Prompt preparation Ollama request Ollama response JSON parsing Database operations Output generation Other Medallion processing stages We want to know where the time is actually being spent, rather than relying on assumptions. Server & Infrastructure Investigation Investigate the server environment, including where relevant: NVIDIA GPU utilization GPU memory / VRAM CPU utilization RAM utilization Disk I/O Network/container communication Docker resource limits Docker container behaviour Ollama configuration Model loading/unloading behaviour Concurrent requests Container restart behaviour Process/resource contention SQL Server performance File-system performance We have previously experienced server/container instability, so we also want to determine whether there are infrastructure-level issues contributing to the problem.

  • $400.00

    Fixed-price
  • Expert
    Experience Level
  • Remote Job
  • One-time project
    Project Type
Skills and Expertise
Mandatory skills
Gen AI Development Skills
Activity on this job
  • Proposals:15 to 20
  • Last viewed by client:yesterday
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since May 11, 2026
  • NLD
    Vleuten5:33 AM

Explore similar jobs on Upwork

I need LLMs to recommend my businessFixed-price‐ Posted 2 days ago
Search Engine Indexing Optimization
Search Engine Ranking
AI Developer for Real-Time SolutionsFixed-price‐ Posted 2 weeks ago
Artificial Intelligence
Computer Vision
Edge AI
Model Optimization

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo