You will get Your LLM, 40-70% cheaper

Let a pro handle the details

Buy Other AI & Machine Learning services from Alexander, priced and ready to go.

Let a pro handle the details

Buy Other AI & Machine Learning services from Alexander, priced and ready to go.

Project details

You are almost certainly paying more than you need to for every token your product generates. Most teams pick a model, get it serving, and move on to features. The serving layer is then never revisited, and the bill grows quietly with usage.

In one week I find out exactly where that money goes and what to do about it. I benchmark your current stack, profile GPU utilisation, and test the levers that matter: quantisation, continuous batching, KV-cache reuse, engine choice and routing cheaper models at the requests that do not need your best one. Everything is measured on your workload, against your quality bar. Nothing is recommended on theory.

You end the week with a ranked list of changes, each one priced in dollars saved per month, with the effort and risk stated plainly. You also keep every benchmark, configuration and number I produce, so your engineers can act on it with or without me.

For two years I co-built a GPU-cloud LLM serving platform as one of its two engineers, and inference optimisation was my job. Reductions of 40 to 70 percent are normal, at equal or better latency.
AI Development Type
Software Maintenance
AI Tools
MLflow, NVIDIA AI Platform, PyTorch, TensorFlow
AI Development Language
Python

What's included $4,500

These options are included with the project scope.

$4,500
  • Delivery Time 7 days
  • Number of Revisions 1
    • AI Model Integration

Frequently asked questions

Alexander P.Status: Offline

About Alexander

Alexander P.Status: Offline
LLM Inference Optimisation | On-Prem & Air-Gapped | Cost Cuts
Poznan, Poland - 12:04 am local time
I cut LLM serving costs by 40 to 70% and make inference fast, using vLLM, SGLang, NVIDIA Triton, quantisation, KV-cache reuse and batching on H100s.

For two years I co-built a GPU-cloud LLM-serving platform as one of its two engineers, owning inference optimisation end to end: engine selection and tuning, quantisation (AWQ/GPTQ/FP8), speculative decoding, continuous batching, GPU profiling with Nsight.

I also deploy private LLM systems that never touch an external API. For a European photonics manufacturer I delivered a fully on-premise LLM + RAG system on a single node, inside a tight memory and latency budget, with complete data sovereignty and a named CEO testimonial on my website. If your data cannot leave your building, that is the work I do.

What I do for clients:
✔ Inference cost and latency audits, with the top 10 savings actions priced by $-impact
✔ Serving-stack builds and tuning (vLLM / SGLang / Triton / TensorRT-LLM)
✔ On-prem and air-gapped LLM + RAG deployment (EU-based; short on-site visits possible for installs)
✔ Model routing, caching and quantisation without quality loss

PhD physicist. Two granted US patents in computer vision. NASA algorithm challenge placements (2nd and 3rd). 12+ years shipping production AI. EU time zones.

Steps for completing your project

After purchasing the project, send requirements so Alexander can start the project.

Delivery time starts when Alexander receives requirements from you.

Alexander works on your project following the steps below.

Revisions may occur after the delivery date.

Day 1: I look at what you run today

Your models, serving stack, traffic pattern and current spend. Short kickoff call if you want one.

Day 2-3: I measure it

I benchmark your stack and profile GPU use, so we know where the money actually goes instead of guessing.

Review the work, release payment, and leave feedback to Alexander.