You will get Your LLM, 40-70% cheaper


Project details
You are almost certainly paying more than you need to for every token your product generates. Most teams pick a model, get it serving, and move on to features. The serving layer is then never revisited, and the bill grows quietly with usage.
In one week I find out exactly where that money goes and what to do about it. I benchmark your current stack, profile GPU utilisation, and test the levers that matter: quantisation, continuous batching, KV-cache reuse, engine choice and routing cheaper models at the requests that do not need your best one. Everything is measured on your workload, against your quality bar. Nothing is recommended on theory.
You end the week with a ranked list of changes, each one priced in dollars saved per month, with the effort and risk stated plainly. You also keep every benchmark, configuration and number I produce, so your engineers can act on it with or without me.
For two years I co-built a GPU-cloud LLM serving platform as one of its two engineers, and inference optimisation was my job. Reductions of 40 to 70 percent are normal, at equal or better latency.
In one week I find out exactly where that money goes and what to do about it. I benchmark your current stack, profile GPU utilisation, and test the levers that matter: quantisation, continuous batching, KV-cache reuse, engine choice and routing cheaper models at the requests that do not need your best one. Everything is measured on your workload, against your quality bar. Nothing is recommended on theory.
You end the week with a ranked list of changes, each one priced in dollars saved per month, with the effort and risk stated plainly. You also keep every benchmark, configuration and number I produce, so your engineers can act on it with or without me.
For two years I co-built a GPU-cloud LLM serving platform as one of its two engineers, and inference optimisation was my job. Reductions of 40 to 70 percent are normal, at equal or better latency.
AI Development Type
Software MaintenanceAI Tools
MLflow, NVIDIA AI Platform, PyTorch, TensorFlowAI Development Language
PythonWhat's included $4,500
These options are included with the project scope.
$4,500
- Delivery Time 7 days
- Number of Revisions 1
- AI Model Integration
Frequently asked questions
About Alexander
LLM Inference Optimisation | On-Prem & Air-Gapped | Cost Cuts
Poznan, Poland - 12:04 am local time
For two years I co-built a GPU-cloud LLM-serving platform as one of its two engineers, owning inference optimisation end to end: engine selection and tuning, quantisation (AWQ/GPTQ/FP8), speculative decoding, continuous batching, GPU profiling with Nsight.
I also deploy private LLM systems that never touch an external API. For a European photonics manufacturer I delivered a fully on-premise LLM + RAG system on a single node, inside a tight memory and latency budget, with complete data sovereignty and a named CEO testimonial on my website. If your data cannot leave your building, that is the work I do.
What I do for clients:
✔ Inference cost and latency audits, with the top 10 savings actions priced by $-impact
✔ Serving-stack builds and tuning (vLLM / SGLang / Triton / TensorRT-LLM)
✔ On-prem and air-gapped LLM + RAG deployment (EU-based; short on-site visits possible for installs)
✔ Model routing, caching and quantisation without quality loss
PhD physicist. Two granted US patents in computer vision. NASA algorithm challenge placements (2nd and 3rd). 12+ years shipping production AI. EU time zones.
Steps for completing your project
After purchasing the project, send requirements so Alexander can start the project.
Delivery time starts when Alexander receives requirements from you.
Alexander works on your project following the steps below.
Revisions may occur after the delivery date.
Day 1: I look at what you run today
Your models, serving stack, traffic pattern and current spend. Short kickoff call if you want one.
Day 2-3: I measure it
I benchmark your stack and profile GPU use, so we know where the money actually goes instead of guessing.