You will get audit your LLM inference stack for cost and latency


Project details
What you get:
• Baseline cost and latency report for your current inference stack
• Actionable configuration changes: continuous batching, KV-cache/prefix caching, quantization, multi-GPU or scale-to-zero where relevant
• Before/after benchmark plan you can re-run after changes
• Optional implementation of high-impact fixes within the engagement window
Best for teams with rising GPU bills, p95 latency SLOs, or a PoC that needs production economics.
Stack familiarity: vLLM, TGI, Triton, Kubernetes/GKE, AWS Batch/serverless inference paths, OpenAI-compatible gateways (LiteLLM), observability for LLM spans.
How we work:
1) Kickoff + access to metrics/staging (read-only preferred)
2) Audit + written findings
3) Prioritized recommendations and optional hands-on tuning
4) Handoff notes your team can operate
Not a ChatGPT wrapper consult — this is production inference engineering.
• Baseline cost and latency report for your current inference stack
• Actionable configuration changes: continuous batching, KV-cache/prefix caching, quantization, multi-GPU or scale-to-zero where relevant
• Before/after benchmark plan you can re-run after changes
• Optional implementation of high-impact fixes within the engagement window
Best for teams with rising GPU bills, p95 latency SLOs, or a PoC that needs production economics.
Stack familiarity: vLLM, TGI, Triton, Kubernetes/GKE, AWS Batch/serverless inference paths, OpenAI-compatible gateways (LiteLLM), observability for LLM spans.
How we work:
1) Kickoff + access to metrics/staging (read-only preferred)
2) Audit + written findings
3) Prioritized recommendations and optional hands-on tuning
4) Handoff notes your team can operate
Not a ChatGPT wrapper consult — this is production inference engineering.
Programming Languages
Python, TypeScript, GoCoding Expertise
Performance Optimization, SecurityWhat's included $4,500
These options are included with the project scope.
$4,500
- Delivery Time 10 days
- Number of Revisions 0
About Ismail
Principal Engineer | Production RAG, LLM Infra & Multi-Tenant SaaS
Mississauga, Canada - 11:43 pm local time
What I do for clients
• Enterprise RAG & document intelligence (PDF/drawing pipelines, hybrid retrieval, citations, access control)
• LLM inference cost & latency optimization (vLLM/Triton/K8s; on my last platform: ~25% cost / ~40% latency)
• Multi-tenant SaaS platforms: dual-plane auth, Stripe billing lifecycle, usage metering, audit logs
• Full-stack GenAI products: FastAPI/Node + React/Next.js + AWS Lambda/CDK or GKE
Current work (employment)
Principal Engineer at Infin8 Information Technologies building Takeoff—a multi-tenant construction estimation SaaS on AWS Canada Central (serverless FastAPI, dual Cognito, Stripe, AI drawing detection & document indexing). I own platform, billing, and AI workers end-to-end.
Prior highlights
• Silver Creek Insights — Principal Engineer: multi-GPU GenAI inference (vLLM/Kubernetes) and RAG/agent platforms
• Homewood Health — Senior Engineer (7 yrs): Canada's leading digital mental-health SaaS, 100k+ users, ~99.5% uptime; Next.js modernization, AWS migration, SAML/OAuth enterprise auth
• Scotiabank — zero-downtime IBM OpenPages GRC upgrades
How engagements work
1) Paid discovery/audit (1–2 weeks) or fixed-scope Project Catalog offering
2) Clear SOW, weekly demos, IaC + runbooks on handoff
3) I work remote from Toronto (ET); overlap with US East/West
Stack (highlights)
Python, FastAPI, TypeScript, Node.js, React/Next.js, AWS Lambda, CDK, Aurora, Cognito, Stripe, EventBridge, Batch, Valkey, GKE, vLLM, Qdrant, OpenAI/HF, Terraform, GitHub Actions
Not a fit for: toy ChatGPT wrappers, unlimited-scope $500 "clone X" apps, or anything illegal/unverified-age content.
Steps for completing your project
After purchasing the project, send requirements so Ismail can start the project.
Delivery time starts when Ismail receives requirements from you.
Ismail works on your project following the steps below.
Revisions may occur after the delivery date.
Kickoff & baseline
Review serving stack, metrics, cost baseline, SLOs, and access constraints.
Audit & analysis
Inspect batching, KV-cache, quantization, autoscaling; capture latency and cost data.
