You will get audit your LLM inference stack for cost and latency

Ismail K.Status: Offline
Ismail K.

Let a pro handle the details

Buy Web Application Programming services from Ismail, priced and ready to go.
Ismail K.Status: Offline
Ismail K.

Let a pro handle the details

Buy Web Application Programming services from Ismail, priced and ready to go.

Project details

What you get:

 • Baseline cost and latency report for your current inference stack
 • Actionable configuration changes: continuous batching, KV-cache/prefix caching, quantization, multi-GPU or scale-to-zero where relevant
 • Before/after benchmark plan you can re-run after changes
 • Optional implementation of high-impact fixes within the engagement window

Best for teams with rising GPU bills, p95 latency SLOs, or a PoC that needs production economics.

Stack familiarity: vLLM, TGI, Triton, Kubernetes/GKE, AWS Batch/serverless inference paths, OpenAI-compatible gateways (LiteLLM), observability for LLM spans.

How we work:
1) Kickoff + access to metrics/staging (read-only preferred)
2) Audit + written findings
3) Prioritized recommendations and optional hands-on tuning
4) Handoff notes your team can operate

Not a ChatGPT wrapper consult — this is production inference engineering.
Programming Languages
Python, TypeScript, Go
Coding Expertise
Performance Optimization, Security

What's included $4,500

These options are included with the project scope.

$4,500
  • Delivery Time 10 days
  • Number of Revisions 0
Ismail K.Status: Offline

About Ismail

Ismail K.Status: Offline
Principal Engineer | Production RAG, LLM Infra & Multi-Tenant SaaS
Mississauga, Canada - 11:43 pm local time
I'm a principal software engineer with 15+ years shipping production systems—and the last stretch has been making GenAI and multi-tenant SaaS actually survive production.

What I do for clients
• Enterprise RAG & document intelligence (PDF/drawing pipelines, hybrid retrieval, citations, access control)
• LLM inference cost & latency optimization (vLLM/Triton/K8s; on my last platform: ~25% cost / ~40% latency)
• Multi-tenant SaaS platforms: dual-plane auth, Stripe billing lifecycle, usage metering, audit logs
• Full-stack GenAI products: FastAPI/Node + React/Next.js + AWS Lambda/CDK or GKE

Current work (employment)
Principal Engineer at Infin8 Information Technologies building Takeoff—a multi-tenant construction estimation SaaS on AWS Canada Central (serverless FastAPI, dual Cognito, Stripe, AI drawing detection & document indexing). I own platform, billing, and AI workers end-to-end.

Prior highlights
• Silver Creek Insights — Principal Engineer: multi-GPU GenAI inference (vLLM/Kubernetes) and RAG/agent platforms
• Homewood Health — Senior Engineer (7 yrs): Canada's leading digital mental-health SaaS, 100k+ users, ~99.5% uptime; Next.js modernization, AWS migration, SAML/OAuth enterprise auth
• Scotiabank — zero-downtime IBM OpenPages GRC upgrades

How engagements work
1) Paid discovery/audit (1–2 weeks) or fixed-scope Project Catalog offering
2) Clear SOW, weekly demos, IaC + runbooks on handoff
3) I work remote from Toronto (ET); overlap with US East/West

Stack (highlights)
Python, FastAPI, TypeScript, Node.js, React/Next.js, AWS Lambda, CDK, Aurora, Cognito, Stripe, EventBridge, Batch, Valkey, GKE, vLLM, Qdrant, OpenAI/HF, Terraform, GitHub Actions

Not a fit for: toy ChatGPT wrappers, unlimited-scope $500 "clone X" apps, or anything illegal/unverified-age content.

Steps for completing your project

After purchasing the project, send requirements so Ismail can start the project.

Delivery time starts when Ismail receives requirements from you.

Ismail works on your project following the steps below.

Revisions may occur after the delivery date.

Kickoff & baseline

Review serving stack, metrics, cost baseline, SLOs, and access constraints.

Audit & analysis

Inspect batching, KV-cache, quantization, autoscaling; capture latency and cost data.

Review the work, release payment, and leave feedback to Ismail.