You will get LLM Cost Optimization | Model Routing | Token Reduction | API Cost Audit


Project details
Most teams find out what their AI costs when the invoice arrives, and by then the number is a single total with no way to argue with it.
I break the bill apart. Spend per route, per model, per token, so you can see which feature is expensive and which one you assumed was expensive and is not. Usually one or two routes carry the whole cost, and they are rarely the ones people expect.
Then I cut it, in the order that costs you the least to accept. Routing, so the cheap model handles what it can already handle. Caching, so you stop paying twice for the same answer. Prompts, because tokens you never needed are still tokens you pay for. Ceilings, so a loop or an attacker cannot run your bill overnight.
Every change comes with a measured before and after on your own traffic. No projections, no industry averages, no percentages borrowed from a blog post.
If I go through your setup and find nothing worth cutting, you pay nothing. That offer is only sane because it is rarely the outcome.
Written and asynchronous throughout, and I will tell you which cuts I think are not worth the quality tradeoff rather than pushing the number down for its own sake.
I break the bill apart. Spend per route, per model, per token, so you can see which feature is expensive and which one you assumed was expensive and is not. Usually one or two routes carry the whole cost, and they are rarely the ones people expect.
Then I cut it, in the order that costs you the least to accept. Routing, so the cheap model handles what it can already handle. Caching, so you stop paying twice for the same answer. Prompts, because tokens you never needed are still tokens you pay for. Ceilings, so a loop or an attacker cannot run your bill overnight.
Every change comes with a measured before and after on your own traffic. No projections, no industry averages, no percentages borrowed from a blog post.
If I go through your setup and find nothing worth cutting, you pay nothing. That offer is only sane because it is rarely the outcome.
Written and asynchronous throughout, and I will tell you which cuts I think are not worth the quality tradeoff rather than pushing the number down for its own sake.
AI Algorithms
Large Language Model, Transformer ModelAI Applications
AIOps, Conversational AI, Natural Language GenerationAI Development Language
PythonAI Tools
Azure OpenAI, Hugging FaceAI Models
ChatGPT, GPT-4What's included
| Service Tiers |
Starter
$149
|
Standard
$700
|
Advanced
$1,800
|
|---|---|---|---|
| Delivery Time | 3 days | 7 days | 14 days |
Number of Revisions | 1 | 2 | 3 |
AI Model Integration | - | - | - |
Batch Normalization | - | - | - |
Database Integration | - | - | - |
Detailed Code Comments | - | - | - |
Image Upscaling | - | - | - |
MLOps | - | - | |
Model Deployment | - | - | |
Model Documentation | |||
Model Monitoring | - | - | |
Model Testing & Optimization | |||
Model Tuning | - | ||
Natural Language Processing | - | - | - |
NLP Tokenization | - | - | - |
Pre-Training | - | - | - |
Prompt Engineering | - | ||
Setup File | - | - | - |
Source Code | - |
About Guglielmo
AI Agents, Chatbots & Workflow Automation | Production AI Engineer
Milan, Italy - 12:58 pm local time
I build AI agents, chatbots and automations that run in production: connected to your real data, wired into the tools your team already uses, and still standing after the first thousand messages.
WHAT I BUILD
• AI agents that actually do the work. Multi-step reasoning, tool calling, and the guardrails that stop them going off-script or running up your bill.
• Chatbots that answer from your own documents, with the source attached to every answer and an honest "I don't know" instead of an invented one.
• Workflow automation, n8n or custom, that removes the task nobody on your team wants to do twice.
• AI built into the app you already have, rather than a separate toy nobody opens.
WHY CLIENTS KEEP COMMISSIONING
I ship on rails. A written plan before I touch code. Staging before production. Numbers you can verify instead of adjectives. I name my own bugs before you find them, and I will tell you when a feature is not worth building, which costs me money and saves you more.
My last client commissioned continuously for five months and wrote this: "one of the most substantive technical collaborations of my career."
HOW I WORK
Fully asynchronous, entirely in writing. No calls, no standups, no scheduling across time zones. You get updates you can read at 2am and forward to your team, and every decision stays on the record instead of evaporating in a meeting.
Stack: Python, FastAPI, Next.js, PostgreSQL, pgvector, Docker, OpenAI and Anthropic APIs, Stripe. MIT Professional Education certificate in Large Language Models. Background in security, which mostly shows up as systems that fail closed instead of loudly.
Tell me the outcome you need and what is currently in the way. You get back a scoped plan with a fixed price, usually same day.
Steps for completing your project
After purchasing the project, send requirements so Guglielmo can start the project.
Delivery time starts when Guglielmo receives requirements from you.
Guglielmo works on your project following the steps below.
Revisions may occur after the delivery date.
Find out where the money actually goes
I break your spend down per route, per model and per token, so you see which feature costs what instead of one number at the end of the month. Most bills have one or two routes carrying the whole cost, and they are rarely the ones people expect.