You will get an AI agent audit: cost, security & reliability (Claude)

Project details
Your AI agents are running. Do you actually know what they're doing? Most teams ship agents and never look back: nobody knows which agent burns the budget, what tools each one can really reach, or what happens when a malicious prompt walks in.
I audit AI agent systems the way an accountant audits books - evidence, not opinions. I examine cost attribution (including what you pay for failed runs), security (real tool permissions, credential exposure, prompt-injection surface, isolation), reliability (memory and state integrity, silent failures) and your audit trail. You get a written findings report ranked by severity, a concrete fix for every finding, and a walkthrough call showing the evidence.
Why me: I recently completed a scored technical assessment of a ten-agent deployment for a US client - 1,601 logged read-only commands, a 16-page report, highest score of all candidates evaluated. I run my own multi-agent fleet in production with code-enforced allowlists, approval gates and append-only audit logs. 7 Anthropic Academy certifications; 16 years of enterprise delivery. Read-only by design: your system stays untouched.
I audit AI agent systems the way an accountant audits books - evidence, not opinions. I examine cost attribution (including what you pay for failed runs), security (real tool permissions, credential exposure, prompt-injection surface, isolation), reliability (memory and state integrity, silent failures) and your audit trail. You get a written findings report ranked by severity, a concrete fix for every finding, and a walkthrough call showing the evidence.
Why me: I recently completed a scored technical assessment of a ten-agent deployment for a US client - 1,601 logged read-only commands, a 16-page report, highest score of all candidates evaluated. I run my own multi-agent fleet in production with code-enforced allowlists, approval gates and append-only audit logs. 7 Anthropic Academy certifications; 16 years of enterprise delivery. Read-only by design: your system stays untouched.
AI Development Type
Model Tuning, Software MaintenanceAI Development Language
PythonWhat's included
| Service Tiers |
Starter
$495
|
Standard
$950
|
Advanced
$1,800
|
|---|---|---|---|
| Delivery Time | 5 days | 7 days | 14 days |
Number of Revisions | 1 | 1 | 2 |
AI Model Integration | - | - | - |
Detailed Code Comments | - | - | - |
Knowledge Graph | - | - | - |
Model Documentation | |||
Ontology | - | - | - |
Source Code | - | - | |
Taxonomy | - | - | - |
Frequently asked questions
About Fabian
AI Agent Audits & Automation | Claude, MCP, RAG | In Production
Mexico City, Mexico - 2:56 pm local time
Most AI "experts" can prototype. I ship systems that operate daily — and I audit the ones other people shipped.
What I do:
→ AI agent audits: cost attribution, security, permissions, reliability — evidence, not opinions
→ AI agents & assistants (tool-use, memory, multi-model routing) on Claude, MCP, and OpenClaw
→ RAG systems and LLM apps over your own data
→ Document & process automation (OCR + AI, n8n, CRM, API integrations)
Recent work:
• Scored security & cost audit of a ten-agent AI fleet for a US client — highest score of all candidates evaluated; found that 49.4% of a month's API spend was buying failed runs.
• TimbraBot — a multi-tenant LLM engine that reconciles financial transactions in daily production, reaching up to 80% automation.
• Multi-agent fleet with code-enforced tool permissions, approval gates, and append-only audit logs — running my own studio's operations.
7 Anthropic Academy certifications — Claude Code, MCP, subagents, agent skills, Claude Cowork, Claude on Amazon Bedrock, AI fluency (verifiable on this profile).
I translate business problems into AI that works — not demos. Typical engagement: from problem to a system in production in 6–8 weeks; smaller automations in days.
Stack: Claude/OpenAI · Python · RAG · agents & tool-calling · MCP · n8n · Ollama · Next.js · Supabase · Docker.
Based in Mexico — US timezone overlap, bilingual (EN/ES), enterprise quality at nearshore rates.
Steps for completing your project
After purchasing the project, send requirements so Fabian can start the project.
Delivery time starts when Fabian receives requirements from you.
Fabian works on your project following the steps below.
Revisions may occur after the delivery date.
Scope call & access setup
45-minute call to map your system. We agree what is in and out of scope, sign an NDA if you want one, and set up read-only access.
Evidence sweep
Systematic read-only review: tool permissions, credentials, cross-agent isolation, memory and state, log integrity, and per-agent cost attribution including failure spend. Every command logged.