You will get AI Agent Security Audit | LLM Pentest | Prompt Injection | OWASP Top 10


Project details
Most AI systems in production were never tested against the thing that actually breaks them: input the builder did not write.
I audit LLM applications, RAG pipelines and tool-using agents for the failures specific to them. Prompt injection through documents, retrieved content and tool output. Tool permissions wider than the job requires. Data leaking across tenants or into the model provider. Model spend with no ceiling, on the exact surface an attacker can loop.
You get every finding with a reproduction you can run yourself, ranked by what it actually costs you, and a fix. Not a scanner report. Not a list of theoretical risks copied from a framework. If a finding has no reproduction, it does not go in the report.
Background: I build these systems, which is why I know where they break. Most recently a production cognitive platform with a memory graph, retrieval over a real corpus, multi-persona agents and streaming, delivered across fifteen approved milestones. Before that, independent security research and blockchain protocol audits.
Everything is written and asynchronous. You get a plan before I touch anything, and numbers instead of adjectives.
I audit LLM applications, RAG pipelines and tool-using agents for the failures specific to them. Prompt injection through documents, retrieved content and tool output. Tool permissions wider than the job requires. Data leaking across tenants or into the model provider. Model spend with no ceiling, on the exact surface an attacker can loop.
You get every finding with a reproduction you can run yourself, ranked by what it actually costs you, and a fix. Not a scanner report. Not a list of theoretical risks copied from a framework. If a finding has no reproduction, it does not go in the report.
Background: I build these systems, which is why I know where they break. Most recently a production cognitive platform with a memory graph, retrieval over a real corpus, multi-persona agents and streaming, delivered across fifteen approved milestones. Before that, independent security research and blockchain protocol audits.
Everything is written and asynchronous. You get a plan before I touch anything, and numbers instead of adjectives.
AI Algorithms
Large Language Model, Transformer ModelAI Applications
AIOps, Conversational AI, Natural Language UnderstandingAI Development Language
PythonAI Tools
Azure OpenAI, Hugging FaceAI Models
ChatGPT, GPT-4What's included
| Service Tiers |
Starter
$299
|
Standard
$1,200
|
Advanced
$2,500
|
|---|---|---|---|
| Delivery Time | 3 days | 7 days | 14 days |
Number of Revisions | 1 | 2 | 3 |
AI Model Integration | - | - | - |
Batch Normalization | - | - | - |
Database Integration | - | - | - |
Detailed Code Comments | - | ||
Image Upscaling | - | - | - |
MLOps | - | - | |
Model Deployment | - | - | |
Model Documentation | |||
Model Monitoring | - | - | |
Model Testing & Optimization | |||
Model Tuning | - | - | - |
Natural Language Processing | - | - | - |
NLP Tokenization | - | - | - |
Pre-Training | - | - | - |
Prompt Engineering | |||
Setup File | - | - | - |
Source Code | - |
About Guglielmo
AI Agents, Chatbots & Workflow Automation | Production AI Engineer
Milan, Italy - 7:10 am local time
I build AI agents, chatbots and automations that run in production: connected to your real data, wired into the tools your team already uses, and still standing after the first thousand messages.
WHAT I BUILD
• AI agents that actually do the work. Multi-step reasoning, tool calling, and the guardrails that stop them going off-script or running up your bill.
• Chatbots that answer from your own documents, with the source attached to every answer and an honest "I don't know" instead of an invented one.
• Workflow automation, n8n or custom, that removes the task nobody on your team wants to do twice.
• AI built into the app you already have, rather than a separate toy nobody opens.
WHY CLIENTS KEEP COMMISSIONING
I ship on rails. A written plan before I touch code. Staging before production. Numbers you can verify instead of adjectives. I name my own bugs before you find them, and I will tell you when a feature is not worth building, which costs me money and saves you more.
My last client commissioned continuously for five months and wrote this: "one of the most substantive technical collaborations of my career."
HOW I WORK
Fully asynchronous, entirely in writing. No calls, no standups, no scheduling across time zones. You get updates you can read at 2am and forward to your team, and every decision stays on the record instead of evaporating in a meeting.
Stack: Python, FastAPI, Next.js, PostgreSQL, pgvector, Docker, OpenAI and Anthropic APIs, Stripe. MIT Professional Education certificate in Large Language Models. Background in security, which mostly shows up as systems that fail closed instead of loudly.
Tell me the outcome you need and what is currently in the way. You get back a scoped plan with a fixed price, usually same day.
Steps for completing your project
After purchasing the project, send requirements so Guglielmo can start the project.
Delivery time starts when Guglielmo receives requirements from you.
Guglielmo works on your project following the steps below.
Revisions may occur after the delivery date.
Map the attack surface
I read the code, the prompts and the tool definitions, then map every path where untrusted input reaches a model, a tool or your data. You get the map before I test anything, so you can tell me what I misread.