You will get RAG Hallucination Audit — Measure & Fix Your AI's Accuracy


Project details
Your RAG or AI chatbot gives confident answers that aren't in your documents? That's a grounding failure — and it's measurable and fixable.
I'll measure whether your system actually tells the truth:
• I build a small golden dataset from YOUR real documents
• I measure your current groundedness and hallucination rate — real numbers, not "it seems better"
• I deliver a prioritized, reproducible fix plan you can act on
Why me: I built SavoirBase AI, a self-hosted document-AI platform with a MEASURED hallucination rate of 0.0 on a golden dataset. PhD in Computer Science, 3 IEEE publications, former Director General of a national cybersecurity agency.
You get a clear report with numbers — the first step to a RAG your users can actually trust.
I'll measure whether your system actually tells the truth:
• I build a small golden dataset from YOUR real documents
• I measure your current groundedness and hallucination rate — real numbers, not "it seems better"
• I deliver a prioritized, reproducible fix plan you can act on
Why me: I built SavoirBase AI, a self-hosted document-AI platform with a MEASURED hallucination rate of 0.0 on a golden dataset. PhD in Computer Science, 3 IEEE publications, former Director General of a national cybersecurity agency.
You get a clear report with numbers — the first step to a RAG your users can actually trust.
AI Algorithms
Large Language Model, Transformer ModelAI Applications
Conversational AI, Natural Language Generation, Natural Language UnderstandingAI Development Language
PythonAI Tools
Azure OpenAI, Hugging FaceAI Models
ChatGPT, GPT-4What's included
| Service Tiers |
Starter
$149
|
Standard
$299
|
Advanced
$499
|
|---|---|---|---|
| Delivery Time | 3 days | 5 days | 7 days |
Number of Revisions | 1 | 1 | 2 |
AI Model Integration | - | - | |
Batch Normalization | - | - | - |
Database Integration | - | - | - |
Detailed Code Comments | - | - | - |
Image Upscaling | - | - | - |
MLOps | - | - | - |
Model Deployment | - | - | - |
Model Documentation | |||
Model Monitoring | - | - | - |
Model Testing & Optimization | |||
Model Tuning | - | - | |
Natural Language Processing | - | ||
NLP Tokenization | - | - | - |
Pre-Training | - | - | - |
Prompt Engineering | - | ||
Setup File | - | - | - |
Source Code | - | - |
Optional add-ons
You can add these on the next page.
10 additional evaluation questions
(+ 1 Day)
+$75
Executive summary for non-technical stakeholders
(+ 1 Day)
+$60
Additional limited corrective fix
+$100Frequently asked questions
About Mahamat Nassour
AI Engineer | RAG, LLM & Document AI | PhD, Ex-Cybersecurity Director
N'Djamena, Chad - 9:01 pm local time
Most RAG systems hallucinate and almost nobody measures it. I do. On my own document-AI platform, on a golden dataset of real documents: groundedness 1.0, hallucination rate 0.0, appropriate refusal 1.0, deep-vision document quality 0.90 vs 0.44 for OCR alone. Real numbers with a reproducible method, not marketing claims.
WHAT I DO
- RAG systems and chatbots over your documents (PDF, Word, Excel, scans): hybrid search, reranking, grounded answers with page-level citations
- Audits of RAG/LLM apps that hallucinate: measured baseline, prioritized fix plan, regression tests you can re-run
- AI agents and LLM integrations (OpenAI, Claude, Ollama, OpenRouter) with cost tracking and no vendor lock-in
- Self-hosted and air-gapped AI where data cannot leave your infrastructure
- Security audits of AI applications: prompt injection, auth and role checks, exposed secrets
Stack: Python, FastAPI, Next.js, TypeScript, PostgreSQL 16 + pgvector, Ollama, OpenRouter, Docker, Keycloak, Redis, MinIO
I did not wrap a framework. I wrote the retrieval pipeline myself, which is why I can debug it when it fails.
WHAT I BUILT
SavoirBase AI, a complete self-hosted document intelligence platform I designed and built solo: 152 API operations, 21 screens, 391 tests. Full pipeline from ingestion and OCR through deep vision, adaptive chunking, bge-m3 embeddings, hybrid search (dense + full-text, RRF fusion), reranking, and strictly grounded generation that refuses to answer when no source supports it.
Also delivered: a sovereign state communication suite (Matrix, Keycloak, Hyperledger, K3s), a ministry-scale document management system, and a university ERP serving 2,450 students with handwritten OCR (I have worked on handwriting recognition since 2018).
BACKGROUND
- PhD in Computer Science, Universite de Reims Champagne-Ardenne, France
- Former Director General of a national cybersecurity and e-certification agency (3 years)
- Chaired the technical committee for the National Cybersecurity Strategy and for the Cybersecurity Maturity Model conducted with the University of Oxford
- 4 peer-reviewed publications (3 IEEE)
- MSc, INSEEC Paris
I build AI and I secure it, a combination very few people offer.
Languages: French (native), English (professional), Arabic
Tell me what your system is getting wrong, and I will tell you how I would measure it before touching anything.
START SMALL
You should not hand a four-figure project to someone with no reviews on this platform. I would not either. So start with a fixed-price audit: I build a golden dataset from your own documents, measure your system's real groundedness and hallucination rate, and hand you a prioritized fix plan with regression tests you can re-run yourself. Three days, fixed price, and you own everything I produce.
If the numbers come back good, you do not need me. If they do not, you will know exactly what is broken and what it costs to fix.
Steps for completing your project
After purchasing the project, send requirements so Mahamat Nassour can start the project.
Delivery time starts when Mahamat Nassour receives requirements from you.
Mahamat Nassour works on your project following the steps below.
Revisions may occur after the delivery date.
Build a golden dataset from your documents
I create 10-15 real question/answer pairs grounded in your actual documents, to serve as an objective ground truth.
Measure groundedness & hallucination rate
I run your system against the dataset and compute real numbers: how often it stays grounded vs. invents answers.



