AI Platform Audit Engineer — Agent Loop Optimization and Runtime Observability
Worldwide
I have an existing AI platform with agent workflows, RAG/document retrieval, API integrations, task execution, and runtime services. I’m looking for an experienced AI systems engineer to audit the current AI loop, identify performance bottlenecks, and build a lightweight audit and observability system. This is not a from-scratch chatbot project. You will work with an existing codebase and help us understand why certain AI workflows are slow, expensive, unreliable, or difficult to debug. Scope of work: - Trace the full request lifecycle from user input to retrieval, prompting, model calls, tool execution, retries, and final response; - Identify slow steps, unnecessary model calls, repeated retrieval, oversized context, inefficient prompts, and retry loops; - Measure latency, token usage, model cost, failure rates, retry counts, and tool execution time; - Build an audit system with correlation IDs, timestamps, step status, duration, model usage, errors, and trace relationships; - Protect sensitive data and avoid storing API keys or unnecessary raw user content; - Produce a prioritized bottleneck report with evidence from real traces; - Implement and verify at least one high-impact optimization; - Document the architecture, metrics, findings, and recommended next steps. Expected deliverables: - AI loop architecture and trace map; - Working audit/observability layer; - Trace data for representative workflows; - Bottleneck and root-cause report; - Before-and-after measurement for one optimization; - Setup instructions and technical documentation. Acceptance criteria: - One request can be traced across all major AI workflow steps; - Each step exposes status, duration, errors, and relevant metadata; - The top three bottlenecks are supported by trace data; - At least one optimization demonstrates measurable improvement; - The audit system is maintainable and documented. Required experience: Python, FastAPI, LLM APIs, AI agents, RAG, PostgreSQL, API integrations, debugging, and production observability. Experience with OpenTelemetry, Prometheus, Grafana, LangChain, LangGraph, or custom AI tracing systems is a plus.
- Less than 30 hrs/weekHourly
- 1-3 monthsDuration
- IntermediateExperience Level
$19.00
-
$40.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:20 to 50
- Last viewed by client:2 hours ago
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- ChinaQingdao5:47 PM
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by