Need help fixing latency bottlenecks & context issues in our Langgraph/Qdrant RAG setup
Worldwide
Looking for a developer to help audit and speed up our RAG pipeline and multi-agent workflows. We're running FastAPI on the backend with LangGraph for agent orchestration, Qdrant for vector retrieval and self-hosted vLLM endpoints. Everything works functionally, but as we've scaled up, we're hitting a few pain points that's hurting performance and response times: - Our LangGraph multi-agent execution is taking way too long on complex steps and state feels heavy - Longer retrieval passes sometimes blows past our context window or just starts pulling noisy chunks - Sometimes function calling break during LLM outputs - Response streaming via SSE stutters/blocks under concurrent user requests. Looking for someone who can jump into the codebase and pinpoint where the bottlenecks are coming from and clean up the implementation.
$400.00
Fixed-price- IntermediateExperience Level
- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:15 to 20
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- Pakistan12:41 AM
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by