Developer Needed for Private AI Chat System (API-based)
Worldwide
Developer Needed for Private AI Chat System (API-based) Project Overview I need a private, self-hosted AI assistant that connects in real-time to third-party AI APIs (Anthropic Claude and/or Google Gemini) — NOT a local/offline LLM (no Ollama, no local model weights). The system must be accessible via a web interface from multiple devices (laptop, phone, tablet) and multiple locations, with my data staying under my control. What I Need Built 1. Backend server that securely calls the Anthropic API and/or Google Gemini API - API keys stored server-side only (never exposed to the browser/client) - Support for switching between or combining providers 2. Simple web-based chat interface - Responsive design (usable on desktop and mobile browsers) - User login/authentication (so only I — and anyone I authorize — can access it) 3. Document storage & retrieval (RAG - Retrieval-Augmented Generation) - Ability to upload my own documents/files (PDF, Word, text, etc.) through the web interface - Documents automatically split into chunks and converted into embeddings (numerical representations for semantic search) - Embeddings stored in a self-hosted vector database on my own server (e.g., Qdrant, Weaviate, or pgvector — developer to recommend based on scale) - At query time, the system retrieves only the most relevant chunks (semantic similarity search) and sends those — not the full documents — to the AI API along with my question - Full documents and the vector database must remain on my own infrastructure at all times; only the retrieved text snippets are sent externally to Claude/Gemini per query - Ability to update/delete documents from the knowledge base (re-indexing when content changes) - Source attribution: responses should indicate which document(s) the answer was drawn from 4. Content ingestion pipeline (for large personal knowledge sources) - Ability to bulk-ingest a large personal library of trading education content (e.g., YouTube video transcripts) into the RAG knowledge base - Pipeline should: fetch/extract transcripts, clean up the text (remove filler, fix formatting), chunk by topic/concept rather than fixed length, generate embeddings, and index into the vector database - Should also support ingesting reference/technical documentation (e.g., the official Pine Script language reference) as a separate, always-available knowledge source used specifically to ground any code generation and reduce errors/hallucinations - Reusable pipeline: I should be able to add new sources (new videos, new documents) later without needing the developer each time 5. Live web search capability - The assistant should be able to search the web in real time for current information (e.g., market news, recent Pine Script/TradingView updates) when a question needs up-to-date data - Implement via a web search tool/plugin connected to the AI API (e.g., Claude's built-in web search tool, or a search API such as Google/Bing/Perplexity Sonar integrated into the backend) - Should be clearly distinguished from the RAG knowledge base: RAG = my own static documents/strategies; web search = live, current information from the internet - Results should include source links so I can verify information 6. Hosting & deployment - Deployed on a VPS (I will provide access — see below) or recommend one with justification - Dockerized setup preferred, for portability and easy maintenance - Must remain accessible 24/7 via a secure URL (HTTPS) 7. Logging & basic security - Request logs (who asked what, when) - Basic protection against unauthorized access (rate limiting, authentication) 8. Documentation - Clear instructions on how to maintain, update, and restart the system - How to add/rotate API keys - How to add new users if needed Requirements for the Developer - Proven experience with backend development (Python/FastAPI or Node.js/Express) - Experience integrating LLM APIs (Anthropic, OpenAI, or Google Gemini API) - Experience with vector databases / RAG pipelines (e.g., Pinecone, Qdrant, Weaviate, or pgvector) - Experience deploying and managing applications on a VPS using Docker - Understanding of basic security practices (secrets management, authentication, HTTPS/SSL setup) - Able to communicate clearly in English and explain technical choices in plain language - Portfolio or examples of similar past projects (chatbots, AI integrations, RAG systems) - Experience with text extraction/ingestion pipelines (e.g., transcript extraction, document parsing, bulk chunking strategies) is a plus - Experience integrating web search tools/APIs (e.g., Claude's web search tool, Google Search API, Bing API, or Perplexity Sonar) is a plus This will be a project based contract (Upwork fixed contract with milestones), where gradual payments will be released upon milestones achievements. It will be considered finalized when everything will be done/completed. This is not a "pay-by-the-hour" contract. The total budget & exact milestones for this project will be discussed and agreed upon.
- Hours to be determinedHourly
- < 1 monthDuration
- ExpertExperience Level
- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:40 minutes ago
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- United Arab EmiratesDubai9:02 PM
- $9.5K total spent3 hires, 0 active
- 125 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by