AI Data Engineer, FastAPI / PostgreSQL / RAG Pipeline / LLM Data Quality
Worldwide
We operate a growing automotive fault database, currently receiving approximately 500-1,000 organic search clicks per day. The platform allows users to search for common faults, solutions, and technical information by vehicle make, model, and variant. It holds data on over 64,000 vehicles and 94,000 recorded faults. The data was originally generated using an AI/RAG pipeline. We are now at a stage where we need to improve data accuracy, expand the dataset, and add new features. We want expert advice on the best technical approach before any implementation begins. The site URL will be shared with shortlisted candidates. We are posting this job with restricted visibility to protect commercially sensitive details. Tech Stack Layer Technology Frontend React 19, Vite, TailwindCSS, shadcn/ui Backend FastAPI, SQLAlchemy, Uvicorn Database PostgreSQL on AWS RDS Auth AWS Cognito, AWS Amplify Deployment Docker Compose on EC2, Nginx CDN / Static AWS S3 + CloudFront What We Are Looking For We are not looking for someone to start building immediately. We want a developer or technical consultant with hands-on experience in AI/RAG pipelines, PostgreSQL data modeling, and FastAPI to review the four problem areas below and provide clear recommendations on the best approach for each, including realistic effort estimates and any risks we should be aware of. If you have a strong view on the right solution, we want to hear it. If you think a different framing of the problem would lead to a better outcome, please tell us that too. Problem 1: Vehicle Data Structure Is Inaccurate The original RAG pipeline consolidated too many vehicle generations into single database records. For example, all generations of the BMW M3 (E36, E46, E90, F80) have been merged into one record, when in reality they are completely different vehicles with different engines, different faults, and different fault frequencies. The same problem exists across many other makes and models. Additionally, a significant number of modern variants (roughly 2022 onwards) are missing from the database entirely. We need your advice on: • What is the best way to restructure the vehicle taxonomy in PostgreSQL so that Make, Model, Generation, and Variant are properly separated and queryable? • Would linking to the DVLA Vehicle Enquiry Service API provide a better authoritative foundation for the vehicle make/model/variant structure? If so, how would this work alongside our existing data? • How do we add missing modern variants without rebuilding the entire dataset from scratch? • How do we handle the SEO impact? Our existing URLs are already indexed by Google and driving daily traffic. Any restructuring must not result in 404 errors or loss of ranking. Problem 2: Number Plate Lookup (UK Market) We want to allow UK visitors to enter their vehicle registration number to instantly identify their exact make, model, year, and variant — replacing the current manual dropdown selector. The DVLA VES API and DVSA MOT History API are both free, official government sources that return structured vehicle data including engine size, fuel type, and full MOT history. We need your advice on: • How should a number plate lookup integrate with our existing React frontend and FastAPI backend? • Should DVLA/DVSA API calls be made server-side (FastAPI) or client-side (React)? What are the security and rate-limit implications of each approach? • How do we reliably map the DVLA response (which returns make and model as plain text strings) to our internal vehicle database records? • What is the best UX pattern for handling cases where the DVLA returns a vehicle not yet in our database? Problem 3: Fault Data Quality - Correcting Existing Data and Minimising Hallucination The existing fault data was generated by an LLM/RAG pipeline and contains inaccuracies where some faults are attributed to the wrong generation, some descriptions are vague or incorrect, and some content may be hallucinated. We need a systematic way to review and correct this data. We need your advice on: • What is the best workflow for auditing and correcting LLM-generated fault data at scale? We are open to using AI tools to assist, but accuracy is paramount. We need a process that minimises hallucination rather than compounding it. • What authoritative data sources exist for vehicle fault data (e.g., DVSA technical bulletins, manufacturer TSBs, NHTSA) that could be used to validate or replace AI-generated content? • Should corrections be made directly in the database, or is there a better content management approach that allows non-technical team members to review and approve changes? • How do we prevent future data ingestion from repeating the same consolidation and hallucination errors? Problem 4: Enriching Vehicle Pages with Additional Content Each vehicle variant page currently shows fault data only. We want to add richer content to make each page more useful, including: • Recommended lubricants and engine oils (with specific grade and manufacturer spec for each engine variant) • Recommended fuel additives and treatments • Common modifications and tuning options We previously attempted to populate this data using automated AI scraping and found the results unreliable. We need your advice on: • What is the most reliable way to populate structured product recommendation data (oils, additives) for thousands of vehicle variants? Are there existing databases or APIs for oil specifications by engine code or manufacturer spec? • How should this additional content be structured in PostgreSQL to support the vehicle pages cleanly? • What is a realistic, scalable approach to populating modification and tuning data, given that this information is highly variant-specific and scattered across forums and third-party sources? To Apply Please briefly describe your experience with RAG pipelines, LLM-based data extraction, and PostgreSQL data modelling. We are particularly interested in candidates who have worked on data quality problems with AI-generated content.
- Less than 30 hrs/weekHourly
- 1-3 monthsDuration
- ExpertExperience Level
$20.00
-
$30.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:yesterday
- Interviewing:7
- Invites sent:12
- Unanswered invites:5
About the client
- United KingdomAttleborough1:20 AM
- $19K total spent79 hires, 3 active
- 506 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by