AI Data Engineer, FastAPI / PostgreSQL / RAG Pipeline / LLM Data Quality

Posted yesterday

Worldwide

Summary

We operate a growing automotive fault database, currently receiving approximately 500-1,000 organic search clicks per day. The platform allows users to search for common faults, solutions, and technical information by vehicle make, model, and variant. It holds data on over 64,000 vehicles and 94,000 recorded faults. The data was originally generated using an AI/RAG pipeline. We are now at a stage where we need to improve data accuracy, expand the dataset, and add new features. We want expert advice on the best technical approach before any implementation begins. The site URL will be shared with shortlisted candidates. We are posting this job with restricted visibility to protect commercially sensitive details. Tech Stack Layer Technology Frontend React 19, Vite, TailwindCSS, shadcn/ui Backend FastAPI, SQLAlchemy, Uvicorn Database PostgreSQL on AWS RDS Auth AWS Cognito, AWS Amplify Deployment Docker Compose on EC2, Nginx CDN / Static AWS S3 + CloudFront What We Are Looking For We are not looking for someone to start building immediately. We want a developer or technical consultant with hands-on experience in AI/RAG pipelines, PostgreSQL data modeling, and FastAPI to review the four problem areas below and provide clear recommendations on the best approach for each, including realistic effort estimates and any risks we should be aware of. If you have a strong view on the right solution, we want to hear it. If you think a different framing of the problem would lead to a better outcome, please tell us that too. Problem 1: Vehicle Data Structure Is Inaccurate The original RAG pipeline consolidated too many vehicle generations into single database records. For example, all generations of the BMW M3 (E36, E46, E90, F80) have been merged into one record, when in reality they are completely different vehicles with different engines, different faults, and different fault frequencies. The same problem exists across many other makes and models. Additionally, a significant number of modern variants (roughly 2022 onwards) are missing from the database entirely. We need your advice on: • What is the best way to restructure the vehicle taxonomy in PostgreSQL so that Make, Model, Generation, and Variant are properly separated and queryable? • Would linking to the DVLA Vehicle Enquiry Service API provide a better authoritative foundation for the vehicle make/model/variant structure? If so, how would this work alongside our existing data? • How do we add missing modern variants without rebuilding the entire dataset from scratch? • How do we handle the SEO impact? Our existing URLs are already indexed by Google and driving daily traffic. Any restructuring must not result in 404 errors or loss of ranking. Problem 2: Number Plate Lookup (UK Market) We want to allow UK visitors to enter their vehicle registration number to instantly identify their exact make, model, year, and variant — replacing the current manual dropdown selector. The DVLA VES API and DVSA MOT History API are both free, official government sources that return structured vehicle data including engine size, fuel type, and full MOT history. We need your advice on: • How should a number plate lookup integrate with our existing React frontend and FastAPI backend? • Should DVLA/DVSA API calls be made server-side (FastAPI) or client-side (React)? What are the security and rate-limit implications of each approach? • How do we reliably map the DVLA response (which returns make and model as plain text strings) to our internal vehicle database records? • What is the best UX pattern for handling cases where the DVLA returns a vehicle not yet in our database? Problem 3: Fault Data Quality - Correcting Existing Data and Minimising Hallucination The existing fault data was generated by an LLM/RAG pipeline and contains inaccuracies where some faults are attributed to the wrong generation, some descriptions are vague or incorrect, and some content may be hallucinated. We need a systematic way to review and correct this data. We need your advice on: • What is the best workflow for auditing and correcting LLM-generated fault data at scale? We are open to using AI tools to assist, but accuracy is paramount. We need a process that minimises hallucination rather than compounding it. • What authoritative data sources exist for vehicle fault data (e.g., DVSA technical bulletins, manufacturer TSBs, NHTSA) that could be used to validate or replace AI-generated content? • Should corrections be made directly in the database, or is there a better content management approach that allows non-technical team members to review and approve changes? • How do we prevent future data ingestion from repeating the same consolidation and hallucination errors? Problem 4: Enriching Vehicle Pages with Additional Content Each vehicle variant page currently shows fault data only. We want to add richer content to make each page more useful, including: • Recommended lubricants and engine oils (with specific grade and manufacturer spec for each engine variant) • Recommended fuel additives and treatments • Common modifications and tuning options We previously attempted to populate this data using automated AI scraping and found the results unreliable. We need your advice on: • What is the most reliable way to populate structured product recommendation data (oils, additives) for thousands of vehicle variants? Are there existing databases or APIs for oil specifications by engine code or manufacturer spec? • How should this additional content be structured in PostgreSQL to support the vehicle pages cleanly? • What is a realistic, scalable approach to populating modification and tuning data, given that this information is highly variant-specific and scattered across forums and third-party sources? To Apply Please briefly describe your experience with RAG pipelines, LLM-based data extraction, and PostgreSQL data modelling. We are particularly interested in candidates who have worked on data quality problems with AI-generated content.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $20.00

    -

    $30.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
React 19
TailwindCSS
PostgreSQL
Activity on this job
  • Proposals:50+
  • Last viewed by client:yesterday
  • Interviewing:
    7
  • Invites sent:
    12
  • Unanswered invites:
    5
About the client
Member since Mar 29, 2007
  • United Kingdom
    Attleborough1:20 AM
  • $19K total spent
    79 hires, 3 active
  • 506 hours

Explore similar jobs on Upwork

Data Governance- Atlan, Unity CatalogHourly‐ Posted 4 weeks ago
Data Engineering
Data Engineer for API PipelinesHourly‐ Posted 4 days ago
Python
Data Integration
Database Architecture
Data Transformation
ETL Pipeline
Data Preprocessing
SQL
Database Design
Data Engineering
Data Migration
API Development
Tableau
Google Cloud Platform

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo