Node.js Developer to Build Backend + API for Scraped Data Pipeline

Posted 3 hours ago

Worldwide

Summary

We're building a matching system. On one side we have a database of providers and their capabilities, built from scraped public data. On the other side, a user submits their requirements through a form. The system's job is to score every provider against those requirements and return a ranked list of the best fits, with a clear reason for each match. We need a Node.js developer to own the backend: getting the scraped data into a usable shape, designing the capability model, and building the matching engine that sits on top of it. The data sources will be finalized as we go, so the ingestion layer needs to be source-agnostic — adding a new source should mean writing a small adapter, not rebuilding the pipeline. The core challenge Scraped provider data is messy and unstructured. Two providers might describe the same capability in completely different words, or bury it in a paragraph of marketing copy. Meanwhile, users describe what they need in their own language. The hard part of this project is bridging that gap: turning free-form provider data into a structured capability model that user requirements can actually be matched against. If you have opinions on how to do that well, we want to hear them. What you'll build Data pipeline Source-agnostic ingestion layer with an adapter pattern for adding new sources Normalization and cleaning of inconsistent, free-text provider data Deduplication and record merging across sources and across runs Scheduled refresh jobs with retries, backoff, and failure alerting Capability model A structured schema/taxonomy for provider attributes and capabilities Mapping logic that converts raw scraped text into that structure A way for us to review, correct, and improve mappings over time Matching engine Hard filters for non-negotiable requirements (disqualify providers outright) Weighted scoring for soft preferences, with tunable weights we can adjust without a redeploy Ranked results with a match score and a short explanation of why each provider scored the way it did Graceful handling of thin results — near-misses and relaxed criteria rather than an empty page API & infrastructure REST API (or GraphQL — open to your recommendation) for submitting requirements and retrieving ranked matches Auth and rate limiting for API consumers Queue-based background workers Logging, health checks, and alerts when a source breaks or data volume drops unexpectedly API documentation and a README clear enough to hand off internally Required skills Strong Node.js (Express, Fastify, or NestJS) [PostgreSQL / MongoDB] — real schema design and query optimization, not just CRUD Experience wrangling messy, inconsistent, real-world data Queue systems (BullMQ, Redis, or similar) Docker and deployment to [AWS / GCP / Railway / Render / DigitalOcean] Clear Git hygiene and readable code Nice to have Prior work on search, ranking, recommendation, or matching systems Text processing, fuzzy matching, or embedding-based similarity Scraping experience (Puppeteer, Playwright, Cheerio) TypeScript Elasticsearch or a vector database Deliverables Working backend deployed to [our infrastructure / a provider you recommend] Source code in our GitHub repo, with tests covering the matching logic API documentation A short handoff call or Loom walkthrough To apply Please skip the generic proposal. I'll only respond to applications that include: A similar system you've built — matching, search, ranking, or a data pipeline. What was the hardest part? How you'd approach the capability model. How do you get messy provider text into something matchable? One paragraph is plenty. Your recommended stack and one sentence on why. Start the first line of your proposal with the word "Matching" so I know you read this. Rough timeline and cost estimate appreciated. Happy to answer questions before you commit to a number.

  • More than 30 hrs/week
    Hourly
  • 3-6 months
    Duration
  • Intermediate
    Experience Level
  • $3.00

    -

    $5.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
API
Node.js
API Development
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:2 hours ago
  • Interviewing:
    3
  • Invites sent:
    10
  • Unanswered invites:
    5
About the client
Member since Mar 12, 2025
  • USA
    Waldwick4:00 AM
  • $215 total spent
    5 hires, 2 active

Explore similar jobs on Upwork

Zero Trust Access Blog Content WritingFixed-price‐ Posted 1 month ago
Java
Golang
Zero Trust Architecture
Content Writing
Article Writing
Python
Firebase Cloud Firestore
Firebase
FastAPI
Asynchronous I/O
Python Asyncio

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo