Playwright + Firebase Developer for Production Data Ingestion System

Posted yesterday

Worldwide

Summary

We are looking for an experienced independent developer to build a production web-scraping and data-ingestion system that monitors approximately 30 public websites. This is not a one-time scraping project. The system will repeatedly monitor a known set of sources, remember previously observed state, detect meaningful changes, and generate structured proposals for human review before approved data reaches the production application. A detailed technical architecture has already been designed. We are looking for someone who can implement it pragmatically, challenge unnecessary complexity, and favor simple, maintainable solutions over over-engineering. WHAT YOU WILL OWN + Playwright scraper framework + Site-specific scraper adapters + TypeScript / Node.js ingestion logic + Listing-level and detail-page extraction + Data normalization + Source-state storage + Delta detection for new, changed, missing, and unchanged records + Source-health safeguards to prevent bad, partial, blocked, or incomplete scrapes from creating false changes + Partition-aware handling for cases where one section, city, page, or navigation branch fails while the rest of a source succeeds + Entity matching and deduplication across sources + Multi-source provenance + Structured proposal generation + Idempotency and duplicate protection + Rejection / reconciliation handling so the same unchanged observation does not repeatedly create proposals + Synthetic fixtures and regression tests + Scraper diagnostics, failure artifacts, and maintenance + Rapid site-specific adaptation once the real target sites begin launching EXISTING APPLICATION DEVELOPER The existing application developer owns: + Production Firebase / Firestore data model + Admin and human-review workflow + Production publishing pipeline + Client synchronization + Mobile and web application behavior + Canonical production entity schemas and matching rules The ingestion developer will work against a clearly defined integration contract and apply the agreed matching rules against a stable interface to the production data. The ingestion developer is not expected to redesign the production application or independently build the admin/publishing system. REQUIRED EXPERIENCE You should have strong experience with: + Playwright + Node.js / TypeScript + Firebase / Firestore + Recurring, stateful production scraping systems + Dynamic JavaScript websites + Data normalization and entity matching + Idempotent background jobs + Change / delta detection + Failure handling and retry logic + Testing and debugging scrapers as websites change + Observability and diagnostics for production jobs Experience building a one-time scraper is not enough for this project. We are especially interested in candidates who have built systems that run repeatedly over time and must safely distinguish between: + New records + Changed records + Missing records + Unchanged records + Partial or unhealthy scrape results AI-ASSISTED DEVELOPMENT We strongly value developers who use modern AI coding tools effectively, such as: + Claude Code + Codex + Cursor + Similar coding agents or AI-assisted development tools We are not looking for someone who simply delegates the entire project to AI. We are looking for a developer who can use these tools to accelerate repetitive implementation work while still exercising strong engineering judgment, reviewing generated code, writing tests, and validating production behavior. PROJECT PHILOSOPHY This should be a lean, maintainable v1. We do not want enterprise infrastructure for its own sake. The architecture describes required behavior and safeguards, but the developer is encouraged to simplify implementation when the same reliability and data integrity can be achieved with a simpler approach. The core workflow is: Source Websites → Extraction → State Comparison → Structured Proposal → Human Review → Production Publish Human review is intentional. The goal is to automate repetitive discovery and data entry, not eliminate human oversight. IMPORTANT TIMING CONTEXT Most target websites are seasonal and are not expected to expose their current production data for approximately 1–2 months. We have screenshots and reference material from last year’s versions of the sites, which provide useful examples of likely layouts, navigation patterns, and data presentation. Initial development should therefore focus on: + Reusable ingestion framework + Standard scraper-adapter contract + Synthetic fixtures + Source-state logic + Delta detection + Source-health and partition-health handling + Failure handling + Idempotency + Proposal generation + End-to-end test harness Once the live target sites begin appearing, the developer will adapt the site-specific Playwright adapters to actual production behavior and convert important real-world cases into regression fixtures. The project therefore has two major stages: + Build and validate the core system now using fixtures and reference material + Perform rapid live-source adaptation and hardening as the real sites launch IMPLEMENTATION PHASES 1. Shared integration foundation + Define ingestion and proposal contracts + Establish source-state storage + Establish dependency and matching interfaces + Confirm Firebase integration boundaries 2. Playwright framework and test harness + Build reusable execution framework + Define site-adapter contract + Build synthetic fixtures + Implement health validation, retries, idempotency, and delta detection 3. Synthetic end-to-end vertical slice + Prove one complete workflow against deterministic fixture data + Validate extraction, state comparison, proposal generation, review, and non-production publication behavior 4. First live production source + Adapt the framework to the first representative live source + Validate actual navigation, DOM behavior, dynamic loading, access behavior, and real-world extraction 5. Representative source expansion + Expand to several structurally different sources + Use real failures and layouts to improve regression coverage 6. Remaining source rollout + Add adapters for the remaining target sources + Tune source-specific matching, health rules, cadence, and maintenance behavior 7. Stabilization and handoff + Final testing + Documentation + Regression coverage + Operational handoff BUDGET Fixed-price budget: $4,000–$6,000, divided into milestones. PLEASE SEE SCREENING QUESTIONS Please keep your proposal concise and specific to this project. In your proposal, include your fixed-price estimate within the stated $4,000–$6,000 budget and an approximate timeline broken into: + Core framework and test harness + Synthetic end-to-end vertical slice + First complete live source + Expansion to approximately 5 representative sources + Remaining adapters to approximately 30 total sources + Final stabilization, documentation, and handoff We are looking for an individual freelancer, not an agency. The detailed technical architecture brief will be provided to shortlisted applicants. Generic scraping proposals, proposals that do not answer the screening questions, and applicants whose expected budget is substantially above the stated range will not be considered.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Intermediate
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
TypeScript
Firebase Cloud Firestore
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:2 seconds ago
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Aug 16, 2010
  • United States
    Los Angeles11:59 PM
  • $238K total spent
    38 hires, 8 active
  • 5,388 hours

Explore similar jobs on Upwork

Docker
Python
PostgreSQL
HIPAA
Healthcare IT
Cloud Computing
Software Architecture & Design
DevOps

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo