Competitor Intelligence Engine Developer

Posted yesterday

Worldwide

Needs to hire 2 Freelancers
Summary

Python developer — scale a competitor-monitoring engine (scraping, FastAPI, Apify) Comment the codeword in if you have read the full article. The codeword is hidden in the description below About us & the product: We're a Dutch online-marketing agency building an internal intelligence platform for our e-commerce clients. One module monitors the webshops of each client's competitors: it periodically snapshots their pages, diffs each snapshot against the previous one, and surfaces only meaningful changes as findings in an inbox — price moves, assortment changes, delivery-terms changes, positioning shifts, and CRO/conversion-element changes. Each finding gets an AI-written interpretation and concrete action points. This is not a greenfield project. The module is live in pilot and the hard parts exist: a pure, unit-tested diff engine with noise suppression (timestamps, stock counters, session IDs and A/B variants are stripped before comparison), deterministic structured-data extraction (JSON-LD → microdata → dataLayer → OpenGraph, merged per product, with an LLM fallback only for interpretive fields), idempotent change fingerprinting so re-runs never duplicate findings, multi-tenant row-level security, and cost tracking on every scrape and LLM call. The problem you'll solve: Today the scanner follows only ~7 hand-seeded pages per competitor, picked by arbitrary sitemap order — while these shops have catalogs of hundreds to thousands of products. Every page is fetched with a full browser render, one run at a time, which creates both a cost and a throughput ceiling. Our working plan: split monitoring into a broad cheap layer (whole catalog via batched plain-HTTP crawls + deterministic extraction, no LLM, stored as price/product time series) and a narrow expensive layer (a few key funnel pages via browser render + LLM, feeding the findings inbox). About that plan — and your role in it: The plan is worked out in detail, code-verified, and you'll receive the full document with acceptance criteria per step. But it was written by a small team and we know the difference between a plan and the truth: milestone 0 exists specifically to test its core assumptions against real data, and every milestone boundary is a point where we revise based on what we learned. We want a developer who challenges the plan where it's weak — if you see a better route, make the case with evidence and we'll change course. What we don't want is the opposite: silently building around something you disagree with. Disagreement goes on the table, before the code. One thing genuinely is settled: the stack. This is a production codebase, so we're not migrating frameworks or rewriting the engine — proposals need to work within it. Milestones (each independently shippable, fixed price per milestone; scope of later milestones is revisited as findings come in) M0 — Measurement day (1 day, paid trial). Run four defined measurements against our data and one pilot competitor — including batch-fetching 100 product URLs from a sitemap and reporting what percentage yields name/price/EAN through our existing extractor — plus a small, precisely specified config fix with extended tests. Explicitly a go/no-go: if the numbers contradict the plan, we adjust the plan. This milestone doubles as the paid trial for both sides. M1 — The broad layer (3–5 days). Scale an existing production-proven batch-crawl pattern to full catalogs: sitemap harvesting, batched cheerio crawls, deterministic extraction, per-hostname rate limiting and block-page detection, crawl budgets per client tier, provenance-flagging so auto-discovered products don't pollute hand-curated data. M2 — Smarter page selection (1–2 days). Deterministic ranking instead of "first 3 in sitemap order", plus audit events when discovery silently finds nothing. M3 — Sitemap as change oracle (2–3 days). Parse lastmod + full URL sets; URL-set diffing as a whole-catalog assortment signal. Additive only, with sanity checks — sitemaps lie. M4 — CRO detection (3–4 days). On HTML we already fetch: element counting (form fields, payment icons, trust badges), a normalized DOM-skeleton hash for restructure detection, A/B-test detection via repeated sampling with pinned proxy geo, urgency-signal normalization. M5 — Throughput & monitoring (1–2 days). Higher concurrency for the cheap layer, scheduling that fits ~40 clients, alerting on overruns. Codeword: apple pie Stack: Python 3.12, FastAPI, SQLAlchemy 2.0 async, Alembic, Postgres with RLS, Redis + RQ, APScheduler, Apify (cheerio & Playwright crawlers), Anthropic API behind provider interfaces, pytest, GitHub Actions. Conventions: boring over clever, tests ship with the change, small PRs per milestone. What we provide: repo access, the detailed plan document with code-level references, scraping budget for dry-runs, and a responsive technical owner (the codebase author) for questions and same-day review. One rule has a story behind it: an error page was once interpreted as "competitor removed everything", so production baselines are sacred here — changes near the diff engine need a written case first, never a quiet workaround. Who we're looking for: strong async Python with real production scraping experience — you know why JSON-LD beats LLM extraction on product pages, what a 200-OK block page is, and why cheerio-vs-headless is a cost decision, not a preference. You push back with arguments when you disagree, commit small and push daily, flag blockers early, and you're transparent about how you work — including which AI coding tools you use and where you draw the line. Dutch is not required; the target pages are Dutch-language webshops, so comfort with non-English content matters. Sizing: 11–17 working days total. M0 first as the paid trial; continuation per milestone after review.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Intermediate
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
Marketing
Software
Python
Nice-to-have skills
Agile Software Development
Activity on this job
  • Proposals:10 to 15
  • Last viewed by client:2 hours ago
  • Interviewing:
    5
  • Invites sent:
    16
  • Unanswered invites:
    7
About the client
Member since Jan 8, 2026
  • Netherlands
    Volendam1:55 PM
  • $125 total spent
    1 hire, 1 active

Explore similar jobs on Upwork

JavaScript
Node.js
PHP
Web Application
AI App Development
DevOps
API
Git
MySQL
Cs2 Gambling SiteFixed-price‐ Posted 4 weeks ago
Gambling
Unity
Counter Strike
AR & VR
Online Gambling Website
Card Game
Board Game
Unreal Engine
MetaMask
Mystery Box
iGaming
WebGL
Game Development
Gaming
Multiplayer
Game UI/UX Design
UI/UX Prototyping
Steam API
AI Development
PixiJS

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo