Competitor Intelligence Engine Developer
Worldwide
Python developer — scale a competitor-monitoring engine (scraping, FastAPI, Apify) Comment the codeword in if you have read the full article. The codeword is hidden in the description below About us & the product: We're a Dutch online-marketing agency building an internal intelligence platform for our e-commerce clients. One module monitors the webshops of each client's competitors: it periodically snapshots their pages, diffs each snapshot against the previous one, and surfaces only meaningful changes as findings in an inbox — price moves, assortment changes, delivery-terms changes, positioning shifts, and CRO/conversion-element changes. Each finding gets an AI-written interpretation and concrete action points. This is not a greenfield project. The module is live in pilot and the hard parts exist: a pure, unit-tested diff engine with noise suppression (timestamps, stock counters, session IDs and A/B variants are stripped before comparison), deterministic structured-data extraction (JSON-LD → microdata → dataLayer → OpenGraph, merged per product, with an LLM fallback only for interpretive fields), idempotent change fingerprinting so re-runs never duplicate findings, multi-tenant row-level security, and cost tracking on every scrape and LLM call. The problem you'll solve: Today the scanner follows only ~7 hand-seeded pages per competitor, picked by arbitrary sitemap order — while these shops have catalogs of hundreds to thousands of products. Every page is fetched with a full browser render, one run at a time, which creates both a cost and a throughput ceiling. Our working plan: split monitoring into a broad cheap layer (whole catalog via batched plain-HTTP crawls + deterministic extraction, no LLM, stored as price/product time series) and a narrow expensive layer (a few key funnel pages via browser render + LLM, feeding the findings inbox). About that plan — and your role in it: The plan is worked out in detail, code-verified, and you'll receive the full document with acceptance criteria per step. But it was written by a small team and we know the difference between a plan and the truth: milestone 0 exists specifically to test its core assumptions against real data, and every milestone boundary is a point where we revise based on what we learned. We want a developer who challenges the plan where it's weak — if you see a better route, make the case with evidence and we'll change course. What we don't want is the opposite: silently building around something you disagree with. Disagreement goes on the table, before the code. One thing genuinely is settled: the stack. This is a production codebase, so we're not migrating frameworks or rewriting the engine — proposals need to work within it. Milestones (each independently shippable, fixed price per milestone; scope of later milestones is revisited as findings come in) M0 — Measurement day (1 day, paid trial). Run four defined measurements against our data and one pilot competitor — including batch-fetching 100 product URLs from a sitemap and reporting what percentage yields name/price/EAN through our existing extractor — plus a small, precisely specified config fix with extended tests. Explicitly a go/no-go: if the numbers contradict the plan, we adjust the plan. This milestone doubles as the paid trial for both sides. M1 — The broad layer (3–5 days). Scale an existing production-proven batch-crawl pattern to full catalogs: sitemap harvesting, batched cheerio crawls, deterministic extraction, per-hostname rate limiting and block-page detection, crawl budgets per client tier, provenance-flagging so auto-discovered products don't pollute hand-curated data. M2 — Smarter page selection (1–2 days). Deterministic ranking instead of "first 3 in sitemap order", plus audit events when discovery silently finds nothing. M3 — Sitemap as change oracle (2–3 days). Parse lastmod + full URL sets; URL-set diffing as a whole-catalog assortment signal. Additive only, with sanity checks — sitemaps lie. M4 — CRO detection (3–4 days). On HTML we already fetch: element counting (form fields, payment icons, trust badges), a normalized DOM-skeleton hash for restructure detection, A/B-test detection via repeated sampling with pinned proxy geo, urgency-signal normalization. M5 — Throughput & monitoring (1–2 days). Higher concurrency for the cheap layer, scheduling that fits ~40 clients, alerting on overruns. Codeword: apple pie Stack: Python 3.12, FastAPI, SQLAlchemy 2.0 async, Alembic, Postgres with RLS, Redis + RQ, APScheduler, Apify (cheerio & Playwright crawlers), Anthropic API behind provider interfaces, pytest, GitHub Actions. Conventions: boring over clever, tests ship with the change, small PRs per milestone. What we provide: repo access, the detailed plan document with code-level references, scraping budget for dry-runs, and a responsive technical owner (the codebase author) for questions and same-day review. One rule has a story behind it: an error page was once interpreted as "competitor removed everything", so production baselines are sacred here — changes near the diff engine need a written case first, never a quiet workaround. Who we're looking for: strong async Python with real production scraping experience — you know why JSON-LD beats LLM extraction on product pages, what a 200-OK block page is, and why cheerio-vs-headless is a cost decision, not a preference. You push back with arguments when you disagree, commit small and push daily, flag blockers early, and you're transparent about how you work — including which AI coding tools you use and where you draw the line. Dutch is not required; the target pages are Dutch-language webshops, so comfort with non-English content matters. Sizing: 11–17 working days total. M0 first as the paid trial; continuation per milestone after review.
- Less than 30 hrs/weekHourly
- 1-3 monthsDuration
- IntermediateExperience Level
- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:10 to 15
- Last viewed by client:2 hours ago
- Interviewing:5
- Invites sent:16
- Unanswered invites:7
About the client
- NetherlandsVolendam1:55 PM
- $125 total spent1 hire, 1 active
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by