I build modular Python data collection and ETL systems for public or authorized business data sources. I extract, normalize, validate, deduplicate, and enrich records from websites and APIs, then deliver reliable outputs to Google Sheets, CSV files, SQLite, or PostgreSQL.
I completed a paid Upwork project covering five public data sources, a shared schema, multi-key deduplication, Google Sheets export, validation, tests, documentation, and source-code handoff.
I can help with:
• Python web scraping and multi-source data collection
• Data extraction from websites and APIs, with CSV, Excel, or PDF processing when required
• Data cleaning, normalization, validation, and deduplication
• Google Sheets, CSV, Excel, SQLite, and PostgreSQL output
• API data enrichment and scheduled ETL pipelines
• Source tracking, error handling, tests, and documentation
My typical workflow uses a separate adapter for each source, a shared normalized schema, deterministic validation rules, and quality checks before export. This keeps the pipeline maintainable and makes future sources easier to add.
For new projects, I prefer starting with one representative source and a small paid milestone. We confirm access, required fields, output quality, and edge cases before expanding the full pipeline.
Please include the target sources, required fields, sample output, update frequency, and preferred delivery format.
Scrapy
Python
Web Scraping
Data Extraction
Data Scraping
Data Cleaning
Google Sheets
API Integration
PostgreSQL
ETL Pipeline
Data Mining
Beautiful Soup
pandas
Microsoft Excel
Data Processing
Yu Hang L.
Chongqing, China
$25/hr
4.9
17 jobs
Anyone can scrape a page. The hard part — and what I do — is turning messy market data into something you can actually trade on: clean pipelines, point-in-time-correct databases, and bots that survive running 24/7.
I build trading-adjacent data systems end to end:
- Market data & pipelines — login-gated and JS-heavy sources, multi-source collection with failure isolation, normalized into Postgres/DuckDB
- Point-in-time correctness — bitemporal databases with as-of replay and zero look-ahead leakage, so your backtests aren't lying to you
- Signals & automation — rules-based market scanners, exchange API integration, Lightning/crypto automation, crash-resilient 24/7 deployment
Recent work (real, with code & tests):
- A bitemporal horse-racing trading database — Betfair historical pricing (600K+ price rows) + 280K+ race-card index, as-of replay validated leak-free, 176 passing tests (PostgreSQL/SQLAlchemy)
- A Bitcoin (RoboSats) selling-automation system — 18-state order lifecycle, Lightning payments via Core Lightning, Tor routing, Docker API
- A prediction-market arbitrage radar — scans 5,000+ live markets per cycle for YES/NO mispricing, DuckDB snapshots, 17 passing tests
I care about correctness — bad data quietly ruins a trading system, so I build in reconciliation, idempotency and tests. Clear communication in English and Chinese. I'll tell you upfront if an edge isn't realistic rather than overpromising.
Scrapy
Python
Web Scraping
Selenium
REST API
API Integration
Automation
Data Extraction
Browser Automation
pandas
Task Automation
ETL
Python Script
PostgreSQL
Data Engineering
Time Series Analysis
API Development
Financial Analysis
Bitcoin
SQL
Zhendong L.
Huizhou, China
$20/hr
5.0
9 jobs
🟩 Overview
I help businesses turn complex, login-protected websites into clean, structured datasets ready for analysis.
I specialize in extracting reliable data from dynamic websites, including pages with anti-bot protection and heavy JavaScript rendering.
Recently completed projects:
Healthcare data project: extracted 11,000+ records from a login-protected platform
Investment dataset: structured 17,000+ Crunchbase records into clean CSV format
I focus on delivering accurate, structured, and ready-to-use data — not raw or incomplete outputs.
🟨 What I deliver
Clean, structured data (CSV / Excel / JSON)
Reliable extraction from dynamic and protected websites
Handling of login, pagination, and complex page structures
Fast and consistent delivery (typically 24–48 hours depending on complexity)
Scripts may be delivered depending on project requirements.
🟧 Tools & Technologies
Python (Scrapy, Playwright, Requests, BeautifulSoup)
Dynamic websites (JavaScript / React / Vue-based pages)
Data extraction and cleaning workflows
Automation for repetitive data collection tasks
🟦 Sample & Process
I can provide a free sample (5–10 rows) within 12 hours after reviewing the target website, so you can evaluate the data quality before starting the full project.
🟩 Why me
✔ 100% Job Success Score
✔ Experience with complex and dynamic websites
✔ Structured, analysis-ready datasets (not raw scraping output)
✔ Fast and reliable communication
✔ Focus on accuracy and long-term usability
🟥 Next step
Send me your website — I’ll quickly check feasibility and provide a sample before we start.
Scrapy
Web Scraping
Python
Data Cleaning
pandas
Microsoft Excel
Automation
Data Processing
JSON
CSV
Data Mining
Selenium
Jackey C.
Fuzhou, China
$20/hr
5.0
7 jobs
Data & AI Solutions Engineer | Lead Generation & Web Data
I build data pipelines and AI-powered tools that turn messy public web data into clean, decision-ready assets — and when it makes sense, into RAG-powered agents that answer questions from that data.
What I solve:
• Lead Generation at Scale — prospect databases with verified contacts (names, emails, phones, LinkedIn), enriched and deduplicated, ready for your sales team.
• Market & Competitive Intelligence — pricing monitoring, product catalogs, review mining, market research.
• Document Intelligence — parsing complex PDFs (tables, formulas, mixed-language) into structured Excel/CSV, and into chunked, embeddable formats for RAG.
• AI Agents & RAG Pipelines — knowledge-base Q&A agents (WhatsApp, web, internal tools) on vector databases; document ingestion → chunking → embeddings → retrieval → LLM answer, with moderation and audit layers.
• Anti-Bot & Hard Targets — Cloudflare, AWS-WAF, aggressive rate limiting: I know when to engineer around it and when to tell you it's not worth it.
How I work:
• Feasibility-first: I tell you what's realistic before you commit — including when the answer is "don't do this."
• Accuracy over volume: every record is verified or clearly flagged. No fabricated data, ever.
• Documented & reusable: scripts, schemas, pipelines you can run again without me.
• AI done right: generated content is moderated and human-reviewed — I don't ship hallucination-prone outputs.
Selected outcomes:
• Built a 50,000+ record physician directory from publicly available health registries, deduplicated and URL-verified — delivered as a structured database for client's internal use.
• Processed 60,000+ facility records (clinics, hospitals, labs) from an open government registry, with ~85% phone and ~75% email completeness — cleaned, normalized, and export-ready.
• Extracted 15,000+ product reviews from a Cloudflare-protected e-commerce site in 3 days with dual-pass validation.
• Delivered a 5,000+ record Google Maps enrichment pipeline (phone/website/email matching, 23-28% verified-match rate).
• Processed formula-heavy, bilingual PDFs into structured Excel — eliminating days of manual re-entry.
Skills: lead generation, prospect list, B2B data, list building, contact enrichment, data scraping, web scraping, Python, Playwright, Selenium, API integration, RAG, vector databases, PDF parsing, data cleaning, data mining, market research
Languages: Fluent English & Chinese.
Message me with your use case. I'll reply within 24 hours with a feasibility assessment and a realistic plan — including what I can't do, so you never waste budget on false promises.
Data Extraction
Web Scraping
PDF Conversion
Image Processing
OCR Algorithm
Computer Vision
API Integration
Selenium
Automation
AI Agent Development
B2B Lead Generation
Xiaowei L.
Guangzhou, China
$25/hr
5.0
1 jobs
I help e-commerce stores, businesses, and operators turn manual, time-consuming data tasks into fast, reliable, and error-free automated workflows.
Whether you need to extract thousands of product listings from supplier websites, batch-download and rename image assets for auction/e-commerce catalogs, or sync messy data across platforms, I deliver clean, structured results.
How I can help your business:
• E-Commerce Data Extraction & Web Scraping:
Extract product details, prices, specs, and variants from any protected or dynamic website using Python (Playwright/Scrapy/Selenium), Bright Data, or custom automated scrapers.
• Catalog Preparation & Asset Processing:
Format and map bulk inventory directly into target marketplace templates (eBay, LiveAuctioneers, Shopify, Zoro, Amazon). Automatically batch-download, organize, and rename thousands of high-res images to exact file-naming conventions in minutes.
• Data Cleaning & Automated Formatting:
Transform messy, inconsistent raw exports, PDFs, and spreadsheets into clean, structured CSV/Excel files with automated quality control checks to ensure zero missing fields or formatting errors.
• Workflow Automation & API Integrations:
Connect data across Google Sheets, Airtable, CRM systems, and e-commerce backends so your repetitive tasks run automatically.
Why work with me:
• Fast & Accurate Turnaround: Focused on 100% data integrity with no errors.
• Simple Handoff: I deliver clean end-results (ready-to-use CSVs/organized assets) or lightweight automated scripts that you can easily run with one click.
• Responsive & Clear Communication.
Send over your target website, sample CSV/spreadsheet, or workflow description—let's make your data workflow fast and effortless.
Data Extraction
ETL
ETL Pipeline
Artificial Intelligence
Machine Learning
Data Mining
Python
Web Scraping
Data Cleaning
Data Processing
API Integration
Automation
Google Apps Script
pandas
Microsoft Excel
Seemab Y.
Changsha, China
$15/hr
5.0
12 jobs
I am Seemab - a detail-oriented AI Research Engineer who utilizes the power of AI to solve complex problems.
I specialize in developing efficient custom AI, Data, and software solutions. With a strong foundation in Python programming, I excel at playing with AI models and evaluating them for specific use cases. Proven experience in cloud technologies like Server-less computation, Linux Servers, and databases.
Skills: AI/ML Model, Python Engineer, N8N Workflows, Automation, Data Mining
Projects:
OCR WebApp
Snowflake warehouse remodeling.
Education: BS in Computer Science
Scrapy
Python
Data Scraping
Selenium
Data Processing
Data Manipulation Language
Data Engineering
Automation
Data Extraction
ETL Pipeline
Spreadsheet Automation
Web Scraping
SQL
DevOps
pandas
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a Scrapy Developer in China on Upwork?
You can hire a Scrapy Developer in China on Upwork in four simple steps:
Create a job post tailored to your Scrapy Developer project scope. We'll walk you through the process step by step.
Browse top Scrapy Developer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top Scrapy Developer profiles and interview.
Hire the right Scrapy Developer for your project from Upwork, the world's largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a Scrapy Developer?
Rates charged by Scrapy Developers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a Scrapy Developer in China on Upwork?
As the world's work marketplace, we connect highly-skilled freelance Scrapy Developers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream Scrapy Developer team you need to succeed.
Can I hire a Scrapy Developer in China within 24 hours on Upwork?
Depending on availability and the quality of your job post, it's entirely possible to sign up for Upwork and receive Scrapy Developer proposals within 24 hours of posting a job description.