Svetlana isn't taking new orders for this project right now. Here are some similar projects to explore.
You will get a custom Python data collection tool for public or authorized sources
Top Rated

Top Rated

Project details
I will build a custom Python data collection tool and deliver a clean, structured dataset in CSV, Excel, or JSON. This service is limited to publicly available data, official APIs, and client-owned sources with documented authorization.
I extract data from business directories, e-commerce catalogs, public records, location sources, restaurant listings, search forms, APIs, sitemaps, PDFs, and JavaScript-rendered websites.
For each source, I choose the simplest reliable approach: direct HTTP requests, APIs, Scrapy, aiohttp, BeautifulSoup, Playwright, Camoufox, or Selenium. I use browser automation only when the site requires it.
I can handle pagination, infinite scrolling, data normalization, duplicate removal, missing-record detection, retries, checkpoints, and recovery after interruptions.
For larger projects, I send an initial batch after the contract starts so you can approve the fields, structure, and quality before full collection.
Before ordering, send the target website, required fields, approximate volume, and preferred output format. Pricing covers public or client-authorized data from websites with a reasonably consistent structure.
I extract data from business directories, e-commerce catalogs, public records, location sources, restaurant listings, search forms, APIs, sitemaps, PDFs, and JavaScript-rendered websites.
For each source, I choose the simplest reliable approach: direct HTTP requests, APIs, Scrapy, aiohttp, BeautifulSoup, Playwright, Camoufox, or Selenium. I use browser automation only when the site requires it.
I can handle pagination, infinite scrolling, data normalization, duplicate removal, missing-record detection, retries, checkpoints, and recovery after interruptions.
For larger projects, I send an initial batch after the contract starts so you can approve the fields, structure, and quality before full collection.
Before ordering, send the target website, required fields, approximate volume, and preferred output format. Pricing covers public or client-authorized data from websites with a reasonably consistent structure.
Data Tool
PythonWhat's included
| Service Tiers |
Starter
$90
|
Standard
$200
|
Advanced
$500
|
|---|---|---|---|
| Delivery Time | 3 days | 5 days | 7 days |
Number of Pages Mined/Scraped | 10000 | 20000 | 30000 |
Number of Sources Mined/Scraped | 1 | 1 | 3 |
Number of Revisions | 1 | 2 | 3 |
Optional add-ons
You can add these on the next page.
Additional Page Mined/Scraped
(+ 1 Day)
+$5
Additional Source Mined/Scraped
(+ 5 Days)
+$90
Additional Revision
+$20Frequently asked questions
20 reviews
(19)
(1)
(0)
(0)
(0)
This project doesn't have any reviews.
JO
Johnathan O.
Mar 18, 2026
Data mining for Turkey
AB
Abdallah B.
Mar 8, 2026
Shopify theme edit
JO
Johnathan O.
Oct 14, 2025
Scraping data
Great person to work with, I will definitely hired her again
PR
Pascual R.
Sep 24, 2025
Python Developer for Web Scraping Chilean News Websites into Excel Database
Svetlana is a great professional, she took the time to analyse carefully the request before accepting in order to make sure that she could handle the job and/or how much time it would take. She ended up doing it perfectly and timely, always communicating advances and clarifying everything in order to develop exactly what I needed. I defenitely can trust in her and I would work with her again.
JO
Johnathan O.
Aug 10, 2025
Scrape 6 websites
Reliable, great communication, I will definitely hire here again
About Svetlana
Python Automation, Data Scraping Expert
97%
Job Success
Astana, Kazakhstan - 8:12 am local time
I build custom Python scrapers and data pipelines for websites where extensions, no-code tools, and simple scripts are not enough.
My projects cover directories, e-commerce, healthcare, public records, location data, restaurant listings, news archives, and document extraction.
I manage each project from collection through final delivery.
SELECTED RESULTS
• Processed 134,000+ restaurant listing URLs for a national dataset
• Consolidated 43,000+ healthcare professionals and organizations from 26 sources across four countries
• Processed 116,000+ company-record rows using website extraction, PDF processing, and OCR
• Collected 35,000+ e-commerce products with prices, descriptions, images, and source URLs
• Prepared products, variants, images, pricing, descriptions, and metafields for Shopify import
SERVICES
• Business directories and company databases
• Products, prices, SKUs, variants, and inventory
• Location, healthcare, restaurant, and hospitality data
• Real estate, court, probate, and public records
• News and publication archives
• Search forms, filters, pagination, and profile pages
• XML sitemaps and large URL collections
• Public APIs and client-authorized website endpoints
• Authorized login-based workflows
BROWSER AUTOMATION
I work with websites that require JavaScript, infinite scrolling, browser sessions, cookies, interactive forms, or multi-step navigation.
My stack includes Playwright, Camoufox, Selenium, Selenium Wire, Scrapy, Requests, aiohttp, curl_cffi, and BeautifulSoup. I use direct requests when possible and browser automation when required.
PDF AND OCR
I download documents, extract text from PDFs, process scanned pages, apply OCR, locate fields with validation rules and regular expressions, and combine document data with website records.
My document stack includes Tesseract, OpenCV, pdf2image, Pillow, and NumPy.
DATA QUALITY
• Deduplication
• Address and phone normalization
• Missing-record detection and source validation
• Country and location filtering
• Multi-source consolidation and consistent schemas
• Recovery passes for failed records
• Translation with local caching
• Checkpoints and periodic saving
DELIVERY
• Excel, CSV, and JSON
• Shopify-compatible CSV
• Files prepared for database import
For long collections, I use asynchronous requests, multiprocessing, retries, caching, checkpoints, crash recovery, and chunk-based processing. Temporary errors do not force a full restart.
I work with publicly available data and client-authorized sources, using access methods appropriate for each project.
HOW I WORK
I inspect the source and choose the most reliable collection method.
I confirm the fields, volume, output format, and whether you need a one-time dataset or reusable code.
For larger projects, I deliver a small initial batch after the contract begins so you can verify the structure and data quality before full collection.
I build the scraper and check the output against the agreed requirements.
I deliver reusable Python code and setup instructions when included in the project scope.
GOOD FIT
• Large datasets
• JavaScript websites or interactive forms
• Scrapers that stop or miss records
• Extraction combined with data cleanup
• Reusable Python code
Send me the target website. I will review the source and recommend a practical approach before development begins.
Feedback from one of my client:
📍"Svetlana is AMAZING! I have a fair amount of data background and I am very picky about data quality. She delivers excellent quality data in a timely fashion. I spot checked the data that she delivered and it all checked out well. In addition, when we hit an issue, she offered proactive and creative solutions. I highly recommend her. She is your data wizard, don’t go any further!"
Mary, US/ Jul 2025
I am always available to discuss all the details of your project!
Steps for completing your project
After purchasing the project, send requirements so Svetlana can start the project.
Delivery time starts when Svetlana receives requirements from you.
Svetlana works on your project following the steps below.
Revisions may occur after the delivery date.
Source review and scope confirmation
I review the source, required fields, volume, output format, and technical complexity.
Data sample and schema approval
I prepare an initial batch so you can approve the schema and data quality.