Data & AI Solutions Engineer | Lead Generation & Web Data
I build data pipelines and AI-powered tools that turn messy public web data into clean, decision-ready assets — and when it makes sense, into RAG-powered agents that answer questions from that data.
What I solve:
• Lead Generation at Scale — prospect databases with verified contacts (names, emails, phones, LinkedIn), enriched and deduplicated, ready for your sales team.
• Market & Competitive Intelligence — pricing monitoring, product catalogs, review mining, market research.
• Document Intelligence — parsing complex PDFs (tables, formulas, mixed-language) into structured Excel/CSV, and into chunked, embeddable formats for RAG.
• AI Agents & RAG Pipelines — knowledge-base Q&A agents (WhatsApp, web, internal tools) on vector databases; document ingestion → chunking → embeddings → retrieval → LLM answer, with moderation and audit layers.
• Anti-Bot & Hard Targets — Cloudflare, AWS-WAF, aggressive rate limiting: I know when to engineer around it and when to tell you it's not worth it.
How I work:
• Feasibility-first: I tell you what's realistic before you commit — including when the answer is "don't do this."
• Accuracy over volume: every record is verified or clearly flagged. No fabricated data, ever.
• Documented & reusable: scripts, schemas, pipelines you can run again without me.
• AI done right: generated content is moderated and human-reviewed — I don't ship hallucination-prone outputs.
Selected outcomes:
• Built a 50,000+ record physician directory from publicly available health registries, deduplicated and URL-verified — delivered as a structured database for client's internal use.
• Processed 60,000+ facility records (clinics, hospitals, labs) from an open government registry, with ~85% phone and ~75% email completeness — cleaned, normalized, and export-ready.
• Extracted 15,000+ product reviews from a Cloudflare-protected e-commerce site in 3 days with dual-pass validation.
• Delivered a 5,000+ record Google Maps enrichment pipeline (phone/website/email matching, 23-28% verified-match rate).
• Processed formula-heavy, bilingual PDFs into structured Excel — eliminating days of manual re-entry.
Skills: lead generation, prospect list, B2B data, list building, contact enrichment, data scraping, web scraping, Python, Playwright, Selenium, API integration, RAG, vector databases, PDF parsing, data cleaning, data mining, market research
Languages: Fluent English & Chinese.
Message me with your use case. I'll reply within 24 hours with a feasibility assessment and a realistic plan — including what I can't do, so you never waste budget on false promises.
Data Extraction
Web Scraping
PDF Conversion
Image Processing
OCR Algorithm
Computer Vision
API Integration
Selenium
Automation
AI Agent Development
B2B Lead Generation
Rajib B.
Rajbari, Bangladesh
$10/hr
5.0
10 jobs
Are you looking for a reliable expert to scrape Instagram data efficiently and accurately? You’re in the right place!
I specialise in Instagram scraping and automation, delivering clean, structured data from one of the most challenging platforms. With 3+ years of hands-on experience, I’ve developed custom solutions that bypass common limitations, including dynamic content, login restrictions, and bot detection.
💼 Instagram Scraping Services I Offer:
🔹 Scrape followers/following lists
🔹 Scrape public post data (captions, likes, hashtags, etc.)
🔹 Scrape profile info (bio, website, follower count)
🔹 Extract post/comment metadata from profiles or hashtags
🔹 Monitor and update data over time (daily/weekly scraping)
🔹 Handle automated login/session-based scraping
🔹 Save data in structured formats for analytics
🛠️ Tools & Technologies:
Python-based scraping:
Selenium, Playwright, Beautiful Soup, undetected-chromedriver, requests
Advanced Features:
Multi-profile rotation & proxy support
Mobile and desktop user-agent spoofing
Scroll automation for complete list extraction
Session management to reduce blocks
Output Formats:
.csv, .xlsx, .json, .sql, .txt
✅ Why Work With Me?
✔️ Proven success with Instagram scraping projects
✔️ Familiarity with anti-bot mechanisms and workarounds
✔️ Fast turnaround & clean, customizable scripts
✔️ Long-term support for recurring or real-time scraping tasks
Let me help you extract the Instagram data you need—reliably and at scale. Whether it’s for research, marketing, analytics, or lead generation, I can build the perfect solution tailored to your goals.
📩 Send me a message and let’s talk about your project!
DeepCrawl
Scrapy
Beautiful Soup
Data Extraction
Automation
Web Crawling
Data Mining
Data Scraping
Python Script
Browser Automation
Data Collection
Tesseract OCR
Muttayyab A.
Abbottabad, Pakistan
$10/hr
5.0
1 jobs
I help businesses cut manual work and unlock value from their data through AI-powered automation, intelligent document processing, and custom LLM solutions.
As an AI Engineer with hands-on production experience, I've built systems that deliver measurable results — including an OCR + LLM document processing pipeline that reduced manual data entry by 70%, and semantic product-matching systems using vector embeddings that automated reconciliation across messy, inconsistent datasets.
I specialize in turning complex, repetitive processes into reliable automated workflows — whether that's extracting structured data from invoices and forms, building RAG-based chatbots that answer questions from your private documents, or designing end-to-end n8n pipelines that scrape, analyze, and organize data from sources like Google Maps, TikTok, Instagram, and CSV files into actionable insights.
What I Can Do For You:
Document & Data Automation — Build OCR pipelines (PaddleOCR, AWS Textract) combined with LLMs to extract, validate, and structure data from invoices, receipts, and forms, cutting manual processing time dramatically.
RAG Systems & AI Chatbots — Design retrieval-augmented chatbots using LangChain, ChromaDB/FAISS, and models like GPT, Claude, or local LLMs (Mistral) for accurate, document-grounded conversations.
AI Workflow Automation — Create end-to-end n8n automations integrating OpenAI/Anthropic APIs, web scraping, transcription (Whisper), and data storage into Sheets or databases.
Machine Learning & Forecasting — Build LSTM/GRU time-series models and ensemble classifiers (XGBoost, Random Forest) for prediction, classification, and signal detection, with FastAPI deployment for real-time use.
Computer Vision — Develop CNN-based classification and YOLO object detection systems with transfer learning on custom datasets.
Semantic Search & Embeddings — Implement vector similarity search (FAISS, ChromaDB) for product matching, content discovery, and intelligent search experiences.
Core Tech Stack:
Languages: Python, JavaScript
AI/ML: TensorFlow, PyTorch, Scikit-Learn, Keras, Hugging Face Transformers, OpenAI API, Claude
NLP & RAG: LangChain, Embeddings, Vector Similarity, FAISS, ChromaDB
Automation: n8n, API Integration, Workflow Automation
Computer Vision: OpenCV, YOLO, Object Detection, Image Segmentation
MLOps: Docker, FastAPI
Web: MERN Stack (MongoDB, Express, React, Node.js), RESTful APIs
Why Work With Me:
Proven results — Real production systems with quantifiable impact (70% reduction in manual workload).
End-to-end delivery — From data pipeline design to deployment, I handle the full stack.
Clear communication — Regular updates, transparent progress, and responsive collaboration.
Quality-focused — Production-ready code following software engineering best practices (reinforced through code review work evaluating AI-generated solutions for top AI labs).
Let's discuss your project — I'm ready to help automate your processes, build intelligent systems, and turn your data into results.
Python
TensorFlow
Machine Learning
Exploratory Data Analysis
Matplotlib
Seaborn
Deep Learning
Computer Vision
Keras
Natural Language Processing
Large Language Model
Docker
LangChain
Google Cloud Platform
FastAPI
Dung N.
Ho Chi Minh City, Vietnam
$8/hr
5.0
2 jobs
I design and build AI-powered automation systems n8n workflows, LLM agents, and custom Python automation that eliminate repetitive work and connect your business data directly into the tools you already use. Backed by 5+ years of Python engineering and a 100% Job Success Score across 30+ delivered projects.
✅ What I Do Best
⚡ AI & n8n Automation (core expertise)
- Building AI automation workflows with n8n, OpenAI, and Claude APIs
- AI agents for lead capture, CRM sync (GoHighLevel and similar), and customer support
- AI-powered document processing, classification, and summarization
- Debugging and fixing existing n8n workflows payload mapping, webhook, and API errors
- Automation combined with Make, Zapier, and custom webhooks
- RAG pipelines, LangGraph, CrewAI for multi-agent systems
🕸️ Web Scraping & Data Extraction (feeding automation pipelines)
- Python: Playwright, Scrapy, Selenium, SeleniumBase, Crawl4AI, BeautifulSoup, Scrapling
- Scraping dynamic, JavaScript-heavy, login-protected, and anti-bot websites (Cloudflare bypass, CAPTCHA handling)
- Delivered: Glassdoor company data (Cloudflare bypass), Capterra reviews, social media data extraction (400+ hours)
🤖 Automation, Bots & Browser Extensions
- Custom Python bots for monitoring, notifications, reporting
- Chrome Extensions, Telegram/Discord bots, webhook integrations
- Excel, Google Sheets, Airtable, CRM automation
🔗 APIs & Backend
- REST API & GraphQL integrations, FastAPI, Flask, Django
- Connecting SaaS platforms, CRMs, databases, and third-party services
☁️ Cloud & Deployment
- AWS EC2, Docker, Linux servers, GitHub Actions
- PostgreSQL, MySQL, MongoDB
✅ Why Clients Hire Me
- Hands-on experience designing production AI automation workflows (n8n, LangGraph, CrewAI, RAG) as a working AI Automation Engineer
- 5+ years Python engineering background, 100+ scrapers/bots/automation systems delivered
- 100% Job Success, Top Rated, fast 0-4h response time
- Clean, documented, production-ready code with honest scoping I'll tell you if automation isn't the right fit before we start
If you have a repetitive process, messy data pipeline, or need an AI agent/chatbot built into your business, message me happy to review your workflow and turn it into a production-ready automation system.
n8n
AI Agent Development
OpenAI API
Automation
Automated Workflow
API Integration
Python
CRM Automation
LangChain
Bot Development
Web Scraping
Selenium
Scrapy
FastAPI
Django
PostgreSQL
Docker
Google Chrome Extension
Browser Automation
Chatbot Development
Prashant P.
Pune, India
$10/hr
4.9
190 jobs
AI Automation & Python Engineer — production AI agents, browser automation, RAG systems, data extraction pipelines, and backend applications that solve real business workflows.
I build production-ready automation and AI systems, not just prototypes.
My work combines Python, Playwright, FastAPI, OpenAI/Claude/Gemini, RAG, n8n, APIs, databases, and cloud infrastructure to automate workflows that would otherwise require hours of repetitive manual work.
I have 8+ years of experience with Python, web scraping, browser automation, APIs, data processing, and backend development, and I now specialize heavily in AI-powered automation.
WHAT I BUILD
→ AI Agents & Business Automation
AI agents that can research, extract data, process documents, make decisions, call APIs, update databases, generate reports, and trigger downstream workflows.
Tech includes OpenAI, Claude, Gemini, n8n, Make, Python, FastAPI, MCP, structured outputs, tool calling, and custom APIs.
→ RAG & Knowledge-Base Chatbots
Production RAG systems that answer questions from private company documents and knowledge bases.
Typical architecture:
Documents → Parsing → Chunking → Embeddings → Vector Database → Retrieval → LLM → Citations
I work with FAISS, pgvector, Supabase/PostgreSQL, OpenAI embeddings, Claude, Gemini, and custom retrieval pipelines.
→ Browser Automation & Web Scraping
Complex automation for authenticated portals, dashboards, document downloads, data collection, and repetitive browser workflows.
I work extensively with:
Playwright • Selenium • Scrapy • requests • aiohttp • BeautifulSoup • lxml
Including dynamic JavaScript websites, pagination, authenticated sessions, downloads, retries, proxy infrastructure, and anti-bot environments.
→ Document & PDF Automation
Systems that turn PDFs, receipts, statements, images, spreadsheets, and business documents into structured data.
Examples include:
PDF/Image → OCR/Vision → Extraction → Validation → Structured JSON/CSV/Excel → Database/API
I build workflows for invoices, receipts, financial documents, reports, forms, and large document collections.
→ Python / FastAPI Backend Systems
Production APIs and backend services using:
Python • FastAPI • Django • Flask • Celery • Redis • PostgreSQL • MySQL • MongoDB • Docker
I can own the complete backend — API design, asynchronous processing, database architecture, background jobs, integrations, logging, testing, and deployment.
→ n8n / Make / API Automation
I automate workflows across CRMs, Google Sheets, Google Drive, email, databases, AI models, internal APIs, and third-party SaaS platforms.
When low-code is enough, I use it.
When reliability or complexity requires code, I build the missing pieces in Python.
RECENT SYSTEMS I’VE BUILT
AI Healthcare Research Platform
Built a complete pipeline for collecting public patient discussions, extracting symptoms, interventions and outcomes with AI, normalizing thousands of signals, maintaining source-level traceability, and presenting the results through an interactive dashboard.
Bank Document Automation
Built Python + Playwright workflows for authenticated financial portals where the user handles login/MFA and automation takes over afterward — discovering accounts, downloading historical PDFs, organizing documents, detecting duplicates, and maintaining audit manifests.
AI Document Extraction Platform
Built systems that accept PDF, Excel and image documents, extract structured information using traditional parsing + OCR/Vision/LLMs, validate the results, and export clean datasets.
Large-Scale Web Data Pipelines
Built and maintained scraping infrastructure across hundreds of websites with proxy rotation, retry chains, anti-bot handling, automated QA, structured extraction, and production monitoring.
AI-Powered Data & Reporting Systems
Built structured datasets and semantic layers designed specifically for LLM reasoning, reporting, search, and downstream AI applications.
MY CORE STACK
AI
OpenAI • Claude • Gemini • RAG • AI Agents • MCP • Function Calling • Structured Outputs • Embeddings • Vector Search
Automation
Playwright • Selenium • Scrapy • n8n • Make • Apify
Backend
Python • FastAPI • Django • Flask • Celery • Redis
Data
PostgreSQL • Supabase • MySQL • MongoDB • Pandas • FAISS • pgvector
Frontend
React • Next.js • Tailwind • Streamlit
Infrastructure
Docker • AWS • GCP • Linux • Nginx • GitHub Actions • CI/CD
WHY CLIENTS HIRE ME
A lot of automation projects work perfectly in a demo and fail as soon as they encounter unexpected data, a changed webpage, an API timeout, or a long-running job.
I design around those failures from the beginning.
That means:
Retry and recovery strategies
Structured logging
Duplicate prevention
Validation and QA
Error reporting
Maintainable code
Human-in-the-loop workflows where appropriate
Documentation and handover
Production deployment support
Data Extraction
Web Scraping
Selenium
Web Crawling
Data Mining
Data Annotation
Supabase
Claude
OpenAI Codex
Web Scraping Software
AI Agent Development
Python
TypeScript
AWS Development
Data Collection
Machine Learning
Data Science
Artificial Intelligence
Prompt Engineering
LLM Prompt Engineering
Barkilign M.
Addis Ababa, Ethiopia
$15/hr
5.0
2 jobs
Hello, I'm Barkilign Mulatu,
A Data Scientist and AI/ML Engineer specializing in building production-ready AI systems. From developing Advanced RAG pipelines that "chat with your data" to fine-tuning Multilingual NLP models for niche markets, I bridge the gap between complex research and business value.
My approach isn't just about training models; it's about building end-to-end systems that are accurate, interpretable and scalable.
Core Areas of Expertise:
Generative AI & RAG: Building custom chatbots using LangChain, OpenAI, and Open-Source LLMs. Expert in optimizing retrieval with ChromaDB, FAISS, and Cross-Encoder re-ranking.
Natural Language Processing (NLP): Specialized in Named Entity Recognition (NER) and text classification, including multilingual support for languages like Amharic (XLM-Roberta/mBERT).
Predictive Analytics & FinTech: Developing credit risk models and scoring systems using XGBoost, Random Forest, and Basel II compliance standards.
Full-Stack AI Deployment: Turning models into interactive tools using Streamlit and FastAPI.
Technical Toolbox:
Languages: Python (Pandas, NumPy, Scikit-learn)
AI Frameworks: PyTorch, Hugging Face, LangChain
Databases: ChromaDB, FAISS, PostgreSQL
Tools: Streamlit, Docker, Git, Telethon (Scraping)
My Goal: To provide clear, documented, and high-performing AI solutions that solve your specific challenges—whether that’s automating customer insights or building a custom financial risk engine.
Let's hop on a quick message to discuss how we can bring your AI project to life!
Python
Machine Learning
Large Language Model
Generative AI
Natural Language Processing
Automation
Retrieval Augmented Generation
n8n
LangChain
Data Science
Vector Database
Hugging Face
Streamlit
Ecommerce
Python Scikit-Learn
Data Visualization
Predictive Modeling
API Integration
Data Scraping
Web Scraping
How it works
Post a job for freePost a job
Tell us what you need. Create your own job post or generate one with AI then filter talent matches.
Hire top talent fast
Consult, interview, and hire quickly, so you can meet the freelancers you're excited about.
Collaborate easily
Use Upwork to chat or video call, share files, and track project progress right from the app.
Payment simplified
Manage payments in one place with flexible billing options. Only pay for approved work, hourly or by milestone.
Don't just take our word for it
“Upwork provides an umbrella-level of security. I can see a talent’s work history and ratings. I can hold payments in escrow. I can communicate through Upwork Messages instead of working through my email address.”
KD
Kim Darling
Emerald Tiger
“Upwork is the best platform to hire skilled professionals when we're not looking for a full-time employee. All the companies in our portfolio use Upwork to find talent across a wide range of fields.”
DM
David Merry
Kinetic Investments
“Our very specific requirements can be a challenge—With Upwork, we’re able to access a bigger community to ensure the success of our projects.”
KK
Katja Krohn
Summa Linguae
How do I hire a DeepCrawl Freelancer on Upwork?
You can hire a DeepCrawl Freelancer on Upwork in four simple steps:
Create a job post tailored to your DeepCrawl Freelancer project scope. We’ll walk you through the process step by step.
Browse top DeepCrawl Freelancer talent on Upwork and invite them to your project.
Once the proposals start flowing in, create a shortlist of top DeepCrawl Freelancer profiles and interview.
Hire the right DeepCrawl Freelancer for your project from Upwork, the world’s largest work marketplace.
At Upwork, we believe talent staffing should be easy.
How much does it cost to hire a DeepCrawl Freelancer?
Rates charged by DeepCrawl Freelancers on Upwork can vary with a number of factors including experience, location, and market conditions. See hourly rates for in-demand skills on Upwork.
Why hire a DeepCrawl Freelancer on Upwork?
As the world’s work marketplace, we connect highly-skilled freelance DeepCrawl Freelancers and businesses and help them build trusted, long-term relationships so they can achieve more together. Let us help you build the dream DeepCrawl Freelancer team you need to succeed.
Can I hire a DeepCrawl Freelancer within 24 hours on Upwork?
Depending on availability and the quality of your job post, it’s entirely possible to sign up for Upwork and receive DeepCrawl Freelancer proposals within 24 hours of posting a job description.