Fix Broken Data Pipeline: Two Python/Node APIs Not Updating AWS RDS Postgres Database
Worldwide
I have two scraper/API jobs that are supposed to continuously pull data from two different websites and write it into a PostgreSQL database hosted on AWS RDS (Aurora/RDS instance called "fincrawl", running in us-east-2 / Ohio). The database has not received any new data in about a year, and I need an experienced engineer to diagnose and fix the pipeline so it's reliably populating the DB again — plus set up monitoring so this doesn't silently break again. Environment: Database: AWS RDS PostgreSQL instance ("fincrawl") in us-east-2 Compute: An EC2 instance ("FinCrawl") in the same VPC runs the two data-pull jobs/APIs Security groups: rds-ec2-1 (DB side) and ec2-rds-1 (EC2 side) control traffic between the two Access: I can provide AWS console access (or scoped IAM credentials), SSH/EC2 Instance Connect access to the EC2 instance, and DBeaver is already set up on my end for direct DB inspection What I need done: Diagnose why the two ingestion jobs stopped writing to the database — could be a crashed process, a cron/scheduler failure, a broken scraper due to a website layout change, an expired credential/API key, a network/security group misconfiguration between EC2 and RDS, disk/storage issue on the instance, or something else entirely Fix whatever is broken so both jobs reliably pull data from their respective source websites and write it into the correct tables in the fincrawl database Verify data is actually landing correctly — check for gaps, duplicate rows, or malformed data from whatever time the pipeline was last working Add basic monitoring/alerting (e.g., a simple health check, log review process, or alert via email/Slack) so if either job fails again, I find out within a day instead of a year Document what was wrong and what you changed, plus any recommendations for making the setup more resilient going forward (e.g., moving from a manually-run script to a managed scheduler, containerizing the jobs, adding retry logic, etc.) Nice to have: Experience with web scraping resilience (handling site structure changes, rate limiting, rotating IPs/user agents if relevant) Experience with AWS (RDS, EC2, security groups, IAM) and Postgres administration Comfort setting up lightweight monitoring/alerting (CloudWatch, cron + email, UptimeRobot, etc.) Deliverables: Working, verified data pipeline actively populating the fincrawl RDS database again Short written summary of root cause and fix Basic monitoring/alerting in place Any updated scripts/configs handed over with clear instructions to run/maintain them To apply, please include: A brief note on how you'd approach diagnosing a "silently stopped for a year" pipeline like this Relevant experience with AWS RDS/EC2 and/or web scraping pipelines Your estimated timeline and rate (fixed price)
- Less than 30 hrs/weekHourly
- < 1 monthDuration
- IntermediateExperience Level
- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:50+
- Last viewed by client:2 weeks ago
- Hires:1
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- CanadaToronto6:43 AM
- $10K total spent25 hires, 3 active
- 574 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by