Databricks Engineer - Audience & Campaign Data
Worldwide
We build and deliver the person-level data behind large marketing campaigns — the files that determine who gets contacted, who gets suppressed, who lands in which test cell, and what comes back after the fact. Those pipelines run in production today. Parts of them are solid Databricks work; other parts still depend on large Excel workbooks where campaign rules accumulated over time. We're moving the whole thing toward an automated, repeatable framework while campaigns continue to ship on schedule. You'd be working on both sides of that at once. The work *Audience & segmentation logic* Eligibility rules, prioritization, deduplication, suppression. Test/control cells, waves, and the final production files that go to print and outbound vendors. Tracing a rule from a spreadsheet through to the SQL or Python that enforces it, and confirming it still does what it was meant to do. *Pipelines & platform* Databricks notebooks, jobs and workflows. PySpark at scale. Snowflake alongside it. Ingestion failures, partial SFTP transfers, silently failing scheduled jobs, permission drift — you own the diagnosis, not just the ticket. *Quality & release* Reconciliation, QC, counts that tie out, validation before anything ships. A wrong campaign file is expensive and public. We want someone whose instinct is to prove a file is correct rather than assume it. You'll also support downstream response matchback, attribution, and the reporting built on them. Requirements Databricks / PySpark — hands-on and recent Advanced SQL: complex joins, windowing, performance tuning Python Snowflake A demonstrable practice around data QC and reconciliation Also useful: Git and CI/CD, AWS, Linux and SFTP, and the patience to reverse-engineer business logic out of a 40-tab workbook. Strong plus — experience with individual-level marketing or communications data: direct mail and production file prep, large third-party consumer datasets, person- and household-level identity resolution, address hygiene, seed records, holdouts and control groups, suppression and eligibility processing, matchback and attribution. Healthcare payer experience helps but isn't required; comparable scale and regulatory weight from another industry counts. When you apply, answer these 1. Describe a pipeline you didn't build but ended up owning. What did it do, what did you find once you were inside it, and what did you personally change? 2. A person-level campaign file is about to be released. How do you confirm it's safe to send? Name the specific checks you'd run to catch a broken rule, an unexpected duplicate, a suppression that didn't apply, or counts that don't reconcile. 3. Where your Databricks/PySpark and SQL work has been most substantial. 4. Your experience with large person-level datasets, marketing data, or regulated data.
- Less than 30 hrs/weekHourly
- < 1 monthDuration
- ExpertExperience Level
$5.00
-
$20.00
Hourly- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:15 to 20
- Last viewed by client:3 days ago
- Interviewing:4
- Invites sent:3
- Unanswered invites:0
About the client
- ITAAgira3:12 PM
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by