Senior GCP Data Architect | Audit & Simplify BigQuery/GCS Platform

Posted 5 days ago

Worldwide

Summary

We are looking for a Senior GCP Data Architect / Principal Data Engineer to independently audit and simplify an existing quantitative research/data platform. This is not initially a development project. The first engagement is a paid 4–8 hour read-only architecture assessment. Our platform has evolved through: DuckDB → SQL Server → local Windows storage → GCS → BigQuery We also use Python, SQL, Dagster, Dataform, Git/GitHub and Parquet. The platform works and contains a substantial data estate, but it has grown through multiple iterations. We now need someone senior enough to understand the whole system, determine what is actually authoritative, identify unnecessary complexity/duplication, and design the simplest sensible target architecture. We specifically want someone willing to say: “You don't need this layer.” “These datasets are duplicates.” “This should be authoritative in GCS.” “SQL Server/DuckDB no longer serves a purpose.” “This model has a grain or temporal-join problem.” “Do not rebuild this—simplify it.” We are not looking for someone who automatically recommends more infrastructure. The data SYSTEM35 is a quantitative research platform containing: longitudinal SERP/ranking observations from multiple providers; thousands of domains, queries and URLs; repeated HTML/web-crawl observations; internal-link/graph data; semantic/custom extraction data; Google Search Console; provider/API datasets; experimental/change-event data; statistical/ML research outputs. Different providers have very different domain coverage, query coverage, cadence, historical depth and missingness. For example, 600,000 observations across 20 domains are analytically very different from 300,000 observations across 15,000 domains. A critical requirement: point-in-time correctness Our observations occur at different times. Example: Ranking observation: 10 June HTML crawl: 7 June Next HTML crawl: 14 June A prediction made on 10 June may use the 7 June state. It must never use 14 June. We therefore strongly value experience with: temporal/bitemporal data modelling; as-of / latest-valid-prior joins; event time vs ingestion/system time; SCD Type 2 / effective-dated models; late-arriving/corrected observations; point-in-time ML feature engineering; longitudinal/panel data; leakage prevention. Experience with causal inference, experiments, event studies or survival/transition analysis is advantageous because the data platform ultimately supports research, not just dashboards. What we want from the first 4–8 hours We expect a first-pass architecture assessment, not a production-quality audit of every table. You should inspect the actual systems and establish enough evidence to recommend the target direction. Please cover: Current architecture — what actually exists and how it connects. Sources of truth — what should be authoritative for each major data family. Duplication/technical debt — unnecessary copies, layers, pipelines and dependencies. Canonical grains/identities — domain, query, URL, provider, observation, crawl, etc. Temporal correctness — whether historical/ML datasets can be built without future leakage. Target architecture — what should remain in GCS, BigQuery, SQL Server, local storage, Dagster/Dataform, etc. Retirement plan — KEEP / CONSOLIDATE / MIGRATE / DEPRECATE / DELETE LATER. ML architecture recommendation — BigQuery ML, Vertex AI, Python or another justified approach. Prioritised implementation roadmap with approximate effort. A plain-English explanation for a non-technical owner. An exhaustive table-by-table audit is not expected in 4–8 hours. Where deeper investigation is required, identify it and estimate the additional work. Who we want Strong experience in: Data Architecture · Data Modelling · GCP · BigQuery · GCS · SQL · Python · Data Engineering · Data Warehousing · Data Lineage Particularly valuable: Bitemporal/Temporal Modelling · SCD2 · Point-in-Time ML · Dagster · Dataform/dbt · SQL Server · DuckDB · Parquet · BigQuery ML · Vertex AI · Longitudinal/Panel Data Search/SEO expertise is not required. Please do not apply if your main experience is: basic SQL/reporting; dashboards/BI; SEO/WordPress; generic AI-agent development; simply implementing cloud tools without architecture ownership. We want someone who has inherited complicated data estates before and simplified them. Screening questions Please answer these directly. 1. The same underlying data exists across SQL Server, local files, GCS and BigQuery. How would you determine which copy should become authoritative? 2. A dataset has 650,000 observations across 20 domains; another has 350,000 across 15,000 domains. Why does this matter when designing analytical/ML datasets? 3. A ranking observation is dated 10 June and HTML snapshots exist on 7 June and 14 June. Which can be used, and how would you make the join point-in-time correct? 4. A provider observation represents 15 August, arrives on 20 August and is corrected on 25 August. How would you preserve both: what we now believe happened on 15 August; and what a model could actually have known on 22 August? 5. What is the difference between temporal ordering and causal identification? Give an example where X occurring before Y still does not establish X caused Y. 6. Our architecture evolved through DuckDB → SQL Server → local storage → GCS → BigQuery. What evidence would you examine to decide whether SQL Server, DuckDB or local storage still deserve a place in the active architecture? 7. Describe one data platform you personally simplified. What did you remove? 8. What would you realistically expect to establish during the first 4–8 hours of read-only access? This is initially a small paid discovery engagement. If the assessment demonstrates that you understand the estate and can simplify it, there is substantial follow-on implementation work available.

  • Less than 30 hrs/week
    Hourly
  • < 1 month
    Duration
  • Expert
    Experience Level
  • Remote Job
  • One-time project
    Project Type
Skills and Expertise
Mandatory skills
Data Engineering
Data Integration
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:4 days ago
  • Hires:
    2
  • Interviewing:
    9
  • Invites sent:
    16
  • Unanswered invites:
    2
About the client
Member since Jun 26, 2023
  • United Kingdom
    Bolton9:05 AM
  • $39K total spent
    86 hires, 15 active
  • 728 hours

Explore similar jobs on Upwork

Varicent/SPM Consultant for GuidanceHourly‐ Posted 4 weeks ago
SAP
ETL
Databricks PySpark YML IntegrationHourly‐ Posted 3 weeks ago
Databricks Platform
Apache Spark
PySpark
YAML

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo