AWS/Python Data Engineer — Secure Student Data Storage + Pseudonymization MVP

Posted yesterday

Worldwide

Summary

Unlock Education is an education analytics company working with U.S. public schools and districts. We analyze student-level assessment and instructional data and need to establish a lightweight, secure workflow for handling FERPA-protected student records. We already have an active AWS environment. We are not looking to build a new cloud platform or data lake. We need an experienced AWS/Python engineer to extend our existing environment with a small, well-designed secure data area and a simple pseudonymization utility. Our immediate goal is: District secure SharePoint → restricted AWS storage → pseudonymization → analysis-ready files The engagement should take approximately 5–10 hours and leave us with a simple system we can operate ourselves. SCOPE 1. Review our existing AWS environment Briefly review the current AWS account/configuration and determine the cleanest way to add secure storage for district student data without disrupting our existing application infrastructure. We expect this will primarily involve S3, IAM and appropriate encryption/logging rather than new application infrastructure. 2. Create three access-controlled data areas Set up logically and technically separated storage for: Raw / Restricted — original district files as received, potentially containing district student IDs, Florida Student IDs and other direct identifiers. Identity / Crosswalk — mapping between district identifiers and randomly generated Unlock analytical IDs. This should have the most restrictive access. Analysis — pseudonymized datasets containing only the fields needed for analysis. Access should be technically enforced through AWS IAM. Routine analysts should be able to access Analysis without access to Raw or Identity. We want a simple implementation appropriate to our current scale—not enterprise infrastructure. 3. Build a lightweight Python pseudonymization/preprocessing utility Create a documented Python script that: • reads CSV/XLSX district files; • identifies district student IDs; • generates a cryptographically random opaque Unlock student ID for each new student; • maintains the same Unlock ID for that student across files and subsequent runs; • maintains the district-ID ↔ Unlock-ID crosswalk in the restricted Identity area; • replaces district student IDs with Unlock IDs in analytical outputs; • removes configured direct identifiers such as student name, district ID, Florida Student ID, DOB, email, etc.; • optionally supports the same approach for teacher and section identifiers; • writes pseudonymized files to the Analysis area. We do not want student tokens generated using a simple deterministic hash of the original ID. The script should be configuration-driven enough that we can specify which columns are identifiers/removable fields without rewriting the code for every district. 4. Add basic validation/QA The process should report basic checks such as: • input/output row counts; • unique student counts; • new versus existing IDs mapped; • missing/null student IDs; • duplicate or problematic IDs; • removed fields; • output files generated. The purpose is simply to give us confidence that pseudonymization has not broken joins or unexpectedly altered the analytical data. 5. Test using synthetic data The contractor should not require access to actual student records. We will provide synthetic files approximating our district schemas. The solution should be developed and demonstrated using those records. After handoff, an authorized Unlock user will run the workflow against the real files. ACCESS MODEL At minimum we want: Access to Raw/Identity/Analysis layers for Unlock data custodian/admin. But we want access restricted to Analysis layer for Analyst + Future AI/Analysis Roles. We are not asking for Claude/Anthropic integration as part of this engagement. DELIVERABLES At the end of the engagement, we should have: 1. Secure Raw, Identity and Analysis storage within our existing AWS environment. 2. IAM policies/roles enforcing the access boundaries. 3. Appropriate S3 encryption/public-access/security settings. 4. Lightweight Python pseudonymization utility. 5. Stable random student-token/crosswalk functionality. 6. Configurable direct-identifier removal. 7. Basic QA output. 8. Synthetic test demonstrating the full workflow. 9. Short README explaining how to run the process. 10. Simple one-page architecture diagram. 11. 30-minute handoff/walkthrough. The workflow should ultimately be simple enough for us to run ourselves, e.g.: python pseudonymize.py --project sjc --input ./incoming EXPLICITLY OUT OF SCOPE We do not need a data lake, warehouse, database, VPC redesign, automated SharePoint ingestion, Airflow, Snowflake/Redshift, dashboards, web application, Claude/AI integration, hosted analytical environment, enterprise SSO, or full FERPA compliance audit. We are deliberately building a minimum viable secure workflow that can mature later. IDEAL BACKGROUND We're looking for someone senior enough to make good security decisions without overengineering the solution. Strong experience with: • AWS S3 • IAM / least-privilege access • AWS encryption/KMS • Python/pandas • secure data processing • pseudonymization/tokenization Experience with FERPA, HIPAA, education, healthcare or other sensitive-data environments would be particularly valuable. We value simple, secure and well-documented implementation over enterprise complexity.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $80.00

    -

    $150.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
Data Engineering
Amazon Web Services
Activity on this job
  • Proposals:50+
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Feb 19, 2021
  • United States
    St. Augustine2:10 AM
  • $5.9K total spent
    13 hires, 1 active
  • 169 hours

Explore similar jobs on Upwork

Data Governance- Atlan, Unity CatalogHourly‐ Posted 4 weeks ago
Data Engineering
Data Engineer for API PipelinesHourly‐ Posted 5 days ago
Python
Data Integration
Database Architecture
Data Transformation
ETL Pipeline
Data Preprocessing
SQL
Database Design
Data Engineering
Data Migration
API Development
Tableau
Google Cloud Platform

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo