Dataset Curator
Worldwide
Overview We're building a labeled evaluation corpus for a PII detection system that scans files across many formats. We need a meticulous data researcher to find, verify, and organize files from open-licensed public datasets — images of identity documents, tax forms, medical records, structured data exports, code files, and logs. This is a research and organization role. You are not generating, creating, or modifying files. You are finding files that already exist under open licenses and delivering them in a structured, documented format. What You'll Do - Search open data repositories (Zenodo, HuggingFace Datasets, Kaggle, data.gov, SEC EDGAR, PhysioNet, IRS.gov, USCIS.gov, GitHub) for specific document types - Verify each file has an acceptable open license (CC0, CC-BY, or public domain) — document the source URL for every file - Rename files descriptively and organize into a structured folder hierarchy - Fill out a sources.csv for each file: path, format, document type, polarity (positive/hard negative), source URL, license, notes - Flag gaps when no licensed file can be found for a deliverable category — document what you searched What You'll Be Finding Organized into 12 deliverable categories across 4 phases: - Images (Phase 1 — highest priority): US driver's licenses (IDNet/Zenodo CC0), passports (MIDV-500/CC0), insurance cards, scanned financial documents, hard negative blanks - PDFs (Phase 2): Financial filings, medical records, blank government forms (IRS/USCIS — public domain) - Structured data (Phase 2): CSV/parquet datasets with personal identifiers - Code & logs (Phase 3): Open-licensed source files and log samples with technical identifiers We provide starting sources for each category — you're extending and filling gaps, not starting from scratch. Required Skills - Experience navigating academic and open data repositories (Zenodo, HuggingFace, Kaggle, government portals) - Solid understanding of Creative Commons licensing — you must verify licenses, not assume them - Strong organizational discipline: consistent file naming, complete CSV documentation, no gaps left undocumented - Comfort working with diverse file types: images, PDFs, CSVs, JSON, code, log files - Clear written English for README notes and gap reports Nice to Have - Familiarity with document datasets used for ML/AI (DocVQA, RVL-CDIP, IDNet, MIDV-500) - Basic Python or command-line experience (useful for batch downloading Zenodo ZIPs) - Understanding of privacy/PII concepts — helps judge whether a file is a positive or hard negative Deliverables Per Batch - ZIP of files with descriptive names organized by format - sources.csv — one row per file with all required columns - README.txt — batch notes, new sources discovered, gap report Scope: ~480–560 files across 12 deliverable categories, delivered in phased batches (batch 1 after S-01 through S-05 are complete). Ongoing engagement.
- Less than 30 hrs/weekHourly
- 1-3 monthsDuration
- IntermediateExperience Level
$8.00
-
$25.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:15 to 20
- Last viewed by client:3 days ago
- Interviewing:4
- Invites sent:4
- Unanswered invites:2
About the client
- USANew York1:32 AM
- $25K total spent10 hires, 5 active
- 1,757 hours
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by