AI / Prompt Engineer: LLM Benchmark & Extraction Test Harness for Legal DPAs (Fixed Price)

Posted 2 weeks ago

Worldwide

Summary

Project Overview: We are building a secure B2B software tool that automates data-privacy record-keeping for corporate legal and compliance teams. We are building an official audit log called a Record of Processing Activities (ROPA). This log details every software vendor used across the business, what types of personal data they handle, where that data is stored, and how long it is retained. Often, corporate privacy teams read 40-page vendor contracts and Data Processing Agreements (DPAs) by hand and manually type these data points into spreadsheets. Our app automates this: users upload vendor contracts, our AI engine extracts key compliance details into a structured table, highlights missing information, and provides exact page citations so a human can quickly review and approve the record. Before building our full web application, we are executing a 3-day technical benchmark. We need an experienced AI engineer to set up an automated test harness (using Promptfoo, Braintrust, or custom Python scripts) to test and optimize system prompts against our gold-standard dataset of 25 real-world legal DPAs. Goal: Achieve over 95 percent accuracy on extracted regulatory fields with zero percent hallucinations on missing fields. Successful completion of this micro-project may lead to an invitation to lead the full platform build (10,000 to 20,000 USD budget). Key Deliverables: 1. Evaluation Test Harness Setup: Configure a test harness running Claude 3.5 Sonnet and OpenAI GPT-4o with temperature set to zero. 2. Prompt Optimization and JSON Schema Enforcement: Refine system prompts using structured outputs to extract strict JSON adhering to our ICO ROPA master template. 3. Citation and Missing-Data Handling: Ensure the prompt extracts the exact paragraph or page source for cited data points and outputs explicit null or NA for unstated fields without guessing. 4. Accuracy Dashboard and Report: Run all 25 PDF DPAs through the pipeline, compare outputs against our human-verified ground-truth dataset, and deliver a pass or fail accuracy matrix. Technical Requirements: - Proven experience with LLM orchestration in Python or TypeScript. - Deep familiarity with Structured Outputs, Pydantic or JSON schemas, and temperature control for deterministic extraction. - Experience evaluating LLM performance using automated benchmarks. Intellectual Property: All prompt files, test scripts, schemas, and configurations created under this project are Work Made for Hire and belong exclusively to the client upon payment. You will be required to sign a standard IP Assignment and Non-Disclosure Agreement prior to receiving the benchmark dataset. Screening Questions (Please answer in your proposal): 1. How do you prevent an LLM from guessing or hallucinating missing data fields when enforcing a rigid JSON output schema? 2. How do you structure system prompts or post-processing to force the LLM to return exact paragraph or page citations alongside extracted values? 3. What framework or tool would you use to measure accuracy across our 25 PDF test dataset?

  • $400.00

    Fixed-price
  • Expert
    Experience Level
  • Remote Job
  • One-time project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Machine Learning
Artificial Intelligence
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:yesterday
  • Interviewing:
    2
  • Invites sent:
    2
  • Unanswered invites:
    0
About the client
Member since Aug 15, 2026
  • United Kingdom
    1:25 AM

Explore similar jobs on Upwork

Artificial Intelligence
Job Search Strategy
LinkedIn Recruiting
Resume Writing
Cover Letter Writing
Applicant Tracking Systems
Career Coaching
Virtual Assistance
Data Entry
EEG based emotion recognitionHourly‐ Posted 4 weeks ago
Machine Learning
Python
Keras
PyTorch
Deep Neural Network

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo