Harden an AI scheduling assistant for production — Next.js/Supabase/TypeScript (evals, Playwright, guardrails)

Posted 2 weeks ago

Worldwide

Summary

Rebound is a live student planning app (Next.js 16, React 19, TypeScript, Zustand, Supabase/Postgres with owner-scoped RLS, Vercel, Vitest, Playwright). It has an AI assistant that proposes schedule actions from an LLM (DeepSeek via OpenRouter). Proposals are schema-validated and sanitized server-side, shown to the student, and only applied after explicit confirmation. Placement itself is deterministic engine code, not the model. I need a senior engineer to make that AI layer production-grade and prove it. WHERE THINGS STAND - A 60-case frozen evaluation dataset with a nine-dimension rubric exists in the repo. Baseline from a manual production run: 48.3% strict pass. Planner-scheduling category: 8 of 20 pass. - Four documented failure classes: (1) the assistant claimed it moved work when no action was proposed or applied; (2) long conversations answered a previous prompt or routed to the wrong source; (3) an instruction embedded in course material was reproduced inside generated study artifacts; (4) low-quality generated flashcards/quizzes. - Six release-blocking safety invariants, e.g. no schedule mutation without confirmation; no claim of a mutation without a committed action receipt; no invented deadline stated as fact. - There is no automated eval runner yet; the baseline was hand-run. That is the first gap. MILESTONE 1 — DIAGNOSTIC (est. 6–8 h) 1. Build a replay harness that sends the 20 planner fixtures to the live /api/assist route (restricted dev environment, my OpenRouter key) and scores proposed actions and no-invention automatically against the frozen expectations. Table output; cost per run noted. 2. Reproduce the 9 failing planner cases and write a failure taxonomy: prompt defect / validation gap / engine mismatch / model limitation, with evidence per case. 3. One PR: a Playwright spec proving invariants 1 and 2 end to end (no mutation without confirmation; no "I moved it" without a receipt). MILESTONE 2 — HARDEN (est. 20–30 h) - Extend the fixtures with adversarial planner cases: clock constraints, overlapping deadlines, locked/done blocks, long-conversation carryover, injection via source text, invented deadlines. - Server-side enforcement so a reply cannot assert a mutation unless an action receipt is present. Code checks, not prompt-only fixes. - Raise planner strict pass from 8/20 to at least 15/20 without changing the deterministic engine's contract or the confirm-before-commit flow. - Written report: exact commit SHA, before/after eval table, per-case diffs, residual risks. DEFINITION OF DONE The harness runs from one command and reproduces the numbers in your report. The Playwright spec is in CI and green. PRs pass typecheck, lint, unit and e2e. No test disabled or weakened. No new features. NON-NEGOTIABLES - Preserve the architecture. No rewrites, no platform migrations. - Keep confirm-before-commit for every AI-generated schedule change. - Owner-scoped database access only; fail closed. - No secrets in commits or logs. Production secrets are not shared; you get a restricted dev setup. - Never hard-code a model id; models are read from environment variables. - Work personally. Written, asynchronous communication only, no calls. Small PRs with evidence. HOW I WORK I use AI coding tools heavily myself, so I am not paying for keystrokes. I am paying for judgment, verification and honest reporting. If something cannot be reproduced, say so. TO APPLY Answer the three screening questions directly and link to a repo or PR containing Playwright or Vitest specs you wrote. Short, specific answers beat long ones.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $15.00

    -

    $19.00

    Hourly
  • Remote Job
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:2 weeks ago
  • Hires:
    1
  • Interviewing:
    7
  • Invites sent:
    4
  • Unanswered invites:
    0
About the client
Member since Aug 24, 2026
  • USA
    Austin10:12 PM
  • 1 hire, 1 active

Explore similar jobs on Upwork

OpenAI API
API Integration
Customer Relationship Management
Artificial Intelligence
Business Intelligence
WhatsApp Chatbot Marketplace DevelopmentFixed-price‐ Posted 4 months ago
PHP
JavaScript
HTML5
MySQL

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo