TypeScript / LLM Engineer — Conversational AI Agent (Sports Betting)
Worldwide
We’re building a production conversational AI that texts members over iMessage/SMS about sports betting — picks, odds, live scores, and ongoing banter. The agent has a distinct voice, tool calling into our own data, multi-turn memory, and a golden-eval loop. We need an engineer to deepen conversation quality, tool grounding, and the reply pipeline. This is backend / LLM product work, not a greenfield chatbot. The system already ships; you’ll extend and harden it. What you’ll work on Conversation quality — brevity, honesty, temporal grounding (“today” vs stale memory), multi-turn coherence, group-chat behavior Tool use & grounding — odds board, live scores, player lookup, recall tools; stop invented numbers; keep web search for gaps only Persona / prompts — system voice, prompt versioning, latency-aware reply lanes (fast vs full) Memory & personalization — thread memory, communication style, member model; privacy-safe updates Eval & regression — golden conversation scenarios + LLM judge; unit tests for pure helpers Inbound pipeline — webhooks → gates/routing → context assembly → agent loop → outbound bubbles; acks, rate limits, failure fallbacks Stack TypeScript on Deno (Supabase Edge Functions) Linq for messaging Postgres / Supabase Anthropic Claude (Sonnet for persona + tools; Haiku for routing/judges) iMessage/SMS webhook integrations Deno unit tests + scenario-based persona evals Primary surface is the conversation backend. iOS companion app exists; not the main focus. Requirements Strong TypeScript (async pipelines, testable pure helpers) Shipped LLM agents: system prompts, tool schemas, agent loops, eval harnesses Comfortable with Supabase/Postgres and serverless/edge functions Experience with multi-turn chat: routing, memory, latency UX Sports/betting domain familiarity (odds, slates, scores) is a strong plus Solid testing habits (unit + scenario evals) Nice to have SMS/iMessage or messaging webhook systems Prompt caching, voice/style systems, LLM-as-judge evals Deno experience Engagement Contract / part-time or full-time — open to the right fit Async-friendly; clear PRs against an existing codebase You’ll get access to the conversation modules, eval harness, and recent voice/tool work To apply Link to a shipped conversational or tool-calling LLM project (repo or product) Brief note on how you’ve measured or improved reply quality (evals, human review, etc.) Hours/week available and typical turnaround for a focused PR Your rate Please also type "penguin" in your letter.
- Not SureHourly
- < 1 monthDuration
- ExpertExperience Level
$20.00
-
$47.00
Hourly- Remote Job
- One-time projectProject Type
Skills and Expertise
Activity on this job
- Proposals:5 to 10
- Last viewed by client:5 days ago
- Hires:1
- Interviewing:1
- Invites sent:0
- Unanswered invites:0
About the client
- USANew York7:04 AM
- $11K total spent19 hires, 8 active
- 413 hours
- Tech & ITSmall company (2-9 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by