Voice AI Expert to Make Our Voice Bot Faster & More Human (Fish Audio + Kokoro)

Posted 4 days ago

Worldwide

Summary

Overview We run a product with a real-time AI voice — It works and handles real calls today. However We need an expert voice-AI engineer to obsess over two things: making it sound genuinely human, and making it respond instantly. Everything else is secondary. The mission: humanness + speed Make it human. Callers should forget they're talking to a machine — natural prosody, warmth, and pacing from our TTS; clean barge-in/interruption handling and endpointing so it never talks over people or leaves dead air; natural confirmation of spelled-out emails and phone numbers; no stiff "AI" phrasing. Make it fast. We're around ~5s voice-to-voice and want to get under 2s without hurting naturalness. We have a tuning plan (VAD/endpointing, LLM prefix caching + speculative decoding, STT decode params, TTS warm-up/streaming, pipeline parallelism) — validate, execute, and push it further. Every 500ms cut makes it feel more alive. Our current TTS — this is what you'll tune We are standardizing on two TTS engines and want an expert to get maximum humanness and minimum latency out of these specifically (not to swap in new ones): Local / self-hosted Kokoro TTS on our own GPU (NVIDIA A6000) — low-latency, no per-minute cost. Fish Audio TTS (s2.1-pro) — our most expressive, human-sounding voice. Direct, hands-on Fish Audio and/or Kokoro experience is a major plus — knowing their params, streaming behavior, warm-up, and reference voices. Rest of the stack Pipecat + Daily WebRTC → Whisper STT + Qwen via vLLM → local Kokoro. LiveKit Agents → Inworld STT + LLM → Fish Audio or Kokoro (per call). Python · asyncio · aiohttp · Docker · Postgres (call/turn metrics). Must-have experience Shipped real-time, low-latency voice agents to production — made one sound human AND respond fast, and can prove it with numbers. Deep knowledge of the STT → LLM → TTS streaming pipeline and where both the milliseconds and the naturalness come from. Hands-on with Fish Audio and/or Kokoro (or very close analogues), plus Pipecat, LiveKit Agents, Daily, WebRTC, Whisper/faster-whisper, vLLM. Strong Python + asyncio; comfortable operating self-hosted GPU inference (CUDA/VRAM on an A6000). You measure both speed and human-ness — p50/p95 latency, TTFB budgets, and structured naturalness scoring. Nice to have Prosody/expressiveness & voice-clone/reference-audio tuning · receptionist persona prompt engineering · barge-in/interruption UX · frustration/sentiment detection & human escalation · staged production rollouts. Please answer when you apply Describe one time you made a voice agent sound more human or respond faster — before/after numbers and how you measured it. Your hands-on experience with Fish Audio and/or Kokoro (or the closest engines you've shipped), and what you tuned to improve naturalness and TTFB. In 2–3 sentences: how would you get a ~5s voice-to-voice loop under 2s while making it sound more human, not less? Your availability, hourly rate, and time-zone overlap with US hours. Tip: We read applications that engage with the actual problem far more closely than generic ones. Please skip the templated cover letter — talk concretely about prosody, latency budgets, and Fish/Kokoro tuning.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $20.00

    -

    $40.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
AI Text-to-Speech
AI-Generated Audio
Activity on this job
  • Proposals:50+
  • Last viewed by client:2 days ago
  • Interviewing:
    10
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Jan 15, 2021
  • Philippines
    City Of Cebu12:17 PM
  • $139K total spent
    24 hires, 6 active
  • 10,726 hours
  • Large company (100-1,000 people)

Explore similar jobs on Upwork

Text to speechHourly‐ Posted 4 weeks ago
AI-Generated Audio
AI-Generated Music
Music

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo