Speech Corpus Recording Consultant — Multi-Speaker, Multi-Mic Array Capture

Posted yesterday

Worldwide

Summary

Job Description We are building a large-scale, multi-speaker conversational speech dataset for a commercial AI/speech-separation use case (diarization and speaker-separation model training). We need an experienced audio/acoustic engineer to design our recording protocol, guide our in-house team through execution, and provide ongoing quality review as we scale data collection across multiple locations. This is not a music recording, mixing, or mastering project. We need raw, unprocessed, scientifically clean multi-channel capture — not a "good sounding" mix. Project Background - We are collecting real, co-located, multi-speaker conversational audio to serve as training/evaluation data for speech AI models (speaker diarization / source separation). The dataset must meet strict technical requirements around synchronization, signal isolation, and lack of post-processing, detailed below. Hard Technical Requirements (Non-Negotiable) - 2+ speakers, physically co-located in the same room — no remote calls, no replayed/re-recorded audio One close-talk microphone per speaker (headset, lapel, or binaural) — raw capture, with no noise suppression, AGC, or beamforming applied At least one far-field microphone (room mic or mic array) capturing the full mixed signal Sample-level synchronization across all channels — either a shared clock/single interface, or verified alignment (e.g., clap/tone sync) across independently clocked devices Natural conversational dynamics — spontaneous turn-taking, interruptions, natural overlapping speech (target overlap ratio ≥10%) Uncompressed PCM WAV output, 16-bit or 24-bit, consistent sample rate across channels (16kHz or 48kHz) Environmental variety — multiple room types/acoustic conditions, varying speaker-to-mic distances (1–4m for far-field), a mix of recording hardware (laptop mics, phone mics, studio mics, lapel mics, arrays) What We Need From You - Protocol design: A detailed, written recording protocol covering equipment selection/setup, mic placement guidelines, sync verification method, session structure, and a pre-session checklist our team can follow without prior audio engineering expertise. Equipment recommendations: Specific mic/interface/recorder models (with budget tiers) suited to close-talk + far-field simultaneous capture, appropriate for repeated use across varied rooms in India. Sync verification method: A concrete, repeatable way for our non-expert team to verify sample-level sync at the start/end of every session. QC criteria & rejection checklist: Clear pass/fail criteria (clipping, bleed thresholds, dropout, sync drift) our team can apply to every recorded session before it's accepted into the dataset. Team training: One or more live/recorded training sessions (video call or in-person if Ahmedabad-based) to walk our operations team through correct setup and common failure modes. Ongoing review (optional/extendable): Periodic review of sample recordings as we scale, with feedback on drift from spec. Ideal Candidate Background - -Experience with multi-channel/array microphone recording (not just stereo music tracking) -Background in speech corpus collection, ASR/voice-AI data collection, broadcast/film production sound (multi-lav + boom sync), or acoustic/bioacoustic research -Comfortable explaining why raw/unprocessed capture matters (i.e., understands downstream ML use, not just "clean audio" from a listening perspective) -Prior work with companies or labs in speech AI, telecom, linguistics, or corpus vendors (e.g., LDC/ELRA-style collection, or Indian voice-AI companies) is a strong plus -India-based or willing to work across IST hours; based in or able to travel to Ahmedabad is a plus but remote consulting is acceptable

  • Less than 30 hrs/week
    Hourly
  • 3-6 months
    Duration
  • Intermediate
    Experience Level
  • $10.00

    -

    $25.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Music Production
Audio Mastering
Activity on this job
  • Proposals:10 to 15
  • Last viewed by client:yesterday
  • Interviewing:
    9
  • Invites sent:
    29
  • Unanswered invites:
    14
About the client
Member since Oct 7, 2025
  • United States
    New York4:31 PM
  • $998 total spent
    64 hires, 35 active
  • 44 hours

Explore similar jobs on Upwork

Podcasts ConsultantFixed-price‐ Posted 3 weeks ago
Foley Effects
ADR
Communications
Podcast
Media Relations
English
Public Relations
Marketing Strategy
Report
Administrative Support
VO recording engineer neededHourly‐ Posted 2 weeks ago
Sound Mixing
Audio Production
Audio Mastering
Audiobook
Voice-Over Recording
Audio Editing
Audio Post Production
Voice-Over
Audio Engineering
Music & Sound Design
YouTube
Audio Restoration
Narration
Sound Design

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo