Korean Plain Text Dataset Needed for AI Training

Posted 2 weeks ago

Worldwide

Summary

We are looking for large-scale, original Korean text datasets for AI and LLM training. We welcome both ready-made datasets and vendors who can source and process eligible Korean text data. 📌 ACCEPTED CONTENT CATEGORIES 📚 Literature ▪️ Web novels ▪️ Fiction and fan fiction ▪️ Traditional stories ▪️ Literary works 💬 Forums ▪️ Complete discussion threads ▪️ Original posts and associated replies ▪️ Content reflecting natural Korean language and local usage ✍️ Blogs ▪️ High-quality personal blog posts ▪️ Long-form articles ▪️ Personal essays and opinion pieces ✅ STRICT DATA REQUIREMENTS ▪️ Content must have been originally written in Korean. ▪️ Machine-translated, manually translated, AI-generated, or synthetic content is not accepted. ▪️ Delivery format must be plain text in UTF-8 encoding. ▪️ PDFs, scanned documents, screenshots, images, and OCR-extracted files are not accepted. ▪️ Data must be clean, readable, and directly usable. ▪️ Duplicate and near-duplicate content should be removed or clearly identified. ▪️ The provider must be able to explain the source and collection method. ▪️ Applicants must confirm that the proposed data can legally be collected and provided for the intended use. ▪️ Source URLs or provenance records should be available upon request. 📩 HOW TO APPLY Please provide the following information: 1️⃣ Content categories you can provide: literature, forums, blogs, or a combination. 2️⃣ Whether the data is ready-made or needs to be newly sourced. 3️⃣ Estimated raw data volume in GB / Estimated usable token volume after cleaning and deduplication. 4️⃣ Proposed rate per GB or per 1 million usable tokens / Estimated lead time. 5️⃣ A representative plain-text sample of approximately 1–5MB, or a shareable Google Drive link. ⚠️ IMPORTANT Please do not submit machine-translated, AI-generated, OCR-extracted, duplicated, or unverifiable datasets. We are looking for vendors who can support large-volume delivery. Partial-volume proposals are also welcome.

  • More than 30 hrs/week
    Hourly
  • 6+ months
    Duration
  • Entry level
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
Web Scraping
Korean
Data Collection
Activity on this job
  • Proposals:Less than 5
  • Last viewed by client:2 weeks ago
  • Interviewing:
    3
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Aug 28, 2024
  • South Korea
    Seoul4:59 AM
  • $119K total spent
    376 hires, 292 active
  • Tech & IT
    Mid-sized company (10-99 people)

Explore similar jobs on Upwork

True Crime FootageHourly‐ Posted 3 weeks ago
Data Extraction
Video Editing
Research & Development
ETL
Airtable
Database Design
Automation
No-Code Development
Data Migration
Zapier
Customer Relationship Management

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo