Vietnamese Plain Text Dataset Needed for AI Training

Posted 3 weeks ago

Worldwide

Summary

We are looking for large-scale, original Vietnamese text datasets for AI and LLM training. We welcome both ready-made datasets and vendors who can source and process eligible Vietnamese text data. 📌 ACCEPTED CONTENT CATEGORIES 📚 Literature ▪️ Web novels ▪️ Fiction and fan fiction ▪️ Traditional stories ▪️ Literary works 💬 Forums ▪️ Complete discussion threads ▪️ Original posts and associated replies ▪️ Content reflecting natural Vietnamese language and local usage ✍️ Blogs ▪️ High-quality personal blog posts ▪️ Long-form articles ▪️ Personal essays and opinion pieces ✅ STRICT DATA REQUIREMENTS ▪️ Content must have been originally written in Vietnamese. ▪️ Machine-translated, manually translated, AI-generated, or synthetic content is not accepted. ▪️ Delivery format must be plain text in UTF-8 encoding. ▪️ PDFs, scanned documents, screenshots, images, and OCR-extracted files are not accepted. ▪️ Data must be clean, readable, and directly usable. ▪️ Duplicate and near-duplicate content should be removed or clearly identified. ▪️ The provider must be able to explain the source and collection method. ▪️ Applicants must confirm that the proposed data can legally be collected and provided for the intended use. ▪️ Source URLs or provenance records should be available upon request. 📩 HOW TO APPLY Please provide the following information: 1️⃣ Content categories you can provide: literature, forums, blogs, or a combination. 2️⃣ Whether the data is ready-made or needs to be newly sourced. 3️⃣ Estimated raw data volume in GB/ Estimated usable token volume after cleaning and deduplication. 4️⃣ Proposed rate per GB or per 1 million usable tokens. Estimated lead time. 5️⃣ A representative plain-text sample of approximately 1–5MB, or a shareable Google Drive link. ⚠️ IMPORTANT Please do not submit machine-translated, AI-generated, OCR-extracted, duplicated, or unverifiable datasets. We are looking for vendors who can support large-volume delivery. Partial-volume proposals are also welcome.

  • More than 30 hrs/week
    Hourly
  • 3-6 months
    Duration
  • Entry level
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
Web Scraping
Data Entry
Activity on this job
  • Proposals:10 to 15
  • Last viewed by client:3 weeks ago
  • Interviewing:
    10
  • Invites sent:
    30
  • Unanswered invites:
    16
About the client
Member since Aug 28, 2024
  • South Korea
    Seoul8:37 PM
  • $120K total spent
    376 hires, 292 active
  • Tech & IT
    Mid-sized company (10-99 people)

Explore similar jobs on Upwork

Artificial Intelligence
AI Agent Development
Automation
AI Development
Python
AI App Development
Document AI
Generative AI
AI Bot
AI Chatbot
Data Extraction
LLM Prompt Engineering
LangChain
OCR Algorithm
Prompt Engineering
Claude
n8n
AI Implementation
Retrieval Augmented Generation
AI Model Integration
ETL
Airtable
Database Design
Automation
No-Code Development
Data Migration
Zapier
Customer Relationship Management

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo