Build a Human-Like Conversational AI That Learns From Real Conversations + Continuous Learning

Posted 3 days ago

Worldwide

Summary

*UPDATED* AI/LLM / Conversational AI Engineer — Real Conversation Training, Learning & Intelligence for Fanvue Creator Platform I’m looking for an experienced AI/LLM / Conversational AI engineer to improve the actual conversational intelligence behind my existing platform, Taca. Website: https://tacatech.net/ IMPORTANT: THIS IS NOT A PROMPT-ENGINEERING JOB. Taca is already a substantial working platform with Personas, Controls, conversation modes, memory, automation, Fanvue integration, Vault/content systems, subscriptions, commercial goals, screening, bot detection and other functionality. I can already write and improve prompts myself. What I am looking for is someone who can take the underlying conversational intelligence significantly further — particularly through real conversational data, model selection, training/adaptation where appropriate, evaluation, learning systems and production LLM architecture. The goal is to make Taca genuinely excellent at natural human conversation rather than simply making an LLM follow a larger system prompt. -------------------------------------------------- THE MOST IMPORTANT DISTINCTION -------------------------------------------------- There are TWO different types of learning/training that I want to keep completely separate. 1. PERSONA-SPECIFIC TRAINING Taca has functionality for teaching an individual Persona from that Persona’s own material/conversations. For example, a creator may provide: - screenshots of conversations - pasted messages - chat exports - other Persona-specific information. This is intended to help a particular Persona better reflect that creator/persona’s own voice, personality, communication style and characteristics. This Persona-specific information must remain isolated to the appropriate creator/Persona. 2. TACA-WIDE CONVERSATIONAL INTELLIGENCE This is a completely different problem. I want to investigate how the UNDERLYING CONVERSATIONAL INTELLIGENCE of Taca can become better over time. The objective here is not to make Persona A sound more like Persona A. The objective is to make the underlying AI better at things such as: - natural conversation - conversational flow - context - emotional understanding - conversational timing - variation - humour - slang - relationship-oriented dialogue - knowing when to ask questions - knowing when not to ask questions - understanding implied meaning - avoiding repetitive openings - avoiding repetitive phrases - avoiding generic AI language - maintaining long conversations - understanding changing conversational energy - responding naturally to different situations. The global conversational intelligence should improve from appropriate aggregate evidence, while individual creator Personas, memories and private information remain separate. I do NOT want these two learning systems confused with each other. -------------------------------------------------- REAL CONVERSATIONAL DATA -------------------------------------------------- A major part of this project is investigating whether real human conversation data can be used to improve the underlying conversational intelligence. I am NOT simply looking for someone to create synthetic examples or write better prompts. I want the engineer to investigate the technically appropriate ways of using substantial real conversational material, where the necessary rights, consent, privacy and commercial licensing exist. Potential uses could include: - evaluation datasets - conversational benchmarks - preference data - failure analysis - supervised fine-tuning - model adaptation - training data - comparison datasets - synthetic-data generation based on identified weaknesses - other approaches you recommend. I am NOT assuming that raw conversations should simply be "fed into the AI". I want someone who understands how real dialogue data should be: - selected - cleaned - structured - anonymised where appropriate - filtered - separated into training/evaluation sets - quality-controlled - legally/licensing checked - used safely in a commercial product. If fine-tuning is not the right answer, that is completely fine. If a better model is the answer, tell me. If preference learning is better, tell me. If evaluation data is more valuable than training data, tell me. If retrieval, memory or context construction is the actual bottleneck, tell me. I want the engineer to determine what actually improves conversational intelligence rather than assuming that "more training" automatically means a better model. -------------------------------------------------- THE LONG-TERM LEARNING / IMPROVEMENT GOAL -------------------------------------------------- I also want help designing and improving the way Taca can learn from increasing usage over time. I do NOT want to claim that this is already perfected — it isn't. This is specifically an area where I want an experienced engineer to investigate what we have, identify what is missing or weak, and design a much stronger long-term approach. The long-term vision is that as Taca grows, potentially to tens of thousands of creators/users and a very large volume of conversations, the increasing amount of appropriate data and feedback can provide increasingly strong evidence about what makes conversations better. For example, we may eventually be able to identify: - recurring conversational failures - repetitive behaviours - unnatural responses - poor conversational decisions - situations where the model misunderstands context - situations where it asks unnecessary questions - situations where it over-explains - situations where it fails to maintain continuity - differences between successful and unsuccessful conversations - responses creators frequently edit - responses creators approve - patterns associated with stronger conversation outcomes - model/provider differences - Persona-specific issues - genuinely global conversational weaknesses. I want to investigate whether this can eventually create a controlled improvement loop such as: CONVERSATIONS ↓ APPROPRIATE QUALITY / PERFORMANCE SIGNALS ↓ FAILURE & PATTERN DETECTION ↓ HIGH-QUALITY EVALUATION / TRAINING DATA ↓ MODEL / ARCHITECTURE / BEHAVIOUR IMPROVEMENT ↓ AUTOMATED REGRESSION TESTING ↓ CONTROLLED DEPLOYMENT ↓ MEASURE RESULTS ↓ FURTHER IMPROVEMENT But I want the engineer to determine what parts of this are actually technically valuable and feasible. I do NOT want a system that blindly retrains itself on every conversation. I do NOT want low-quality AI outputs recursively training future AI outputs. I do NOT want creator-specific information leaking into global training. I want a proper, controlled and evidence-based improvement pipeline. At significant scale — potentially 50,000+ creators/users — I want us to investigate whether Taca's growing volume can become a genuine advantage for improving the underlying conversational intelligence. This is a design goal, not a claim that the current system already achieves it. -------------------------------------------------- FANVUE / CREATOR CONVERSATIONS -------------------------------------------------- Taca is designed for Fanvue creators, so the conversational domain is very different from a normal customer-support chatbot. Depending on the creator's settings and applicable platform/model/provider policies, conversations can involve: ordinary everyday conversation → friendship → humour → teasing → flirting → relationship-style conversation → commercial/subscription conversations → adult/explicit conversation where permitted. The underlying conversational intelligence therefore needs to handle natural interpersonal conversation rather than simply answering questions. -------------------------------------------------- COMMERCIAL / VAULT CONTEXT -------------------------------------------------- Taca is also a commercial creator platform. The conversational AI may need to operate alongside systems involving: - followers - lapsed users - subscribers - subscription goals - retention - paid content - free content - Vault content - photos - videos - content sets - pricing - price bands - free gifts - content requests - commercial goals - automated messaging. For example, a conversation may involve a fan asking to see something. The correct behaviour may not simply be: "generate a text response." Depending on the current Taca configuration and available content, the system may need to understand the conversation, determine the appropriate commercial/content behaviour and potentially work with relevant Vault content. This can include situations where a relevant photo or video from the Vault needs to be offered or sent according to the creator's settings, pricing/content rules and conversation context. I want the engineer to investigate how well this currently works and where the conversational intelligence, Controls, commercial logic and Vault/content systems may be interacting incorrectly. I do NOT want applicants to assume that every part of this functionality is already perfect. Part of the job is to audit the existing behaviour and identify gaps, conflicts or limitations. -------------------------------------------------- TACA'S EXISTING ARCHITECTURE -------------------------------------------------- Taca already has substantial systems for: - Personas - Controls - conversation modes - memory - automation - Fan journeys - subscriptions - Vault/content - paid content - free gifts - price bands - commercial goals - screening - bot detection - audience stages - analytics. There are broadly three important areas: 1. CONVERSATIONAL INTELLIGENCE The underlying ability to have natural conversations. 2. GLOBAL CONTROLS Controls that influence how the AI behaves. 3. INDIVIDUAL PERSONAS Creator-specific identities, personalities and characteristics. However, I do NOT want candidates to assume that the existing architecture is already perfect. I have built a substantial amount of the current platform, but I want an experienced AI/LLM engineer to inspect, challenge and test it. There may be: - Controls that conflict - instructions that interact badly - memory/context problems - commercial logic interfering with conversation - Persona information being injected incorrectly - missing decision-making layers - architectural limitations - model/provider limitations - evaluation gaps - learning/data gaps - other problems we have not identified yet. I want the engineer to determine what is actually working correctly. The approach should be: AUDIT → TEST → IDENTIFY THE REAL BOTTLENECKS → IMPROVE WHAT EXISTS → CHANGE ARCHITECTURE WHERE JUSTIFIED. I do NOT want someone to rebuild the entire platform unnecessarily. But I also do NOT want someone to simply accept the existing architecture without questioning it. The primary objective is improving the actual conversational intelligence and its long-term learning/improvement capability. -------------------------------------------------- GLOBAL CONTROLS -------------------------------------------------- Taca has extensive Controls that influence how the AI behaves. Examples include: - Casual - Bestie - Tease - Partner - Explicit - Unrestricted - Short / Medium / Long / Smart replies - Emoji intensity - Followers - Lapsed - Subscribers - commercial goals - subscription behaviour - paid-content behaviour - Vault/content behaviour - automation - reply timing - screening - bot detection. For example: Partner + Short + Some Emoji + Subscription Goal + Specific Paid Content Behaviour should produce a response that naturally follows those requirements. It should not sound like the AI is reading a checklist. However, I want the engineer to audit these interactions rather than assume that they are already perfect. -------------------------------------------------- PERSONA SYSTEM -------------------------------------------------- Creators can have multiple Personas. For example: Creator → Persona A → Persona B → Persona C → Persona D Each Persona can have its own: - identity - personality - voice - character - preferences - relationship style - behavioural characteristics - memories - Persona-specific training material. The underlying conversational intelligence should be reusable across Personas. Changing the Persona should change WHO the AI feels like. Changing the Controls should change HOW that Persona behaves in a situation. But the Persona-specific training/data must remain separate from the global conversational intelligence learning system. -------------------------------------------------- CURRENT MODELS / PROVIDERS -------------------------------------------------- We have experimented with: - Hermes - Dolphin - Venice - OpenRouter - and other models/providers. I am not attached to any particular model. I want the engineer to investigate whether the best solution is: - improving the existing implementation - changing models - changing providers - model routing - specialised conversational models - fine-tuning - training/adaptation using real dialogue - better context construction - better memory - different inference configuration - or a combination. The architecture should ideally remain flexible enough to change models/providers later. -------------------------------------------------- EVALUATION -------------------------------------------------- I don't want someone to produce a few impressive demo conversations and call the project finished. I want proper evaluation. Testing should cover combinations of: - Personas - chat modes - reply lengths - emoji settings - audience stages - commercial goals - content behaviour - emotional situations - short messages - long messages - rapid conversations - long-running conversations - memory - context - tone changes - unusual messages - conflicting instructions. For example: Partner + Short + Some Emoji + Subscribe should behave differently from: Bestie + Smart + Minimal Emoji + Just Chat without either becoming robotic. I want to be able to test hundreds or thousands of combinations automatically. -------------------------------------------------- CONVERSATIONAL QUALITY -------------------------------------------------- I want to measure whether conversations are actually improving. Potential evaluation areas include: - naturalness - repetition - response diversity - conversational continuity - Persona consistency - Control compliance - memory accuracy - emotional appropriateness - conversational progression - unnecessary questions - generic/AI-like phrasing - long-conversation degradation - creator edits - creator approval - human evaluation - relevant performance signals. I am open to better evaluation methodologies if you have them. -------------------------------------------------- COMMERCIAL / LICENSING / PRIVACY -------------------------------------------------- This is a commercial product. Any proposed: - model - dataset - training data - conversation dataset - character/persona data - fine-tuning data - third-party API - open-source component - external service must have licensing/usage terms compatible with the intended commercial use. I also need proper consideration of: - consent - privacy - anonymisation - data separation - commercial rights - platform/provider terms. I am interested in using high-quality existing conversational resources where appropriate, but nothing should be incorporated into a commercial system without verifying that it can actually be used for that purpose. -------------------------------------------------- WHAT I WANT DELIVERED -------------------------------------------------- Ultimately I want a significantly stronger conversational intelligence system integrated with Taca. The result should aim to demonstrate that: 1. Conversations feel substantially more natural. 2. The underlying conversational intelligence is genuinely stronger, rather than simply having a larger prompt. 3. Personas remain recognisably distinct. 4. Global Controls influence behaviour correctly. 5. Multiple Controls can operate together without making responses robotic. 6. Conversations do not become repetitive or scripted. 7. Context and memory are maintained correctly. 8. Long conversations remain coherent. 9. Commercial/Vault-related situations are handled naturally alongside the conversation. 10. Model/provider failures are handled gracefully. 11. The system remains flexible enough to change models/providers. 12. Real conversation data can be used appropriately for evaluation and potentially training/adaptation where justified. 13. There is a strong automated evaluation/regression system. 14. There is a credible long-term approach for improving the global conversational intelligence as Taca grows. 15. Persona-specific learning and global conversational learning remain properly separated. -------------------------------------------------- PLEASE ANSWER THESE QUESTIONS IN YOUR PROPOSAL -------------------------------------------------- 1. REAL CONVERSATIONAL DATA Have you worked with real-world conversational/dialogue datasets? What kinds of conversations? Have you actually used real conversations to improve a conversational AI system? 2. TRAINING / MODEL ADAPTATION Have you worked with: - fine-tuning - supervised fine-tuning - preference learning - dialogue modelling - conversational model training - model adaptation - retrieval/context systems - memory systems? Please explain what you personally built. 3. REAL CHAT DATA If given a large amount of real conversation material, how would you determine: - what is useful - what should be discarded - what should become evaluation data - what could become training data - what should remain Persona-specific - what could potentially contribute to global conversational learning? 4. GLOBAL VS PERSONA LEARNING How would you technically separate: GLOBAL TACA CONVERSATIONAL LEARNING from: INDIVIDUAL CREATOR/PERSONA LEARNING? This is extremely important to us. 5. CONTINUOUS IMPROVEMENT How would you design a system where increasing usage can provide better evidence and improve the underlying conversational intelligence over time? How would you approach this at potentially 50,000+ creators/users? 6. EXISTING ARCHITECTURE How would you audit the current Persona, Controls, Memory, Vault, commercial and conversational systems? What would you look for first? 7. MODEL SELECTION What models would you consider? Would you keep Hermes/Dolphin/Venice/OpenRouter? Why or why not? 8. EVALUATION How would you objectively evaluate natural conversation? How would you test hundreds/thousands of Persona × Control combinations? 9. VAULT / COMMERCIAL CONVERSATION How would you approach the interaction between conversational intelligence, commercial goals and content/Vault systems? 10. PRODUCTION EXPERIENCE Please give specific examples of conversational AI/LLM systems you have personally built and operated in production. 11. COMMERCIAL DATA How would you handle privacy, consent, anonymisation, licensing and commercial use of real conversational data? -------------------------------------------------- ONE FINAL POINT -------------------------------------------------- Please do NOT send me a generic proposal saying: "I can build your AI chatbot." Taca already has the platform. And please don't make the proposal primarily about prompt engineering. I can already do that myself. I am specifically looking for someone who understands: REAL CONVERSATIONAL DATA + CONVERSATIONAL AI + MODEL / LLM ARCHITECTURE + TRAINING / MODEL ADAPTATION + EVALUATION + MEMORY / CONTEXT + PRODUCTION SYSTEMS + CONTINUOUS IMPROVEMENT. I want someone who can look at what already exists and tell me: "Here is what is working." "Here is what is not." "Here is what is causing the conversational problems." "Here is where real conversation data could help." "Here is what should potentially be trained, and what should not." "Here is how Persona-specific learning should remain separate." "Here is how the global conversational intelligence could improve as Taca grows." "Here is what needs changing in the existing architecture, and what should be left alone." If you have experience with real dialogue datasets, conversational model training, fine-tuning, preference learning, conversational agents, creator/fan messaging, relationship-oriented dialogue, roleplay, or large-scale production conversational systems, please make that the focus of your proposal. That is the expertise I am specifically looking to hire.

  • Less than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Intermediate
    Experience Level
  • Remote Job
  • Ongoing project
    Project Type
Skills and Expertise
Mandatory skills
AI Bot
Chatbot
AI Development
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:yesterday
  • Interviewing:
    0
  • Invites sent:
    0
  • Unanswered invites:
    0
About the client
Member since Aug 15, 2026
  • United Kingdom
    10:23 AM

Explore similar jobs on Upwork

Phd technical sideHourly‐ Renewed 2 months ago
Artificial Neural Network
Deep Neural Network
Convolutional Neural Network
Python
Data Science
Neural Network
Machine Learning
Deep Learning
Image Processing
Academic Writing
Technical Writing
Research Papers
Computer Science
C
AI Instructor for Student MasterclassesHourly‐ Posted 4 weeks ago
Adobe Illustrator
Graphic Design
Illustration
Adobe Photoshop

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo