Manual QA Tester for AI/LLM Web Apps — Frontend & Backend

Posted yesterday

Worldwide

Summary

We’re an AI-native venture studio building and launching web-based AI products. We need a Manual QA Tester who can test the complete product—not only the interface, but also the APIs, backend behavior, data flow, and AI-generated results. This is a QA-first role. Full-stack engineering experience is strongly preferred because we want someone who can inspect code, diagnose likely root causes, and potentially fix smaller issues. What you’ll do: - Manually test AI-powered web applications across desktop and mobile browsers. - Test complete user journeys, including registration, authentication, forms, reports, payments, permissions, and account states. - Test frontend behavior, responsive layouts, accessibility, loading states, error messages, and overall usability. - Test backend APIs, request and response data, authentication, validation, timeouts, rate limits, and failure handling. - Validate data across the browser, API, and database when access is provided. - Red-team AI and LLM functionality for hallucinations, fabricated facts, unsupported claims, prompt injection, inconsistent results, bias, and unsafe output. - Verify that AI-generated claims, scores, sources, and recommendations are supported by evidence. - Test low-data and ambiguous cases where the AI should admit uncertainty instead of inventing an answer. - Create and maintain test cases, regression checklists, and launch-readiness reports. - Report bugs with severity, reproducible steps, expected versus actual behavior, screenshots or recordings, logs, and relevant network requests. - Retest fixes and provide a clear launch recommendation: ready, conditionally ready, or not ready. Required qualifications - At least 3 years of hands-on manual QA or software testing experience. - Experience testing modern web applications from end to end. Strong understanding of functional, exploratory, regression, integration, usability, and negative testing. - Ability to test REST APIs using Postman, Insomnia, curl, or similar tools. - Confidence using browser developer tools, including the console, network panel, storage, cookies, and request inspection. - Ability to write concise bug reports that developers can reproduce without additional explanation. - Experience testing responsive behavior across multiple browsers and screen sizes. - Strong written English and reliable communication. - Comfortable working independently in fast-moving, pre-launch environments. Strongly preferred - Direct experience testing AI, generative AI, agents, RAG systems, or LLM-powered products. - Understanding of hallucinations, non-deterministic outputs, context limits, retrieval failures, prompt injection, and model fallbacks. - Full-stack engineering experience with frontend and backend systems. - Ability to read or troubleshoot HTML, CSS, JavaScript, TypeScript, React, Node.js, SQL, and serverless applications. - Experience tracing a frontend defect through an API request to the likely backend cause. - Ability to implement small, well-scoped fixes and submit them through Git and pull requests. - Experience with Cloudflare, serverless functions, databases, authentication, payments, or third-party AI APIs. - Familiarity with Playwright, Cypress, Selenium, or another automation framework. - Accessibility testing knowledge, including WCAG. What success looks like - You will help us catch failures that ordinary UI testing misses: plausible-but-wrong AI answers, unsupported claims, silent backend failures, inconsistent scoring, entitlement problems, exposed credentials, broken mobile flows, and errors that appear to users as successful results. - We care more about judgment and prioritization than the total number of bugs reported. How to apply - Start your proposal with the word LAUNCH to show you've read through the requirements. Submit a response to these questions, we are limited on time and need to do a deep pre-screen, followed by a short introductory video call to the candidate we select. 1. What have you personally delivered? Provide two relevant AI or LLM products you personally tested. For each, include: - Product or redacted project description - Models or AI providers involved - Frontend, backend, API, database, and mobile responsibilities - What you personally discovered or improved - A URL, test plan, bug report, Loom, PR, or redacted work sample (Do not send only a company name or list of responsibilities) 2. Walk us through your end-to-end process from receiving an unfamiliar pre-launch product to issuing a final go/no-go recommendation. Include a sample of the deliverables we would receive. 3. What AI failure have you personally found that normal QA would miss and how did you fix it? 4. How would you test an AI feature that produces different answers for the same input? Explain your exact methodology for determining whether it passes or fails. 5. Describe a defect you traced from the frontend through an API or backend service. 6. How would you test our complete launch flow? 7. Which tools and technologies can you test or troubleshoot? 8. . Can you make small code fixes? If yes, share your GitHub, portfolio, or a relevant example. **We will not select a candidate who can't show previous work, projects, client recommendations or references** 9. What are your hourly rate, timezone, weekly availability, and earliest start date? If NDA prevents naming a client, anonymize the client but explain your personal contribution precisely. **Generic proposals that do not answer these questions will not be considered.** Required: 1. Past work/projects, Linkedin Profile, any customer references/recommendations. 2. NDA will be required before accessing confidential products. We have multiple products in development and many more ideas in the pipeline. This will begin as a paid trial so we can evaluate your work and how well we collaborate. If successful, we plan to hire you on an ongoing, project-by-project basis.

  • More than 30 hrs/week
    Hourly
  • 1-3 months
    Duration
  • Expert
    Experience Level
  • $25.00

    -

    $50.00

    Hourly
  • Remote Job
  • Ongoing project
    Project Type

Contract-to-hire opportunity

This lets talent know that this job could become full time.
Learn more
Skills and Expertise
Mandatory skills
Full-Stack Development
Activity on this job
  • Proposals:20 to 50
  • Last viewed by client:yesterday
  • Interviewing:
    4
  • Invites sent:
    14
  • Unanswered invites:
    9
About the client
Member since Jul 28, 2026
  • USA
    Mckinney8:07 PM
  • Tech & IT
    Small company (2-9 people)

Explore similar jobs on Upwork

Shopify Website Development for Arc GISHourly‐ Posted 1 month ago
HTML
Web Development
Shopify
Shopify Templates
SQL
Microsoft SQL Server
Oracle Database
POS Terminal
POS Terminal Development

How it works

  • Post a job icon
    Create your free profile
    Highlight your skills and experience, show your portfolio, and set your ideal pay rate.
  • Talent comes to you icon
    Work the way you want
    Apply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
  • Payment simplified icon
    Get paid securely
    From contract to payment, we help you work safely and get paid securely.
Want to get started? Create a profile

About Upwork

  • Rating is 4.9 out of 5.
    4.9/5
    (Average rating of clients by professionals)
  • G2 2021
    #1 freelance platform
  • 49,000+
    Signed contract every week
  • $2.3B
    Freelancers earned on Upwork in 2020

Find the best freelance jobs

Growing your career is as easy as creating a free profile and finding work like this that fits your skills.

Trusted by

  • Microsoft Logo
  • Airbnb Logo
  • Bissell Logo
  • GoDaddy Logo