Senior AI/LLM Infrastructure Architect Self-Hosted open source model kimi k3 and qwen 3.8
Worldwide
We run a 100-person agency currently on Claude (Opus 4.x) and Fable for most of our production workflows. We want to run a controlled experiment self-hosting Kimi K3 (2.8T total) and Qwen3's largest open-weight release once available, to see if a self-hosted setup can match our current Claude experience across the whole team during office hours. We need an architect who has actually deployed open-weight LLMs in production for 100+ concurrent users not a prototype, not a demo, a real multi-user serving environment. WHAT WE NEED - Full deployment plan: model serving engine, quantization strategy, load balancing, and how you'd validate output quality against Opus before rollout - GPU/server sizing for ~100 concurrent office-hours users - A plan for matching Claude's response latency under real team load - Ongoing operations plan (monitoring, model updates, failover) Project starts when open weights for the target models are released. We have time we are not in a rush and would rather wait for the right architect than rush a bad hire. TO APPLY, ANSWER ALL FOUR QUESTIONS BELOW IN YOUR OWN WORDS: Answer the question 1. This will be the deciding factor. Create as detailed a plan as you can on how you will achieve this task. Don't use AI; write in your own words. Explain the reasoning for each part and what open-source tools you will use to deploy this. 2. Share your portfolio of AI projects you have done for enterprise clients, your role, the biggest challenge you faced in 1-2 projects, and how you resolved them. If you have a portfolio of fewer than 10 projects, please save your connects and don't apply. 3. How will you achieve the same speed we get with Claude when all 100 of our employees are using the self-hosted LLM during office hours? 4. What resources do we need in terms of servers, GPUs, and memory to run this for all 100 employees? Please don't submit an AI-generated proposal. Take your time answering the questions. If your proposal has everything I ask for, I will definitely interview you. The project starts when the open weights get released, so I have plenty of time to interview. Don't rush by submitting a generic AI proposal and wasting your connects.
- More than 30 hrs/weekHourly
- 6+ monthsDuration
- ExpertExperience Level
$35.00
-
$80.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:5 to 10
- Last viewed by client:3 days ago
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- IndiaGuwahati6:51 AM
- $2.9K total spent17 hires, 1 active
- 16 hours
- Tech & ITSmall company (2-9 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by