Agentic AI Engineer – MCP, A2A, RAG & Production LLM Systems
Worldwide
We are looking for an Agentic AI Engineer to design, build, deploy and operate production-grade AI agent systems. Applicants with different levels of experience are welcome, provided they can demonstrate strong technical ability and practical experience building agentic AI or LLM-based systems. This is not a prompt-engineering-only position. You should be capable of developing agents that use tools, maintain state, retrieve knowledge, coordinate with other agents, execute multi-step workflows and safely escalate sensitive actions to humans when required. Key Responsibilities - Design agent workflows, decision loops, state management and escalation paths. - Build agentic AI systems using LangGraph, LangChain, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, OpenAI Agents SDK or custom orchestration. - Design and implement Agent-to-Agent communication and multi-agent coordination workflows. - Develop and integrate MCP servers that securely expose tools, APIs, databases, documents and internal services to AI agents. - Define structured tool schemas, message formats, capability descriptions and communication contracts. - Integrate models from Azure OpenAI, Anthropic, Mistral, Llama and other providers. - Select suitable models based on quality, latency, cost, context window, data residency and compliance requirements. - Develop secure tool integrations using function calling, structured outputs, APIs, MCP and A2A protocols. - Implement input validation, permissions, least-privilege access, retries, timeouts, rate limits, idempotency, sandboxing and audit logging. - Prevent infinite loops, duplicate actions, conflicting agent decisions and uncontrolled task delegation. - Implement RAG pipelines using embeddings, vector databases, hybrid search, metadata filtering and re-ranking. - Design short-term and long-term memory, conversation summarisation and context-budgeting strategies. - Implement human-in-the-loop approvals for sensitive or high-impact actions. - Build automated evaluation suites covering task completion, factuality, tool accuracy, safety, latency, cost and regression detection. - Create golden datasets, adversarial test cases, red-team tests and production feedback loops. - Implement protections against prompt injection, insecure retrieval, data leakage, PII exposure and unauthorised tool access. - Monitor agent performance, tool calls, latency, token usage, errors, success rates and cost per request. - Optimise costs through prompt caching, model routing, token budgeting, batching and the use of smaller models where appropriate. - Package and deploy agents using Docker, ---Kubernetes and CI/CD pipelines. - Implement secrets management, environment separation, feature flags, canary releases, rollback plans and provider failover. - Work with backend, QA and data engineers to define APIs, test scenarios, data requirements and deployment standards. - Document agent capabilities, limitations, tool permissions, escalation conditions and operational procedures. Required Skills and Experience Applicants should have: - Practical experience in software engineering, machine learning engineering or production AI engineering. - Experience building LLM applications or agentic - AI systems beyond basic prototypes. - Strong Python engineering skills, including asynchronous programming, type hints, API development, validation, testing and secure coding. - Experience with structured outputs, function calling, API integration and tool execution. - Experience with agent orchestration frameworks or custom agent loops. - Understanding of multi-agent patterns, task delegation, agent state, message passing and workflow coordination. - Practical knowledge of MCP concepts and secure agent-to-system connectivity. - Experience designing or integrating RAG systems. - Knowledge of embeddings, vector databases, semantic search, hybrid search and re-ranking. - Understanding of agent evaluation, observability, tracing and production monitoring. - Knowledge of prompt injection, access control, PII handling, sandboxing and audit trails. - Experience deploying applications using Docker, with Kubernetes experience preferred. - Strong communication and technical documentation skills. Relevant Technologies Experience with several of the following is expected: - LangGraph, LangChain, LlamaIndex or Semantic Kernel - AutoGen, CrewAI or OpenAI Agents SDK - MCP servers and MCP clients - Agent-to-Agent communication protocols - Azure OpenAI and Anthropic Claude - Mistral, Llama or other open models - Azure AI Search, Pinecone, Weaviate or FAISS - LangSmith, Braintrust or comparable evaluation and tracing platforms - Azure Application Insights or ELK - FastAPI, Pydantic and asynchronous Python - Docker and Kubernetes - GitHub Actions or Azure DevOps - Hugging Face - vLLM or TensorRT-LLM Additional Relevant Experience The following would be advantageous: - MCP server authoring - Production A2A or multi-agent system development - Agent capability discovery and task delegation - Multimodal agent development - Arabic-language localisation - AI governance within regulated industries - On-premises or local-model deployment - Model routing and provider failover - LLM cost and performance optimisation Expected Deliverables Deliverables may include: - Production-ready agent architecture - A2A or multi-agent communication workflows - MCP servers and MCP client integrations - Secure tool wrappers and API integrations - RAG, memory and context-management components - Human-approval and escalation workflows - Automated evaluation and regression-testing suites - Monitoring, tracing and cost dashboards - Dockerised deployment and CI/CD configuration - Technical and operational documentation Application Questions Please answer the following five questions in your proposal: 1. Describe your practical experience with MCP. Have you built an MCP server or integrated an agent with one? Explain what tools, resources or systems it exposed and your specific contribution. 2. Describe an A2A or multi-agent system you have developed. How did the agents exchange messages, discover capabilities, delegate tasks and coordinate the workflow? 3. How would you secure an MCP and A2A architecture? Cover authentication, authorisation, least-privilege access, input validation, data protection and audit logging. 4. Describe a production LLM or agentic AI system you have built. Which orchestration framework, models, tools, RAG components and deployment technologies did you use? 5. How do you evaluate and monitor agent reliability? Explain how you test task completion, tool-call accuracy, safety, latency, cost and regressions before and after deployment. Please include links to relevant GitHub repositories, architecture diagrams, demonstrations or other technical work samples.
- More than 30 hrs/weekHourly
- 6+ monthsDuration
- IntermediateExperience Level
$15.00
-
$40.00
Hourly- Remote Job
- Ongoing projectProject Type
Skills and Expertise
Activity on this job
- Proposals:20 to 50
- Interviewing:0
- Invites sent:0
- Unanswered invites:0
About the client
- United KingdomLondon8:07 PM
- Tech & ITSmall company (2-9 people)
Explore similar jobs on Upwork
How it works
Create your free profileHighlight your skills and experience, show your portfolio, and set your ideal pay rate.
Work the way you wantApply for jobs, create easy-to-by projects, or access exclusive opportunities that come to you.
Get paid securelyFrom contract to payment, we help you work safely and get paid securely.
About Upwork
- 4.9/5(Average rating of clients by professionals)
- G2 2021#1 freelance platform
- 49,000+Signed contract every week
- $2.3BFreelancers earned on Upwork in 2020
Find the best freelance jobs
Growing your career is as easy as creating a free profile and finding work like this that fits your skills.
Trusted by