- Hourly
- Expert
- Est. time: 1 to 3 months, Not sure
We are looking for an experienced AI Architect / Senior LLM Engineer to design and build an enterprise-grade AI platform for the healthcare industry. You will lead the architecture and implementation of intelligent AI solutions that improve clinical operations, automate administrative workflows, and enable healthcare professionals to access trusted medical knowledge through advanced AI technologies. The ideal candidate has hands-on experience building production-ready Agentic AI systems, Multi-Agent architectures, RAG pipelines, and LLMOps using modern AI frameworks and cloud platforms. Responsibilities Design and develop scalable Agentic AI solutions for healthcare applications. Build Multi-Agent Systems using LangGraph, CrewAI, or AutoGen. Develop enterprise Retrieval-Augmented Generation (RAG) pipelines for medical knowledge retrieval. Create AI agents for clinical knowledge assistance, document intelligence, workflow automation, and care coordination. Build and integrate MCP servers and custom AI tools with internal healthcare systems. Optimize prompt engineering, retrieval strategies, and response quality for high accuracy. Implement AI guardrails, evaluation pipelines, monitoring, and observability for production deployments. Deploy secure, scalable AI infrastructure on AWS using Infrastructure as Code and CI/CD best practices. Collaborate with engineering, product, and healthcare stakeholders to deliver reliable AI solutions. Required Skills 5+ years of experience in AI/ML or Generative AI development. Strong expertise in Python and backend API development. Experience with LangGraph, CrewAI, AutoGen, or similar multi-agent frameworks. Hands-on experience with AWS Bedrock, Azure OpenAI, or Vertex AI. Strong understanding of RAG architectures, vector databases, embeddings, and semantic search. Experience with Pinecone, Weaviate, pgvector, or similar vector databases. Knowledge of LLMOps, evaluation frameworks, prompt engineering, and AI observability tools. Experience with Docker, Terraform, CI/CD, and cloud-native deployments. Familiarity with healthcare compliance, security, and responsible AI practices is highly preferred. Preferred Technologies LangGraph CrewAI AutoGen AWS Bedrock Claude GPT-4o Gemini Pinecone pgvector LangSmith Arize Phoenix FastAPI Docker Terraform GitHub Actions MLflow Nice to Have Experience developing AI-powered healthcare platforms. Knowledge of healthcare workflows, clinical documentation, or medical knowledge systems. Experience integrating AI solutions with enterprise applications through APIs and MCP. Familiarity with AI governance, model evaluation, and production monitoring. If you are passionate about building enterprise-scale AI solutions that transform healthcare through Agentic AI and Generative AI, we'd love to hear from you.
- Hourly: $65.00 - $128.00
- Expert
- Est. time: 1 to 3 months, Less than 30 hrs/week
Lead the development of an AI-agent platform that autonomously analyzes financial transactions, customer activity, cash flow, and risk signals to support FinTech operations. The system will use LLM-based agents, ML models, RAG, and real-time financial data to investigate anomalies, assess risk, generate financial insights, and recommend actions. Key Responsibilities: Define the AI-agent architecture, product roadmap, agent workflows, and evaluation strategy. Design specialized agents for fraud investigation, transaction analysis, risk assessment, cash-flow analysis, and financial reporting. Combine deterministic financial rules with ML predictions and LLM reasoning rather than relying solely on LLM outputs. Build agent orchestration using LangGraph/LangChain, tool calling, structured outputs, memory, and RAG. Develop ML pipelines for anomaly detection, behavioral scoring, transaction classification, and risk prediction. Implement human-in-the-loop approvals, confidence scoring, audit trails, and agent observability. Establish evaluation frameworks for agent accuracy, hallucination detection, tool-use reliability, and financial decision quality. Work with engineering teams to productionize agents using Python, FastAPI, PostgreSQL, AWS, Docker, and MLflow. Core Technologies: Python, PyTorch, Scikit-learn, XGBoost/LightGBM, LangGraph, LangGraph, LLMs, RAG, vector databases, PostgreSQL, FastAPI, AWS, Docker, MLflow, REST APIs, Databricks, and event-driven architectures. Core ML Libraries: - Deep Learning: TensorFlow, PyTorch, Lightning - Classical ML: Scikit-learn, XGBoost, LightGBM - NLP/LLMs: Hugging Face Transformers, spaCy - Hyperparameter Tuning: Optuna, Ray Tune Infrastructure & Tools: - Cloud: AWS/GCP/Azure (S3, BigQuery, Sagemaker) - MLOps: MLflow, Kubeflow, Prefect - Data: SQL, Pandas, PySpark, Dask - Deployment: Docker, Kubernetes, FastAPI Expected Outcome: A production-ready multi-agent FinTech intelligence system where AI agents investigate financial events, combine ML predictions with financial rules, retrieve supporting data, explain their reasoning, and route high-risk decisions to human reviewers. Skills Artificial Intelligence (AI) Machine Learning AI Agent Development RAG PyTorch Scikit-learn Databricks MLOps
- Hourly: $60.00 - $120.00
- Expert
- Est. time: 3 to 6 months, 30+ hrs/week
Stevens is a strategic advisory and technology firm that partners with companies on high-impact transformation and technology initiatives. We work directly with executive and technology leaders to solve complex problems, identify opportunities, and turn strategy into working solutions. We are looking for a highly experienced Enterprise AI Solutions Architect to support enterprise customer engagements on an hourly contract basis. Please read before applying This is a senior enterprise architecture role with a high bar for relevant experience. We are specifically looking for candidates who have worked within major U.S. enterprises, leading technology companies, established consulting firms, systems integrators, or comparable large-scale organizations. Experience consisting primarily of freelance projects, small startups, personal projects, or AI prototypes will not meet the requirements for this role. Candidates must: * Be a U.S. citizen and current U.S. resident. * Have substantial professional experience designing and delivering technology within large enterprise environments. * Have experience working directly with enterprise customers, stakeholders, or technology organizations. * Be able to participate in customer meetings during standard U.S. business hours. * Complete a technical architecture interview as part of our screening process. * Be prepared for Stevens to verify relevant employment history and professional experience before engagement. Please only apply if you meet these requirements. The Role Our engagements often sit at the intersection of business strategy, enterprise architecture, and emerging technology. We help customers identify where AI can create meaningful business value, determine what should actually be built, and provide the technical leadership needed to move from concept to production. This is not a salaried or full-time employment position. Architects work with Stevens on an hourly contract basis as project needs arise. Engagement scope and weekly hours may vary based on active customer projects. This is a senior, highly customer-facing role. You will work directly with customer technology leaders, engineering teams, and Stevens delivery teams to understand complex business problems, shape AI use cases, design solutions, guide proofs of concept, and establish the architecture required to take successful solutions into production. We are specifically looking for someone who understands enterprise application development and architecture, not simply AI models or cloud infrastructure. The model is one component of the system. You should be comfortable reasoning across applications, APIs, data, integrations, identity, security, cloud infrastructure, AI services, and production operations to design complete enterprise solutions. Your work may include * Leading technical discovery and architecture sessions with enterprise customers. * Translating business problems and AI opportunities into practical solution architectures and implementation plans. * Designing enterprise AI applications using LLMs, RAG, agents, tool use, structured and unstructured data, APIs, and modern cloud services. * Integrating AI capabilities into existing enterprise applications, data platforms, SaaS products, and business workflows. * Architecting solutions primarily within AWS, including Amazon Bedrock and SageMaker. * Evaluating models, platforms, frameworks, and architectural approaches against specific business and technical requirements. * Helping customers make build-vs-buy decisions across a rapidly evolving AI technology landscape. * Designing rapid proofs of concept while ensuring successful POCs have a realistic path to production. * Defining architecture across application services, APIs, data, model infrastructure, integrations, identity, observability, and cloud infrastructure. * Accounting for enterprise requirements including security, privacy, governance, scalability, reliability, performance, and cost. * Establishing appropriate approaches to AI evaluation, guardrails, monitoring, and production operations. * Working closely with engineers throughout implementation to ensure architectural decisions translate into high-quality working software. * Reviewing technical approaches, identifying risks, and helping engineering teams navigate architectural tradeoffs. * Presenting architecture recommendations and technical decisions to engineering teams, technology leaders, and senior customer stakeholders. This is not a role where you create an architecture diagram and hand it to someone else. We want architects who stay close to implementation, understand the consequences of their decisions, and can help a team navigate the difficult decisions that emerge as a solution gets built. You will have significant autonomy and ownership. A successful architect at Stevens can walk into an ambiguous problem, quickly understand the customer’s business and technology environment, identify the decisions that actually matter, develop a point of view, and confidently guide both the customer and delivery team toward a solution. Requirements * U.S. citizenship and current U.S. residency are required. * 7+ years of relevant professional experience across software architecture, solutions architecture, application development, engineering, or technical consulting. * Meaningful experience working within or delivering solutions for major U.S. enterprises, leading technology companies, established consulting firms, systems integrators, or comparable large-scale organizations. * Significant experience designing and delivering enterprise applications, platforms, or distributed systems. * Hands-on experience architecting and implementing generative AI applications beyond basic prototypes or isolated model integrations. * Experience taking applications or major technical initiatives from discovery and proof of concept through production. * Strong understanding of modern AI application patterns including RAG, agents, tool use, embeddings, vector search, model orchestration, evaluation, and guardrails. * Strong AWS architecture experience, including practical experience with Amazon Bedrock and SageMaker. * Experience working with commercial AI platforms and models including OpenAI and Anthropic Claude, along with familiarity with open-source models. * Strong understanding of modern application architecture, backend services, APIs, databases, data architecture, distributed systems, and enterprise integration patterns. * Experience integrating new applications and capabilities with existing enterprise systems and data sources. * Familiarity with cloud engineering, DevOps, MLOps, CI/CD, infrastructure as code, and production observability. * Strong understanding of security, identity, privacy, governance, reliability, scalability, and cost considerations in production enterprise systems. * Ability to make pragmatic architectural decisions within complex existing technology environments where replacing everything and starting over is not an option. * Ability to move comfortably between conversations about business requirements and detailed technical architecture. * Experience leading technical discovery, architecture workshops, and customer-facing solution design. * Strong written and verbal communication skills. * Confidence presenting recommendations, defending architectural decisions, and challenging assumptions with senior technology and business stakeholders. * Availability during standard U.S. business hours for customer workshops, architecture sessions, and meetings as needed. Experience at organizations such as major technology companies, cloud providers, Fortune 500 enterprises, nationally recognized consulting firms, or large systems integrators is strongly valued. The company name itself is not the qualification; we care about whether you have actually operated in sophisticated enterprise technology environments and solved problems at meaningful scale. An active U.S. government security clearance is not required but is strongly valued. Additional compensation may be available for individuals who hold and maintain an active U.S. government security clearance. Our Screening Process We maintain a high technical and professional bar because architects in this role represent Stevens directly with enterprise customers. Qualified applicants should expect: 1. Resume and experience review, with particular attention to enterprise architecture and delivery experience. 2. Initial screening conversation focused on your background, customer-facing experience, and recent enterprise work. 3. Architecture interview in which you will be asked to reason through an enterprise technology scenario, make architectural decisions, discuss tradeoffs, and explain how you would take a solution from discovery toward production. 4. Employment and professional experience verification before engagement. Please make sure the employment history and experience described in your Upwork profile, resume, and proposal are accurate and verifiable. Working With Stevens At Stevens, you will have the opportunity to work on sophisticated enterprise AI initiatives across multiple industries, work directly with customer technology leaders and engineering organizations, own meaningful architectural decisions, and stay close enough to implementation to see the solutions you design actually get built. We are building Stevens around the idea that a small, highly capable team can do exceptional work for sophisticated customers without introducing the bureaucracy that often comes with traditional consulting. We value good judgment, technical excellence, clear communication, intellectual curiosity, and pragmatic decision-making. We want people willing to challenge unnecessary complexity, recognize when a simpler solution is better, and say when AI is not the right answer. If your background is in enterprise-scale technology architecture and delivery, and you enjoy solving difficult problems directly with customers, we would like to hear from you.
- Hourly: $60.00 - $128.00
- Expert
- Est. time: More than 6 months, 30+ hrs/week
Principal AI Infrastructure & HPC Engineer We are a PE-backed, AI-native technology services company building a new AI Infrastructure & HPC practice across AWS and Microsoft Azure. We're looking for a deeply technical engineer to help build the practice from the ground up. This is not a traditional cloud or DevOps role. The focus is large-scale GPU infrastructure, distributed AI workloads, high-performance networking, and getting expensive compute environments to perform at their potential. What You'll Work On GPU cluster benchmarking, performance tuning and optimization NCCL benchmarking and tuning AWS EFA and Azure GPU/HPC infrastructure InfiniBand, RDMA, RoCE and GPUDirect RDMA CUDA, NVLink/NVSwitch and NVIDIA GPU environments Distributed AI training and inference optimization Linux, kernel, driver and systems-level performance Kubernetes/EKS/AKS and Slurm-based GPU environments GPU cloud / neocloud infrastructure What We're Looking For We want someone with deep hands-on expertise in GPU/HPC systems who can benchmark an environment, identify where performance is being lost, and fix it. Experience with NCCL, CUDA, InfiniBand/RDMA, distributed training, Linux performance engineering, and large multi-node GPU clusters is particularly relevant. Deep AWS and/or Microsoft Azure experience is a major plus, particularly experience designing or optimizing GPU/HPC workloads using EFA, EC2 accelerated computing, Azure GPU infrastructure, EKS/AKS, and high-performance networking. You don't need to check every box. Depth in this domain matters more than breadth. More Than a Project We're building a practice around this capability. The right person can play an important role in defining our technical offerings, developing repeatable optimization methodologies, working directly with AWS and Microsoft, and helping us build the engineering team as the practice grows. If you've worked deep in GPU infrastructure, HPC, distributed systems, or high-performance networking, we'd like to talk.
- Hourly
- Expert
- Est. time: 3 to 6 months, Less than 30 hrs/week
*** DO NOT APPLY IF YOU CAN NOT DO AT LEAST 95% OF THE JOB DESCRIPTION*** We're looking for a battle-tested engineer with deep, self-earned coding fundamentals who also knows how to leverage AI as a force multiplier. This is not a role for developers who have grown dependent on AI to write code they couldn't write themselves. You should be able to navigate a large, complex brownfield codebase on your own — and when you do bring AI into the work, your engineering foundation is what makes the difference between AI generating noise and AI generating production-ready solutions. This role is deliberately vertical. You'll write product code, and you'll own the AWS environment it runs on. Those aren't two jobs handed to one person to save a headcount — they're one job, because the interesting failures happen at the seam between them. ## About the Role You'll work across a TypeScript codebase with a Next.js frontend and AWS-backed services supporting consumer mobile applications at meaningful scale. AI tools are a deliberate part of the workflow: prompting for implementation plans, critically evaluating those plans against our architecture and business requirements, reviewing generated code for correctness and quality, and shipping with confidence. When AI hits the limits of a complex legacy system — and it will — you'll be the one who knows how to guide it through. On the platform side, you'll own infrastructure defined in code, the network boundaries around our data services, our cloud security posture, and the vulnerability backlog. The systems you'll inherit include an event-driven ingestion pipeline (managed queues and a key-value store behind an API gateway), a columnar analytics warehouse feeding BI dashboards, object storage with a query layer over it, a document database, and webhook integrations with third-party attribution and app-store billing systems. That's real scope, and we're stating it plainly so you can decide whether you want it. If you'd rather not touch infrastructure, this isn't the role. If you've been looking for a job where you own the whole vertical instead of filing tickets across a boundary, it is. ## What You'll Do ### Product engineering - Own the full lifecycle of AI-assisted development: generating plans, stress-testing them against real architectural constraints, and validating that generated code is production-worthy - Write correct concurrent code: reason clearly about async/await, the event loop, promise scheduling and cancellation, and the difference between concurrency and parallelism — and keep blocking work off the request path so a slow upstream API never stalls the thread serving users - Make and defend architectural decisions: know where layered, clean, and hexagonal (ports-and-adapters) designs each earn their complexity, enforce separation of concerns, and keep AI-generated solutions inside the boundaries the codebase already established rather than letting them drift toward whatever pattern the model saw most often in training - Practice test-driven development in earnest — write the failing test first, make it pass, then refactor — and use TDD as the mechanism that keeps AI-generated code honest rather than a box to check afterward - Drive tests past the happy path: use AI to enumerate boundary values, error branches, race conditions, malformed input, and failure modes of dependencies, then verify the generated tests actually assert behavior instead of restating the implementation back at itself - Contain the blast radius of AI-assisted work: small reviewable diffs, plans before code, incremental commits, contract and regression coverage on anything touching shared surfaces, and a bias toward changes you can reason about end to end - Build and maintain the backend services behind our mobile applications, including event ingestion, third-party webhook consumers, and the integrations that feed reporting ### Platform and infrastructure - Define and change infrastructure as code: extend and review Terraform modules, manage state, read a plan critically before applying it, and recover when state and reality disagree - Own cloud networking: private and public subnets, security group and NACL design, and the access patterns for data services that sit inside the network — with a clear view of when putting compute in a VPC is the right call and when it just buys cold starts and NAT charges - Design read and write paths for scale, cost, and exposure: when something is hammering an object store or an API, diagnose whether it's legitimate traffic or an unsecured origin being scanned, and reach for the fix that matches — edge caching, cache-control and conditional requests, presigned URLs with sane TTLs, request collapsing, batching and backpressure on one side; blocking public access, origin access control, bucket policy and IAM scoping, WAF rate limiting, and keeping repository metadata and build artifacts out of served paths on the other - Treat a surprising cloud bill as a security signal, not just a cost problem — know which request outcomes you're billed for, what shows up in access and audit logs, and when the right response is rotating credentials rather than adding a cache - Own and tune cloud security posture management (AWS Security Hub, GuardDuty, Config, or equivalents): configure the standards, suppress the noise so real signal survives, and drive findings to closed - Apply and enforce security best practices aligned with NIST controls, including access control, audit logging, system integrity, and secure configuration management — using automated config checks to continuously verify those controls rather than attesting to them in a document nobody re-reads - Manage IAM as a design problem: least-privilege roles, scoped policies, credential rotation, and no long-lived keys where a role will do - Own CI/CD: pipelines that gate on tests, coverage, and security scans, with deploys that are reproducible and reversible ### Vulnerability management - Remediate CVEs, don't just report them. Take findings from discovery through to a shipped fix, including the unglamorous part where the patched version is a major bump and you absorb the breaking changes across the application - Triage with judgment: knowing whether a finding is reachable in code paths we actually execute or buried in a transitive dependency that never runs is what keeps you from breaking production over a theoretical risk. But the deliverable is a closed finding, not an assessment - Maintain dependency hygiene across a large Node/TypeScript tree, where the vulnerability surface is mostly transitive and the fixes are mostly version bumps with consequences - Remediate infrastructure and configuration findings, not just application dependencies — misconfigured storage, over-permissive policies, unencrypted resources, missing logging ### Operations - Serve as a rapid-response resource for user-facing issues — diagnosing, prototyping, and deploying fixes fast when production is on the line - Leave the environment legible to someone else: documented infrastructure, runbooks for the things that page you, and reproducible deploys. Sole ownership only works if it isn't sole knowledge ## What We're Looking For - 5+ years of proven TypeScript development experience, with work you can walk us through line by line and explain the reasoning behind — you know the language, not just the prompts - Deep proficiency in React and Next.js - Fluency in the JavaScript concurrency model, and enough exposure to how other ecosystems solve the same problem (C# tasks, Python asyncio, Kotlin coroutines, Rust futures, Go goroutines) to explain what async/await actually buys you and where it doesn't help - Demonstrated experience with TDD and a clear point of view on what it's good for and where it isn't worth it - The ability to reason about architectural tradeoffs out loud — not just name patterns, but say what each one costs and when you'd skip it - Strong working knowledge of AWS across compute, managed queues, key-value and relational stores, object storage, CDN, WAF, and IAM — including request-pattern and caching design under load and hardening of publicly reachable origins - Hands-on Terraform experience: writing and reviewing modules, managing state, and recovering from drift. Other IaC backgrounds (CDK, CloudFormation, Pulumi) transfer if the depth is there - Practical cloud networking: VPCs, subnet architecture, security groups, NACLs, and private connectivity to managed data services - Experience running a cloud security posture tool in anger — configuring standards, tuning findings, mapping automated checks to a control framework, and closing items rather than accumulating them - A demonstrated CVE remediation history: specific vulnerabilities you personally fixed, in both application dependencies and infrastructure configuration, including at least one where the fix required meaningful refactoring - Working fluency with SCA tooling (AWS Security Hub, Prowler, Dependabot, Snyk, npm audit, or similar) and a defensible process for prioritizing what gets fixed - The ability to critically read and evaluate AI-generated code — catching architectural drift, security gaps, and subtle logic errors that AI won't flag itself - A specific, experience-backed account of where AI coding tools fail: missing system-wide context, no real model of your codebase's complexity or layering, confidently wrong abstractions, tests that validate the bug, and volume that outpaces review capacity — plus the practices you use to keep that in check - Familiarity with AWS Security Hub, CIS Benchmarks, NIST or comparable security frameworks and the ability to translate controls into practical engineering decisions - Comfort using AI tools (Claude, Copilot, etc.) as a development partner, with the technical depth to steer them effectively in unfamiliar or complex codebases - Strong debugging instincts and the ability to move fast under pressure without cutting corners on security or quality - CI/CD pipeline ownership experience, including gating deploys on tests, coverage, and security scans ## Nice to Have Experience with mobile platforms (iOS/Android). Familiarity with FedRAMP, SOC 2, or other compliance frameworks that map to NIST. Experience with data warehousing or BI tooling. Container or serverless packaging experience. A CS degree or equivalent depth in fundamentals — data structures, concurrency, systems — however you came by it.
- Hourly: $20.00 - $45.00
- Expert
- Est. time: 1 to 3 months, 30+ hrs/week
We're a US-based software consulting firm looking for a skilled full-stack engineer to join us long-term. This starts as an hourly engagement with a paid trial project, and converts to full-time for the right person. We have consistent client work and need someone reliable who ships fast and communicates well. This is not a vibe-coding role. We want a real engineer — strong fundamentals, solid architecture instincts, clean production code — who uses AI tools (Claude Code, Codex, Cursor) to move 5–10x faster. AI is your force multiplier, not your crutch. If you lean on AI to paper over gaps in your actual engineering ability, this isn't the fit. If you're a genuinely strong developer who has mastered AI-augmented workflows to ship more and better, keep reading. What you'll do: Build and ship full-stack web apps, APIs, integrations, and backend systems for our clients Own projects end-to-end: architecture, implementation, testing, deployment Juggle multiple client projects at once (AI leverage makes this realistic) Communicate clearly and proactively — updates, blockers, timelines Requirements: US-based. This is a hard requirement — please do not apply if you're not based in the United States. Native or fluent English, excellent written and verbal communication 4+ years professional software engineering experience Strong across a modern stack (examples: TypeScript/React/Next.js, Node, Python, Postgres, cloud — AWS/GCP) Genuine architecture and system-design skills, not just feature-wiring Fluent with AI-assisted development (Claude Code, Codex, Cursor) and able to speak to how you use it to ship faster without sacrificing quality Self-directed, reliable, and able to own work without hand-holding Nice to have: Experience across multiple client projects or agency/consulting background AI/ML integration experience (LLM APIs, RAG, agentic workflows) DevOps / secure cloud deployment experience How we work: Start: paid hourly trial on a real project so we can both evaluate fit Then: ongoing hourly, scaling toward full-time (40 hrs/week) Long-term, stable relationship — we're building a team, not filling a one-off gig To apply, include: A short note on your engineering background and your strongest projects How you specifically use AI tools in your workflow — be concrete, tell us your setup and where it makes you faster Links to work (GitHub, portfolio, shipped products) Your hourly rate Please start your application with the word "SHIPPED" so we know you read this in full. Applications without it will be ignored.
- Hourly: $70.00 - $130.00
- Expert
- Est. time: More than 6 months, Hours to be determined
CONTACTING OUTSIDE OF UPWORK WILL RESULT IN AUTOMATIC DISQUALIFICATION Description: We're an agency staffing data engineering talent across a portfolio of active health-tech and life-sciences platform builds for confidential enterprise clients. You'll be assigned to one or more concurrent engagements based on current need — current active work includes healthcare data integration connecting external clinical/regulatory data sources into a unified platform, and high-throughput consumer health data pipelines processing behavioral/biometric data at scale. Immediate need — looking to onboard within the next 5-7 business days. What you'll do: Build and maintain batch and streaming data pipelines feeding AI/ML and application layers Implement healthcare interoperability standards (FHIR, HL7, X12, NCPDP as applicable) and system-of-record integrations (Veeva, CTMS, eTMF, EMR) Establish data quality, identity resolution, and terminology normalization frameworks Implement consent-aware data handling, including deletion propagation where required Support SSO/SCIM and access provisioning for governed workspaces Required: Strong SQL plus a major data engineering stack (Spark, Kafka/streaming, dbt, or similar) Experience with healthcare data standards (FHIR, HL7, X12, NCPDP) or enterprise system integration Available to start within the next week Nice-to-have: Veeva CRM/Vault/RIM experience, HIPAA/consent management experience, master patient index/identity resolution
- Fixed price
- Expert
- Est. budget: $5,000.00
Vectech builds AI systems for vector surveillance, specializing in mosquitoes and ticks. Our flagship product, IDX, is a lab tool used by vector control organizations to image specimens and identify their species with computer vision. This engagement concerns the C++ firmware of that device. We are seeking a developer or firm to audit the firmware on our deployed IDX fleet: an independent assessment of correctness, robustness, failure behaviour and security posture, with prioritised recommendations. WHY THIS MATTERS The IDX device has been in production for 6 years without an external reliability assessment. As our operations expand, some devices have shown variability in the field, and field failures have required replacing units, which is expensive. Critically: these devices sit at third-party sites on networks we do not control, have no field serial console, and are expensive to recover. A device that fails before its command loop starts cannot be reached at all without manual intervention. We want the assessment weighted toward that operating profile rather than toward defect count. IN SCOPE - Boot and initialisation, including every failure path before the device is remotely reachable - The camera and capture pipeline - Concurrency, memory safety and lifetime correctness across threads - Network behaviour under degraded, filtered and stalled conditions - OTA update and provisioning - Failure modes: what is observable remotely, what self-heals, what needs physical intervention - Security review of network-facing and privileged paths, proportionate to a firmware audit and not a penetration test OUT OF SCOPE - Hardware, electrical, thermal and mechanical design - The ML model, its training and accuracy - The cloud backend and web app, except where firmware depends on them - Implementing the remediation WHAT WE PROVIDE (on NDA, at kickoff) - The firmware repository with tagged history and a written orientation to its layout and build - Toolchain and build tooling - A completed internal static audit with per-finding severity, confidence, evidence and fix ordering - Release history and per-device version inventory - Anonymised field reports with symptoms and resolutions - Provisioning tooling as needed to understand device identity and configuration - An engineer for design-intent questions - Test hardware On the internal audit: it is supplied so you understand what our internal investigation concluded. We want your independent judgement of it - what you would rank differently, what you think is wrong, what it misses. Forming your own conclusions on the highest-risk subsystems before reading it is welcome. REQUIRED EXPERTISE All required unless noted. A proposal that cannot evidence an area is unlikely to succeed. Embedded Linux on NVIDIA Tegra - Jetson / L4T: device tree, bootloader-supplied configuration, kernel module packaging, BSP versioning - Cross-compilation for aarch64 against a target sysroot, at working proficiency, not just able to read - systemd internals: unit ordering, restart policy and rate limiting, watchdog, journald. The target runs an older systemd; check recommendations against its feature set - Diagnosing a headless device with no network output and no console - Flash and SD behaviour: wear, write amplification, durability under power loss, what survives a reboot Camera, ISP and media - GStreamer: read, modify, instrument and debug a pipeline, including one not delivering buffers - The NVIDIA Argus stack (nvarguscamerasrc, Argus daemon): session lifecycle, daemon state, failure modes - MIPI CSI sensor and ISP behaviour. Exposure, gain and white balance are fixed rather than automatic, and image-quality faults are among the reported symptoms - Separating artefacts introduced by the sensor, the ISP, the encoder and our own processing Modern C++ - C++17 at auditing rather than authoring level; roughly 15 modules as shared libraries - Concurrency and lifetime correctness are the priority. The process runs an MQTT event loop, a capture worker, a multi-worker upload pool, an embedded interpreter thread, a Bluetooth thread and a watchdog. Ownership across threads, object versus callback lifetime, promise/future misuse and data races are the defect classes most likely present and least likely to be caught by reading - UB and memory safety, including the discipline to establish reachability rather than reporting a suspected site as certain - Dynamic analysis on aarch64: sanitizers, valgrind or equivalent, on-device or emulated Mixed-language runtime - Embedded CPython via the C API: GIL discipline, interpreter lifecycle, threading constraints - Error handling across the C++/Python boundary. How this goes wrong matters more than general Python skill. The device interpreter is 3.6-era and shipped Python must stay compatible Networking and cloud device management - MQTT v5: session persistence and expiry, reconnection semantics, will messages, client-identity rules - AWS IoT Core: shadows, policies, certificate auth, session semantics. Fleet-scale operational failure modes valued - libcurl in a multithreaded process: timeout semantics, connection versus transfer behaviour, retry design, thread-safety of global init - TLS/mTLS including clock dependence - NetworkManager (libnm) and D-Bus, including GLib/GObject ownership conventions - BLE/BlueZ/GATT. Bluetooth provisioning is the only pre-network path onto a device, so its correctness gates recoverability - Network fault injection: packet loss, DNS interference, captive portals, DPI middleboxes, and connections that handshake then stall. One of the most valuable skills in this list OTA and fleet management - Update mechanism design: atomicity, interruption, rollback, failure partway. The highest-consequence area in the codebase - the failure outcome is an unbootable device - Package signing and verification - The fleet could span multiple firmware versions concurrently, and devices do not necessarily pass through intermediate releases. Analysis must be version-aware and assume no particular prior release ML runtime - familiarity only. TensorRT engines are built and run on-device; enough to reason about build cost, load-time failure and startup impact. Model, training and accuracy expertise not required. Security - proportionate. Threat modelling a root process at a third-party site reachable over a cloud broker; secret handling at rest; input validation on privileged network-facing paths. Methodology and judgement - weighted heavily. Harder to evidence than the technical items, and matters more. - Auditing fielded devices that are expensive to recover, where a low-probability unrecoverable failure outranks a frequent cosmetic one - Reproduction-first: distinguishing "I reproduced this" from "I believe this follows from the code". We want a meaningful proportion of findings reproduced - Ranked, dependency-aware recommendations rather than a severity-labelled list - Assessing the risk of a fix, not only of the defect. With no rollback and limited remote recovery, some corrections are more dangerous than the problem. We expect that reasoning in your findings - Calibrated confidence. Fewer findings you are confident in, plus an explicit list of what you could not determine, beats a long list of mixed quality. We will act on these, so overstated certainty is harmful Valuable but not required: measurement or scientific instrumentation, or any domain where a silently wrong output is worse than a visible failure - much of this device's risk has that shape. Fleet operations experience. Devices in institutional networks with restrictive egress. LevelDB or similar. DELIVERABLES D1. Findings report. Per finding: description, affected version(s), evidence, reproduction status and method, severity and confidence with reasoning, risk of implementing the fix, consequence of not doing so. D2. Reproduction artefacts we can re-run, including any fault-injection tooling built during the engagement. D3. Independent verification of the supplied internal audit - disputed findings, recommendations you consider harmful. D4. Verification recommendations - what testing would have caught these, and the cheapest worthwhile version. D5. Explicit list of what you could not determine, and what access or time would close each item. D6. Walkthrough with our engineering team, plus 30 days of follow-up availability. PROPOSAL FORMAT Eight focused pages beat forty of boilerplate. Please cover: - Understanding of the problem (max 2 pages) - Approach: how you will decide what to reproduce versus reason about - Named individuals who will do the work, with allocation, and the prior projects that evidence each area of the Required Expertise section - Phase plan, milestones and dependencies on us - Assumptions, risks and exclusions - Two references, ideally embedded or fielded-device work Work begins under NDA. Please note in your proposal that you are able to sign one.
- Hourly: $120.00 - $120.00
- Expert
- Est. time: 1 to 3 months, Less than 30 hrs/week
Building a machine learning platform for an insurance related product, with a focus on data pipelines, feature and label design, model development, deployment planning, monitoring, and business impact. Looking for an experienced MLOps or applied ML engineer to provide weekly mentorship through structured project check ins. The implementation will remain entirely my responsibility. Each meeting will focus on the current state of the project. I will provide detailed context on what has been completed, the decisions being considered, current blockers, and the next stage of work. The mentor will review that specific situation, challenge assumptions, identify gaps, and provide direct feedback based on how the issue would be handled in a real production environment. The goal is not general instruction. Feedback should be specific to the project, its architecture, data, constraints, and business use case. Meeting Structure: - Approximately 60 minutes per week - Review of progress since the previous meeting - Discussion of current technical and business decisions - Review of architecture, pipelines, model design, or deployment planning - Identification of risks, missing requirements, and unnecessary complexity - Clear recommendations and next steps for me to complete independently Scope of Guidance: - Business use case and ROI analysis - Data architecture, quality, lineage, and source contracts - Record grain, joins, and entity matching - Feature engineering and leakage prevention - Label and outcome design - Model evaluation and business metrics - Experiment and dataset versioning - Batch and real time deployment - Monitoring, drift, retraining, and rollback - Reliability, privacy, security, and governance Work Expectations: This is a meeting based mentorship role only. The mentor will not be expected to: - perform implementation work outside scheduled meetings - write or maintain the codebase - prepare separate reports or deliverables between meetings - manage the project - provide ongoing asynchronous support - take ownership of delivery Required Experience: - Professional experience building or operating production ML systems - Strong understanding of MLOps, data engineering, deployment, and monitoring - Experience reviewing real technical systems and making practical recommendations - Ability to explain tradeoffs clearly and give direct, specific feedback - Willingness to challenge weak decisions rather than provide generic advice Working Style: Direct communication and practical feedback are important. I will prepare the project context and questions before each meeting, complete the work independently afterward, and return with results for review. Initial Engagement: The engagement will begin with one paid consultation. The session will be used to review the project, discuss the expected mentorship style, and determine whether recurring weekly meetings are a good fit.
- Hourly: $65.00 - $128.00
- Expert
- Est. time: More than 6 months, 30+ hrs/week
AI Architect & Autonomous Agent Engineer (Full-Time, US-Based) Own a Live Production Agent Fleet WHAT THIS IS I run a small, profitable company with an unusual amount of automation behind it. A fleet of autonomous AI agents runs our internal data operation unattended for roughly 12 hours a day, every day. It is real production infrastructure that the business depends on. This is not a "build me a chatbot" job and it is not greenfield. The system exists, it runs daily, and mistakes cost real money. Multiple independent pipelines run in parallel, each doing multi-stage automated research, each calling paid third-party APIs at several points, each with its own quality gates and delivery step. Tens of thousands of records have moved through it. I have been operating and extending this system myself. I need someone to own it so I can stop being the bottleneck. This is an architect role and a builder role at the same time. You will design the system AND write the code AND debug it at 6pm when an agent has done something confident and wrong. There is no team under you to hand it off to. If that split appeals to you, keep reading. I will describe the domain and the specifics on a call, under NDA. What I can tell you publicly is the engineering problem, which is below and is genuinely the interesting part. WHAT YOU WOULD OWN 1. ARCHITECTURE AND AGENT DESIGN - Own the overall design: how the pipelines fit together, where state lives, what runs where, and what happens when any piece fails - Build and maintain autonomous agents that run for hours without a human watching, using Claude Code and Codex - Design the guardrails: quality gates, fail-closed checks, regression tests,and audit trails so an agent cannot silently ship bad work - Debug agents that did the wrong thing confidently, which is the hard part 2. MULTI-DEVICE FLEET ORCHESTRATION - Scale from one machine to many machines running the same pipelines at once - Solve the coordination problems that come with that: shared claim and lock systems so two machines never do the same paid work twice, distributed state, race conditions, safe failure modes - Build the setup and sync tooling so a new machine can be onboarded quickly and every machine runs identical, current logic 3. INTEGRATIONS AND DATA PLUMBING - Cloud spreadsheets and file storage used as coordination and reporting layers across machines - Several third-party vendor APIs, some of them metered and billed per call - Reporting that a non-engineer can actually read and trust 4. QUALITY AND COST CONTROL - Every paid API call should be justified and never duplicated - Build measurement into the system so we know our unit cost and can improve it deliberately, not by guessing WHO THIS IS FOR You will do well here if: - You have shipped agentic systems that run unattended, not just prompts that work in a demo - You think like a systems engineer: idempotency, locking, retries, race conditions, failing closed, and knowing the difference between "it returned 200" and "it actually worked" - You are comfortable in Python, APIs, and the command line - You test your own work adversarially and assume your first answer is wrong - You can explain a technical tradeoff to me in plain language without making me feel stupid or hiding the risk - You are comfortable working on something you cannot put in a public portfolio You will not do well here if you need tickets written for you, if you have only worked on greenfield projects, if you want to architect without implementing, or if you are more excited about model choice than about whether the pipeline is correct at 2am with nobody watching. LOGISTICS - Full-time, long-term. This is an ownership role, not a one-off project. - US-based required. Significant overlap with US Eastern hours. - NDA before we get into specifics. HOW TO APPLY Skip the generic cover letter. I will read all of these and ignore anything that looks templated. Please answer these three questions: 1. Describe an autonomous system you built that ran without supervision. What broke, how did you find out, and what did you change so it could not happen again? 2. Two machines are running the same pipeline against a shared queue of work items. Each item costs money to process. How do you make sure no item is ever paid for twice, and what happens when one machine dies mid-task? 3. What is a mistake an AI agent made in something you built that you did not catch until it had already caused damage? Short and specific beats long and polished. If your answer to #2 is one paragraph and correct, you are ahead of most applicants.