Lead Software Platform Engineer (AI/ML)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Lead Software Platform Engineer (AI/ML): Building and scaling multi-tenant AI/ML infrastructure and cloud services that enable production-grade models, LLMs, and agents to run against scientific data with an accent on distributed systems, MLOps, security, and observability. Focus on designing model-serving platforms, evaluation and deployment workflows, tenant data boundaries, and cost- and latency-efficient inference for regulated pharmaceutical environments.
Location: Remote within the United States; listed locations include Cambridge, Massachusetts, and San Mateo, California. Visa sponsorship is not currently provided.
Salary: $200,000–$270,000 USD
Company
provides a scientific data and AI cloud with lab data management solutions and AI-enabled tools for life sciences organizations.
What you will do
- Own the architecture and API surface of the multi-tenant AI/ML platform used by customers and internal engineering teams.
- Manage model and prompt lifecycles across Databricks MLflow and AWS Bedrock, including versioning, promotion, rollback, and multi-model serving.
- Design real-time and batch inference infrastructure with routing, batching, caching, concurrency control, accelerator capacity planning, and graceful degradation.
- Build production LLM and agent capabilities using RAG, tool and function calling, MCP, and agent runtimes.
- Establish security, evaluation, observability, reliability, lineage, and auditability for AI systems operating in regulated environments.
- Lead design reviews, define technical direction, document reference architectures, mentor engineers, and support production readiness and incident response.
Requirements
- 10+ years of professional software and infrastructure engineering experience designing and scaling distributed cloud-native systems.
- Technical leadership or architecture experience with accountability for system design, scalability, performance, and cost optimization.
- Production experience building multi-tenant AI/ML infrastructure as a customer-facing product, including model deployment and lifecycle management.
- Hands-on experience taking LLM systems to production with RAG, retrieval and embeddings, prompt and model versioning, and tool or function calling.
- Expert coding skills in TypeScript and Python, plus API-first design with REST and OpenAPI.
- Strong AWS, Docker, infrastructure-as-code, CI/CD, observability, and SLI/SLO/SLA experience; familiarity with sensitive-data security, tenant isolation, and LLM risks is required.
Nice to have
- Experience with advanced LLM orchestration frameworks, production agents, MCP, and multimodal model inputs.
- Knowledge of LLM cost attribution, latency optimization, fine-tuning, distillation, quantization, batching, or KV-cache strategies.
- Experience delivering AI systems in regulated or validated environments such as GxP, 21 CFR Part 11, or SOC 2.
- Background in scientific, life sciences, or laboratory data domains.
Culture & Benefits
- Remote work with flexible working arrangements.
- 100% employer-paid benefits for eligible employees and immediate family members.
- Unlimited paid time off.
- Company-paid life insurance and long-term and short-term disability coverage.
- 401(k) and a culture of continuous improvement with coaching and career development.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →