обновлено 10 дней назад
Lead Software Platform Engineer (AI/ML)
200 000 - 270 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Software Platform Engineer (AI/ML): Building and scaling a multi-tenant AI/ML platform for scientific data, including model serving, LLM agents, and production MLOps infrastructure with an accent on security, evaluation, observability, and regulated pharmaceutical workloads. Focus on designing inference systems, model and prompt lifecycles, tenant data boundaries, cost and latency controls, and reliable deployment under real production traffic.
Location: Remote within the United States; office locations listed in Cambridge, Massachusetts and San Mateo, California.
Salary: $200K–$270K USD per year
Company
Scientific Data and AI Cloud company building lab data management solutions, scientific use cases, and AI-enabled outcomes for pharmaceutical and life sciences organizations.
What you will do
- Own the architecture and API surface of a multi-tenant AI/ML platform used by customers and internal engineering teams.
- Design model and prompt lifecycle management across Databricks MLflow and AWS Bedrock, including registration, versioning, promotion, rollback, and serving.
- Build inference infrastructure for real-time and batch workloads, covering routing, batching, caching, concurrency, accelerator capacity, large binary inputs, and graceful degradation.
- Integrate LLMs, RAG, tool and function calling, MCP tooling, and agent runtimes into production systems.
- Establish security, evaluation, observability, reproducibility, lineage, auditability, and production-readiness practices for regulated AI systems.
- Lead design reviews, technical direction, documentation, incident response, and mentoring across engineering teams.
Requirements
- 10+ years of software and infrastructure engineering experience designing and scaling distributed cloud-native systems.
- Technical leadership or architecture experience with accountability for system design, scalability, performance, and cost.
- Production experience building multi-tenant AI/ML infrastructure and operating LLM systems with RAG, retrieval, embeddings, prompt and model versioning, and tool or function calling.
- Expert coding skills in TypeScript and Python, with experience building robust APIs and backend services.
- Experience with model registries and serving platforms, preferably Databricks MLflow, plus AWS, Docker, CloudFormation or CDK, and CI/CD automation.
- Experience with API-first design, observability, SLI/SLO/SLA practices, sensitive-data security, tenant isolation, and LLM-specific risks.
Nice to have
- Experience with advanced LLM orchestration, agentic systems, MCP, multimodal inputs, fine-tuning, distillation, quantization, batching, or KV-cache optimization.
- Experience with LLM cost attribution, latency optimization, usage analytics, and production AI evaluation.
- Background in regulated or validated environments such as GxP, 21 CFR Part 11, or SOC 2.
- Scientific, life sciences, or laboratory data experience.
Culture & Benefits
- 100% employer-paid benefits for eligible employees and immediate family members.
- Unlimited paid time off.
- 401K, company-paid life insurance, and LTD/STD coverage.
- Remote work and flexible working arrangements.
- Continuous improvement culture with career growth and coaching.
- Visa sponsorship is not currently provided.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →