Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Hybrid, requiring 3 days per week in office in the New York City or Boston metro areas; applicants in the Research Triangle or San Francisco Bay Area may also be considered. Candidates must reside in these locations or be willing to relocate.
Total compensation range: $185,500–$232,000 per year, plus equity, benefits, and perks.
Company
is an AI-driven pharmaceutical company developing technology platforms and capabilities to accelerate drug development and clinical trials.
What you will do
- Own the infrastructure and operational platform for shared engineering workloads, including compute, orchestration, deployment, observability, access controls, and reliability.
- Build and operate secure infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software.
- Develop and maintain AWS infrastructure across development, staging, and production, including networking, databases, load balancers, and secrets management.
- Create and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns.
- Establish SLOs, monitoring, alerting, runbooks, incident response, diagnostics, root cause analysis, and post-incident practices.
- Partner with Product Engineering, Data Engineering, and Data Science while mentoring engineers on infrastructure and SRE fundamentals.
Requirements
- 5+ years of experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline.
- Production experience operating cloud infrastructure and distributed systems with strong reliability judgment.
- Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation.
- Experience with AWS and Snowflake.
- Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking.
- Daily fluency with AI tools, including LLMs and agentic coding systems, with strong judgment for validating their output.
Nice to have
- Experience with Azure, GCP, Vercel, or Terragrunt.
- Experience supporting production ML or AI workloads, MLOps infrastructure, workflow orchestration, or model serving.
- Experience operating infrastructure in a regulated or validated environment.
Culture & Benefits
- Hybrid work model with three office days per week.
- Equity and comprehensive benefits.
- Generous perks.
- Opportunity to use AI tools in daily engineering practice and shape AI-native engineering systems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →