Назад
Company hidden
4 дня назад

Senior Site Reliability Engineer (AI)

185 500 - 232 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Релокация
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI): Building and operating secure, observable infrastructure, delivery systems, and operational platforms for product applications, data systems, and ML and AI workloads with an accent on AWS, Kubernetes, infrastructure as code, observability, and reliability. Focus on designing scalable cloud platforms, automating deployment and incident response, and supporting production model training and inference.

Location: Hybrid, requiring 3 days per week in office in the New York City or Boston metro areas; applicants in the Research Triangle or San Francisco Bay Area may also be considered. Candidates must reside in these locations or be willing to relocate.

Total compensation range: $185,500–$232,000 per year, plus equity, benefits, and perks.

Company

hirify.global is an AI-driven pharmaceutical company developing technology platforms and capabilities to accelerate drug development and clinical trials.

What you will do

  • Own the infrastructure and operational platform for shared engineering workloads, including compute, orchestration, deployment, observability, access controls, and reliability.
  • Build and operate secure infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software.
  • Develop and maintain AWS infrastructure across development, staging, and production, including networking, databases, load balancers, and secrets management.
  • Create and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns.
  • Establish SLOs, monitoring, alerting, runbooks, incident response, diagnostics, root cause analysis, and post-incident practices.
  • Partner with Product Engineering, Data Engineering, and Data Science while mentoring engineers on infrastructure and SRE fundamentals.

Requirements

  • 5+ years of experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline.
  • Production experience operating cloud infrastructure and distributed systems with strong reliability judgment.
  • Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation.
  • Experience with AWS and Snowflake.
  • Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking.
  • Daily fluency with AI tools, including LLMs and agentic coding systems, with strong judgment for validating their output.

Nice to have

  • Experience with Azure, GCP, Vercel, or Terragrunt.
  • Experience supporting production ML or AI workloads, MLOps infrastructure, workflow orchestration, or model serving.
  • Experience operating infrastructure in a regulated or validated environment.

Culture & Benefits

  • Hybrid work model with three office days per week.
  • Equity and comprehensive benefits.
  • Generous perks.
  • Opportunity to use AI tools in daily engineering practice and shape AI-native engineering systems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →