Назад
Company hidden
3 дня назад

Sr. Staff Production Engineer (Cloud Infrastructure)

143 500 - 205 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. Staff Production Engineer (Cloud Infrastructure): Designing and operating highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments with an accent on automation, observability, and distributed systems. Focus on building self-healing systems, reducing Mean Time to Mitigate through incident command and post-incident analysis, and improving reliability across multi-cloud production services.

Location: Hybrid, with three days per week in San Jose, California, or Bellevue, Washington, USA

Salary: $143,500–$205,000 USD base salary per year, excluding bonus, equity, and benefits.

Company

hirify.global provides a cloud-native Zero Trust Exchange platform that uses security data and AI to protect users, devices, and applications from cyberattacks and data loss.

What you will do

  • Design and implement highly available, scalable infrastructure across AWS, Azure, GCP, and bare-metal environments.
  • Write Python and Go code to eliminate manual toil, build self-healing systems, and promote an automation-first operating model.
  • Implement observability with Prometheus, Grafana, and OpenTelemetry; define SLIs, SLOs, and error budgets.
  • Serve as a lead Incident Commander during on-call operations, develop response playbooks, and conduct post-incident analyses.
  • Partner with engineering and other teams on operability reviews and infrastructure standards.

Requirements

  • 8+ years of experience managing reliability, scalability, and availability for large-scale production services.
  • Deep programming expertise in Python, Go, or C/C++.
  • Strong knowledge of networking protocols, Linux or FreeBSD systems, and distributed architecture.
  • Experience with high-stakes incident management and participation in a 24/7 on-call rotation.
  • Experience using ITIL frameworks, incident data, problem management, and technical operability reviews.
  • Foundational understanding of AI/ML technologies and experience applying, securing, or positioning AI-driven solutions.

Nice to have

  • Experience with AI-driven AIOps, anomaly detection, or predictive capacity planning.
  • Expertise in AWS, Azure, GCP, Ansible, Terraform, chaos engineering, or disaster recovery at scale.
  • Knowledge of BGP, GRE, IPSec, HAProxy, DNS at scale, and operating-system networking internals.

Culture & Benefits

  • High-impact, high-accountability environment centered on customer obsession, collaboration, ownership, and transparency.
  • Health plans, vacation and sick time, parental leave, retirement options, and education reimbursement.
  • In-office perks and an inclusive workplace focused on belonging and equal opportunity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →