Назад
Company hidden
обновлено 5 дней назад

Sr. Production Engineer (Cybersecurity)

118 400 - 148 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. Production Engineer (Cloud Infrastructure and Cybersecurity): Building highly available, scalable infrastructure and self-healing systems across AWS, GCP, and bare-metal environments with an accent on automation, observability, and distributed architecture. Focus on reducing Mean Time to Mitigate, leading incident response, defining SLIs/SLOs, and improving reliability across a globally distributed multi-cloud platform.

Location: Remote within the USA or hybrid, with 3 days per week in San Jose, California

Base salary: $118,400–$148,000 USD per year, excluding bonus, equity, commission, and benefits

Company

hirify.global provides a cloud-native Zero Trust Exchange platform that secures users, devices, and applications through cybersecurity and AI-powered systems.

What you will do

  • Implement highly available and scalable infrastructure across AWS, GCP, and bare-metal environments.
  • Write Python or Go code to eliminate manual work and build self-healing systems.
  • Develop observability using Prometheus, Grafana, and OpenTelemetry; define SLIs, SLOs, and error budgets.
  • Lead incident response as an Incident Commander, create response playbooks, and conduct post-incident analyses.
  • Partner with engineering and other teams on operability reviews and architectural standards.

Requirements

  • 3–5+ years of experience managing reliability, scalability, and availability for large-scale production services.
  • Deep programming expertise in Python, Go, or C/C++.
  • Strong knowledge of networking protocols, Linux/RHEL systems, and distributed architectures.
  • Experience with high-stakes incident management and participation in a 24/7 on-call rotation.
  • Experience using ITIL frameworks and incident data for problem management, service maturity, and technical operability reviews.
  • Active use and integration of AI tools to improve workflows and problem-solving.

Nice to have

  • Experience with AWS, Azure, or GCP and infrastructure-as-code tools such as Ansible, Terraform, Helm, or Temporal.
  • Experience with chaos engineering and large-scale disaster recovery planning.
  • Knowledge of BGP, GRE, IPSec, HAProxy, DNS at scale, and operating-system networking internals.

Culture & Benefits

  • Automation-first culture focused on ownership, accountability, customer obsession, and constructive debate.
  • Health plans, vacation and sick time, parental leave, retirement options, and education reimbursement.
  • In-office perks and an inclusive workplace focused on collaboration and belonging.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →