Назад
Company hidden
обновлено 8 дней назад

Principal Production Engineer (AI)

164 500 - 235 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Production Engineer (AI/Cloud Infrastructure): Designing and operating highly available, scalable infrastructure across AWS, GCP, and bare-metal environments with an accent on automation, observability, and reliability engineering. Focus on building self-healing systems, reducing Mean Time to Mitigate, leading incident response, and scaling globally distributed multi-cloud services.

Location: Remote in California, USA, or hybrid in San Jose, California, with three days per week in the office

Salary: $164,500–$235,000 USD base salary per year, excluding bonus, equity, and benefits

Company

hirify.global provides the Zero Trust Exchange, a cloud security platform that protects users, devices, and applications from cyberattacks and data loss.

What you will do

  • Design and implement highly available, scalable infrastructure across AWS, GCP, and bare-metal environments.
  • Write Python and Go code to eliminate manual toil and build automation-first, self-healing systems.
  • Develop observability using Prometheus, Grafana, and OpenTelemetry; define SLIs, SLOs, and error budgets.
  • Lead incident response as an Incident Commander, maintain response playbooks, and conduct post-incident analyses.
  • Partner with engineering teams on operability reviews and improve the reliability and scalability of a globally distributed platform processing more than 200 billion transactions daily.

Requirements

  • 10+ years of experience managing reliability, scalability, and availability for large-scale production services.
  • Deep programming expertise in Python, Go, or C/C++.
  • Strong knowledge of networking protocols, Linux/RHEL systems, and distributed architectures.
  • Experience with high-stakes incident management and participation in a 24/7 on-call rotation.
  • Experience using ITIL frameworks, incident data, problem management, and technical operability reviews.
  • Foundational AI/ML knowledge and experience leveraging, securing, or positioning AI-driven solutions.

Nice to have

  • Experience with AWS, Azure, GCP, and Infrastructure-as-Code tools such as Ansible, Terraform, Helm, and Temporal.
  • Experience with chaos engineering and large-scale disaster recovery planning.
  • Expertise in BGP, GRE, IPSec, HAProxy, DNS at scale, and OS networking internals.

Culture & Benefits

  • Ownership, collaboration, trust through outcomes, and a challenge culture with ongoing feedback.
  • Health plans, vacation and sick leave, parental leave, retirement options, and education reimbursement.
  • In-office perks and an inclusive workplace focused on collaboration and belonging.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →