Назад
Company hidden
3 дня назад

Senior Infrastructure Engineer, SRE (AWS)

150 000 - 185 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Infrastructure Engineer, SRE (AWS) (Cloud Infrastructure and Reliability): Building and evolving reliability, observability, and disaster recovery capabilities for a platform running hundreds of production services and processing billions of transactions, with an accent on SLIs, SLOs, error budgets, and resilient operations. Focus on designing failover and restore paths, improving metrics, tracing, and logs, strengthening incident response, and enabling engineering teams to operate services independently.

Location: Remote within the USA; office locations include San Francisco, Washington, D.C., and New York City.

Salary: $150,000–$185,000 per year plus bonus and benefits.

Company

hirify.global provides financial products and services that help millions of people better manage their finances.

What you will do

  • Build and improve the reliability and resiliency of hundreds of production services operating at large scale.
  • Define SLIs, SLOs, and error budgets for critical services and user journeys in partnership with product engineering teams.
  • Own the disaster recovery strategy, including recovery objectives, failover and restore paths, and regular exercises.
  • Evolve observability standards for metrics, tracing, logs, instrumentation, alert quality, and cost.
  • Strengthen incident response by tuning paging thresholds, maintaining runbooks, and following up on postmortem actions.
  • Contribute to cloud infrastructure work, platform improvements, automation, and a shared on-call rotation of one week every six weeks.

Requirements

  • 5+ years of hands-on cloud or infrastructure engineering experience, including substantial production reliability and operations work at scale.
  • Experience defining SLIs and SLOs for real production services and evaluating their operational impact.
  • Production experience with an observability platform; Datadog is strongly preferred.
  • Ability to write code in Python, Go, TypeScript, or a similar language for tooling, debugging, and automation.
  • Production Terraform experience and strong AWS skills.
  • Experience building or operating disaster recovery plans, conducting failover and restore drills, and participating in on-call rotations.

Nice to have

  • Leadership of reliability or observability modernization projects.
  • Experience building internal tooling, libraries, or instrumentation standards.
  • Experience running game days, chaos experiments, or disaster recovery exercises.
  • Experience reducing observability costs while maintaining coverage.

Culture & Benefits

  • Health, dental, and vision plans.
  • 401(k) matching.
  • Unlimited paid time off.
  • Competitive pay, bonus, and additional benefits.
  • Daily lunch, snacks, coffee, and commuter benefits for in-office employees.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →