Назад
Company hidden
обновлено 2 дня назад

Staff Production Engineer (SRE) (Federal)

119 000 - 170 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Production Engineer (SRE) (Federal) (Linux/Kubernetes): Maintaining and improving the reliability and performance of high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions, with an accent on systems-level troubleshooting, automation, and observability. Focus on incident response across OS and network layers, infrastructure lifecycle automation, SLO enforcement, kernel upgrades, and capacity tuning.

Location: Hybrid, onsite three days a week in San Jose, California, or another hirify.global office in the United States; remote may be considered for exceptional candidates. US citizenship is required.

Base salary: $119,000–$170,000 USD per year, excluding bonus, equity, commission, and benefits.

Company

hirify.global provides a cloud-native Zero Trust Exchange platform that helps organizations securely connect users, devices, and applications and protect against cyberattacks and data loss.

What you will do

  • Own systems-level reliability and performance for high-throughput bare-metal and cloud infrastructure processing tens of billions of daily transactions.
  • Partner with Engineering and Networking teams to maintain high availability across Linux/BSD fleets, Kubernetes clusters, and custom routing stacks.
  • Lead incident response and troubleshoot live production issues across OS and network layers using tools such as strace, lsof, tcpdump, iostat, vmstat, and gdb.
  • Automate infrastructure lifecycle management, service provisioning, configuration workflows, and releases with Ansible, Python, and Bash.
  • Operate metrics, logs, and traces with Prometheus and OpenTelemetry, and define SLOs and error budgets to reduce alert noise.
  • Conduct architectural reviews, OS and kernel upgrades, capacity and performance tuning, and CI/CD validation before production rollouts.

Requirements

  • US citizenship is required due to the nature of assigned customers.
  • 5+ years of experience in Site Reliability Engineering, Production Engineering, or Systems Engineering for high-scale, low-latency production platforms.
  • Ability to write and debug executable code live in Python, Go, or Bash, including core logic and data structures.
  • Hands-on experience writing Ansible playbooks and tasks for infrastructure automation.
  • Deep knowledge of Linux internals, kernel troubleshooting, networking protocols, DNS, TLS, TCP/IP, and packet-level analysis with tcpdump.
  • Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions.

Nice to have

  • Production experience with FreeBSD or other BSD operating systems.
  • Experience running, scaling, and troubleshooting Kubernetes clusters in high-traffic environments.
  • Deep experience with Prometheus and OpenTelemetry ecosystems, AI/ML frameworks, or AIOps tools for automated root-cause analysis.

Culture & Benefits

  • Health plans and time off for vacation and sick leave.
  • Parental leave options and retirement plans.
  • Education reimbursement and in-office perks.
  • Culture centered on customer focus, collaboration, ownership, accountability, and constructive debate.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →