Назад
Company hidden
12 часов назад

Principal Site Reliability Engineer (AWS/Kubernetes)

163 620 - 212 710$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Site Reliability Engineer (AWS/Kubernetes): Building reliable cloud infrastructure, data platforms, self-service developer tools, and CI/CD systems for large-scale TV advertising measurement with an accent on AWS, Kubernetes, Apache Spark, observability, and developer productivity. Focus on designing AIOps capabilities, optimizing distributed data workloads and cloud costs, and driving cross-team reliability and engineering enablement.

Location: Bellevue, WA, United States. Employees near the Bellevue or New York offices work hybrid, typically 1–3 days per week in the office; employees located farther away may work fully remotely. Applicants must already be authorized to work in the United States. Visa sponsorship and transfer of employment-visa sponsorship are not available.

Salary: $163,620–$212,710 USD annually, plus equity and standard benefits.

Company

hirify.global develops TV advertising measurement and impact assessment solutions for brands, agencies, and networks, operating large-scale data infrastructure in AWS.

What you will do

  • Architect, build, and maintain highly available AWS cloud infrastructure and Kubernetes platforms.
  • Improve reliability, performance, and cost efficiency of Apache Spark and other high-volume data processing workloads, including EMR, Databricks, and Glue.
  • Establish observability through SLIs, SLOs, monitoring, alerting, logging, incident response, and post-mortems.
  • Build self-service platforms, internal developer portals, and streamlined CI/CD workflows using Terraform, Kubernetes, Helm, and ArgoCD.
  • Define AIOps and AI developer-tooling practices, including automated remediation, LLM-based toil reduction, root-cause analysis, and governed AI coding standards.
  • Lead the SRE and DevEx roadmap, mentor senior engineers, and align infrastructure, security, and product development teams.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 10+ years of relevant software engineering, cloud architecture, or SRE experience, including at least 3 years in a leadership or lead-contributor role.
  • Deep AWS expertise, including EKS, ECR, RDS, SQS/SNS, VPC, MWAA, and S3, plus strong Terraform or CloudFormation experience.
  • 5+ years of Kubernetes and containerization experience, including kubectl, Helm, and ArgoCD.
  • Production experience tuning Apache Spark workloads for performance, cost, and reliability; knowledge of AWS cost optimization and TCP/IP networking.
  • Experience with CircleCI, shell scripting, Python and/or JavaScript, OTel, Splunk or Datadog, and evaluating GenAI tools for developer productivity.

Nice to have

  • Experience in Ad-Tech or a big-data processing organization.
  • Experience with native AI observability tools.
  • Experience researching developer toolsets and supporting vendor, security, and procurement assessments.

Culture & Benefits

  • Hybrid and flexible workplace with office-based or fully remote arrangements depending on location and responsibilities.
  • Full-time employees are eligible for iSpot’s equity plan and stock options.
  • Eligible roles may include variable compensation, annual bonuses, and pre-approved overtime pay for non-exempt positions.
  • Work alongside experienced engineers with opportunities to influence engineering standards, platforms, and technical strategy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →