Principal Site Reliability Engineer (AWS/Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: Bellevue, WA, United States. Employees near the Bellevue or New York offices work hybrid, typically 1–3 days per week in the office; employees located farther away may work fully remotely. Applicants must already be authorized to work in the United States. Visa sponsorship and transfer of employment-visa sponsorship are not available.
Salary: $163,620–$212,710 USD annually, plus equity and standard benefits.
Company
develops TV advertising measurement and impact assessment solutions for brands, agencies, and networks, operating large-scale data infrastructure in AWS.
What you will do
- Architect, build, and maintain highly available AWS cloud infrastructure and Kubernetes platforms.
- Improve reliability, performance, and cost efficiency of Apache Spark and other high-volume data processing workloads, including EMR, Databricks, and Glue.
- Establish observability through SLIs, SLOs, monitoring, alerting, logging, incident response, and post-mortems.
- Build self-service platforms, internal developer portals, and streamlined CI/CD workflows using Terraform, Kubernetes, Helm, and ArgoCD.
- Define AIOps and AI developer-tooling practices, including automated remediation, LLM-based toil reduction, root-cause analysis, and governed AI coding standards.
- Lead the SRE and DevEx roadmap, mentor senior engineers, and align infrastructure, security, and product development teams.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 10+ years of relevant software engineering, cloud architecture, or SRE experience, including at least 3 years in a leadership or lead-contributor role.
- Deep AWS expertise, including EKS, ECR, RDS, SQS/SNS, VPC, MWAA, and S3, plus strong Terraform or CloudFormation experience.
- 5+ years of Kubernetes and containerization experience, including kubectl, Helm, and ArgoCD.
- Production experience tuning Apache Spark workloads for performance, cost, and reliability; knowledge of AWS cost optimization and TCP/IP networking.
- Experience with CircleCI, shell scripting, Python and/or JavaScript, OTel, Splunk or Datadog, and evaluating GenAI tools for developer productivity.
Nice to have
- Experience in Ad-Tech or a big-data processing organization.
- Experience with native AI observability tools.
- Experience researching developer toolsets and supporting vendor, security, and procurement assessments.
Culture & Benefits
- Hybrid and flexible workplace with office-based or fully remote arrangements depending on location and responsibilities.
- Full-time employees are eligible for iSpot’s equity plan and stock options.
- Eligible roles may include variable compensation, annual bonuses, and pre-approved overtime pay for non-exempt positions.
- Work alongside experienced engineers with opportunities to influence engineering standards, platforms, and technical strategy.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →