Назад
Company hidden
2 часа назад

Senior Site Reliability Engineer (AI/ML)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Israel
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Site Reliability Engineer (AI/ML): Building and scaling high-scale infrastructure across on-premise and public clouds with an accent on AI/ML Kubernetes environments and Linux performance. Focus on eliminating manual tasks through automation and solving deep-stack bottlenecks from CDN edge to kernel tuning.

Location: Hybrid (Tel Aviv, Israel) — 3 days in-office required

Company

A leading performance-driven advertising technology company reaching approximately 600M daily active users worldwide.

What you will do

  • Maintain high availability and cost-efficiency of hybrid infrastructure, including on-prem, public cloud, and AI/ML clusters.
  • Develop internal software tooling and manage IaC pipelines using Go, Python, or Rust to automate repetitive operations.
  • Perform deep-dive troubleshooting across the full stack, from CDN edge configurations to Linux kernel tuning and network layer bottlenecks.
  • Design and maintain monitoring and alerting setups to identify and address system health issues proactively.
  • Participate in on-call rotations, lead incident resolution, and conduct blameless post-mortems.

Requirements

  • 7+ years of experience managing, scaling, and troubleshooting large-scale distributed Linux environments in production.
  • Deep understanding of Linux system internals and network protocols (TCP/IP, DNS, HTTP, gRPC).
  • Hands-on experience with edge/CDN services such as Fastly, Cloudflare, Akamai, or CloudFront.
  • Proficiency with Infrastructure as Code (IaC) and orchestration tools like Terraform, Ansible, Puppet, ArgoCD, or Jenkins.
  • Production experience managing containerized environments using Kubernetes and Docker.
  • Solid programming skills in at least one modern language: Go, Python, or Rust.

Nice to have

  • Experience designing and operating telemetry, metrics collection, and alerting stacks at scale (Prometheus, Grafana, ELK).
  • Practical background in optimizing infrastructure costs and resource efficiency across cloud and on-prem environments.

Culture & Benefits

  • Hybrid work schedule with 3 days in-office and general flexibility.
  • Comprehensive health benefits and well-being perks.
  • Fully stocked kitchen and local perks including gym partnerships and parking.
  • Opportunity to work with global media leaders such as Yahoo, Fox Sports, and NBCU.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →