Sr. Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Sr. Site Reliability Engineer (Kubernetes): Designing, building, and evolving resilient platform infrastructure for creator-focused media and advertising products with an accent on Kubernetes operations, infrastructure as code, observability, and scalable architecture. Focus on leading initiatives from concept through production, debugging complex workloads, improving delivery reliability, and optimizing cloud infrastructure costs.
Location: Remote within the United States or in-office in New York City.
Salary: $120,000–$180,000 annual base salary, plus eligible incentive compensation.
Company
A media and advertising technology company that provides creators and publishers with audience, monetization, and business solutions.
What you will do
- Design, build, and evolve the platform and infrastructure powering core products.
- Operate and scale Kubernetes environments, including Helm-based deployments and autoscaling.
- Lead infrastructure initiatives from concept through production while ensuring resilience, observability, and performance.
- Improve CI/CD pipelines, release processes, and engineering standards across platform operations.
- Partner with product and engineering teams to enable fast, reliable delivery.
- Mentor engineers and contribute to technical direction through hands-on engineering and thought leadership.
Requirements
- 8+ years of experience in site reliability engineering, platform engineering, or DevOps.
- Proven experience designing distributed and scalable architectures.
- Hands-on Kubernetes experience, including authoring Helm charts, operating clusters, tuning autoscaling, and debugging production issues.
- Infrastructure-as-code experience with Terraform, including module authoring, CI validation, and environment promotion.
- Experience improving CI/CD pipelines with GitHub Actions, Flux, Argo, or similar tools.
- Experience with observability tools such as Grafana and PagerDuty, plus Prometheus metrics, distributed tracing, and RUM tooling.
Nice to have
- Experience with cloud cost management and identifying infrastructure cost anti-patterns.
- Experience developing secure software using Agile/Scrum methodologies.
- Strong mentoring, communication, collaboration, and proactive problem-solving skills.
- Experience using AI tools such as Claude or Codex to improve engineering velocity and reliability while maintaining security best practices.
Culture & Benefits
- Choice of remote work within the United States or an in-office setup in New York City.
- Work in a collaborative, inclusive environment focused on creators, publishers, and an open internet.
- Opportunity to influence technical direction and raise engineering standards.
- Additional incentive compensation is available beyond the base salary range.
- Commitment to diversity, equity, and inclusion with equal-opportunity employment practices.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →