Назад
Company hidden
3 дня назад

Lead Site Reliability Engineer (Performance & Scalability)

Формат работы
remote (только USA)
Тип работы
project
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead Site Reliability Engineer (Performance & Scalability) (Distributed Systems): Establishing performance, reliability, and scalability baselines for a production platform with an accent on observability, capacity planning, SLOs, and resilience testing. Focus on identifying bottlenecks, modeling capacity and cost, leading load and failure testing, and preparing systems for high-demand launches.

Location: Remote in the USA

Company

hirify.global is a full-service consulting firm delivering predictable outcomes and high-quality technology solutions to clients.

What you will do

  • Establish performance, throughput, latency, capacity, SLO, error budget, dashboard, alert, and reliability baselines for critical platform workflows.
  • Instrument and analyze end-to-end request paths across applications, infrastructure, databases, networking, caches, queues, DNS, registries, and third-party dependencies.
  • Identify bottlenecks, lead cross-functional remediation, and drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements.
  • Build capacity and cost models, develop demand scenarios, and communicate scaling trade-offs and risks to engineering and executive leadership.
  • Lead load, stress, soak, spike, failure, and recovery testing, including automated performance testing and production release gates.
  • Own technical readiness assessments, incident investigations, and operational runbooks for launches, scale-up events, rollback, recovery, and dependency failures.

Requirements

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline.
  • Experience supporting production systems with substantial scale, traffic, latency, or availability requirements.
  • Deep knowledge of observability, performance analysis, capacity planning, reliability engineering, cloud infrastructure, and distributed systems.
  • Hands-on expertise with databases, networking, caching, queueing, compute, storage, system profiling, bottleneck diagnosis, and architectural tuning.
  • Experience defining and operating SLOs, SLIs, error budgets, reliability metrics, load testing, stress testing, scalability testing, and resilience testing.
  • Applicants must be authorized to work for any employer in the United States; visa sponsorship is not available.

Nice to have

  • Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms.
  • Experience creating capacity-cost models and forecasting infrastructure requirements.
  • Experience building performance and reliability gates into CI/CD pipelines.
  • Experience preparing platforms for significant traffic increases from enterprise customers or strategic partnerships.
  • Experience leading reliability or performance initiatives across multiple engineering teams.

Culture & Benefits

  • Remote contract engagement for a US-based worker.
  • Work is guided by deep expertise, integrity, transparency, and dependability.
  • Equal opportunity workplace committed to a diverse and inclusive environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →