3 дня назад
Lead Site Reliability Engineer (Performance & Scalability)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (Performance & Scalability) (Distributed Systems): Establishing performance, reliability, and scalability baselines for a production platform with an accent on observability, capacity planning, SLOs, and resilience testing. Focus on identifying bottlenecks, modeling capacity and cost, leading load and failure testing, and preparing systems for high-demand launches.
Location: Remote in the USA
Company
is a full-service consulting firm delivering predictable outcomes and high-quality technology solutions to clients.
What you will do
- Establish performance, throughput, latency, capacity, SLO, error budget, dashboard, alert, and reliability baselines for critical platform workflows.
- Instrument and analyze end-to-end request paths across applications, infrastructure, databases, networking, caches, queues, DNS, registries, and third-party dependencies.
- Identify bottlenecks, lead cross-functional remediation, and drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements.
- Build capacity and cost models, develop demand scenarios, and communicate scaling trade-offs and risks to engineering and executive leadership.
- Lead load, stress, soak, spike, failure, and recovery testing, including automated performance testing and production release gates.
- Own technical readiness assessments, incident investigations, and operational runbooks for launches, scale-up events, rollback, recovery, and dependency failures.
Requirements
- Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline.
- Experience supporting production systems with substantial scale, traffic, latency, or availability requirements.
- Deep knowledge of observability, performance analysis, capacity planning, reliability engineering, cloud infrastructure, and distributed systems.
- Hands-on expertise with databases, networking, caching, queueing, compute, storage, system profiling, bottleneck diagnosis, and architectural tuning.
- Experience defining and operating SLOs, SLIs, error budgets, reliability metrics, load testing, stress testing, scalability testing, and resilience testing.
- Applicants must be authorized to work for any employer in the United States; visa sponsorship is not available.
Nice to have
- Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms.
- Experience creating capacity-cost models and forecasting infrastructure requirements.
- Experience building performance and reliability gates into CI/CD pipelines.
- Experience preparing platforms for significant traffic increases from enterprise customers or strategic partnerships.
- Experience leading reliability or performance initiatives across multiple engineering teams.
Culture & Benefits
- Remote contract engagement for a US-based worker.
- Work is guided by deep expertise, integrity, transparency, and dependability.
- Equal opportunity workplace committed to a diverse and inclusive environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AI Accelerator Infrastructure)
14 часов назад
Lead DevOps Engineer (AWS)
4 дня назад
Site Reliability Engineer
124 000 - 170 500$
Latitude
4 дня назад
Senior Site Reliability Engineer (Kubernetes)
6 дней назад
Lead Site Reliability Engineer (AI)
6 дней назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$