Site Reliability Engineer (Big Data)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Site Reliability Engineer (Big Data): Building and optimizing streaming and storage systems with an accent on Kafka architecture and Ceph operations. Focus on automating operational workflows, managing petabyte-scale distributed systems, and ensuring high availability across hybrid infrastructure.
Location: Remote (International). Willingness to work 9am-6pm ET U.S. hours is required.
Company
is a high-scale data platform processing billions of events daily using Kafka, Hadoop, and Ceph.
What you will do
- Manage Kafka architecture, including topic design, partition strategy, throughput, and latency optimization.
- Oversee Ceph operations, pool design, placement optimization, and capacity planning.
- Develop operational automation to reduce manual work and accelerate incident response.
- Maintain SQL Server backup and recovery pipelines and provide basic cluster support.
- Build self-service tooling and observability frameworks for the data team.
Requirements
- 5+ years of experience operating large-scale distributed systems in production.
- Deep expertise in Kafka, Ceph, or similar distributed infrastructure.
- Proven ability to design for scale and reliability at petabyte scale.
- Experience mentoring engineers and making critical technical decisions.
- Must be willing to work 9am-6pm ET U.S. hours.
Nice to have
- Experience with multi-region replication and disaster recovery.
- Cost optimization expertise at infrastructure scale.
- Experience with hybrid on-prem/cloud operations.
Culture & Benefits
- Fully remote work arrangement.
- Opportunity to define platform architecture and influence engineering standards.
- High-impact role as a key technical contributor shaping the future of the platform.
- Exposure to large-scale distributed systems challenges.
Hiring process
- Introductory conversation to discuss background and the role.
- Technical discussion focused on systems engineering and Kubernetes.
- Architecture discussion exploring platform design and technical decision-making.
- Leadership conversation regarding team strategy and long-term direction.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →