11 дней назад
Principal Distributed Systems Engineer (Observability)
222 900 - 334 300$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Distributed Systems Engineer (Observability) (ClickHouse/Tempo): Building Workday’s multi-petabyte distributed tracing platform and big-data pipelines with an accent on low-latency querying, high availability, security, and operational reliability. Focus on designing ingestion and storage architecture, scaling distributed systems, and shaping AI-driven anomaly detection, root-cause analysis, and incident triage.
Location: Pleasanton, California, United States; hybrid Flex Work requires at least 50% of each quarter in the office or in the field
Annual base salary: $222,900–$334,300 USD for Pleasanton; $187,100–$334,300 USD for additional US locations. The role may also include bonus compensation and annual refresh stock grants.
Company
is a Fortune 500 company and AI platform for managing people, money, and agents.
What you will do
- Own the technical vision and architecture for distributed tracing across ’s observability platform.
- Architect and build multi-petabyte tracing infrastructure on ClickHouse and Grafana Tempo with sub-second query performance.
- Own Kafka, Spark/Flink, Iceberg-on-S3 ingestion, processing, storage, schema, partitioning, compaction, and lifecycle design.
- Drive performance, scalability, high availability, disaster recovery, security, monitoring, alerting, and capacity planning.
- Evaluate cloud-native and open-source technologies and participate in the platform on-call rotation.
- Partner with ML and AI stakeholders on anomaly detection, root-cause analysis, and AI-assisted incident management while mentoring engineers and setting architectural direction.
Requirements
- 14+ years of software development engineering experience.
- 6+ years designing, building, and operating complex, highly available and fault-tolerant distributed systems.
- 8+ years working with at least two programming languages such as Java, Python, or Go, including production distributed-systems development.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience; a relevant master’s degree is strongly preferred.
- Expertise in distributed-systems principles, API development, high availability, large-scale data processing, scalable systems design, and system security.
- Ability to lead cross-team collaboration, drive architectural direction, and produce technical documentation and presentations.
Culture & Benefits
- Flexible hybrid work combining in-person collaboration with remote flexibility.
- Opportunities for professional growth, skill development, and meaningful long-term work.
- Potential eligibility for the Bonus Plan or role-specific bonus and annual refresh stock grants.
- Reasonable accommodations are available throughout the application process.
- Equal employment opportunity protections apply, including for individuals with disabilities and protected veterans.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Software Engineer, Quantitative Enablement Platform (Vice President)
13 дней назад
Staff Software Engineer (AIOps)
124 000 - 165 000$
13 дней назад
Staff Software Engineer (Java/AWS)
174 000 - 223 000$
11 дней назад
Distributed Systems Engineer (Data Platform)
151 000 - 191 000$
IBM Watson
13 дней назад
Senior Principal Engineer II (Kafka)
192 000 - 358 000$
12 дней назад
Staff Software Engineer (AI)
250 000 - 295 000$