6 дней назад
AIOps Observability/SRE Lead
187 000 - 270 700$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AIOps Observability/SRE Lead (AIOps, SRE, Cloud Networking): Building and leading enterprise reliability and observability capabilities across critical IT and engineering systems with an accent on monitoring architecture, automation, network transformation, and hybrid cloud connectivity. Focus on defining SLOs and error budgets, managing incident response and 24x7 operations, and designing secure, scalable integrations across AWS, Azure, and on-premises infrastructure.
Location: Onsite in San Jose, California, United States; regular in-office work is required.
Salary: $187,000–$270,700 USD annually for the Bay Area, California.
Company
is a pure-play FPGA solutions provider delivering high-performance, flexible FPGA solutions for complex computing challenges.
What you will do
- Lead, hire, mentor, and develop the SRE team and reliability engineering capabilities.
- Define SRE practices, including SLIs, SLOs, error budgets, reliability targets, and operational maturity reporting.
- Establish unified observability standards for monitoring, alerting, logging, tracing, and incident management.
- Drive automation, toil reduction, platform scalability, and continuous reliability improvement.
- Design reliable connectivity and network transformation solutions across cloud, on-premises infrastructure, engineering, integration, and security environments.
- Provide technical leadership for architecture, migration, optimization, documentation, security controls, and operational procedures.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related technical field.
- 10+ years of professional experience in enterprise networking and 10+ years in SRE or DevOps focused on platform reliability.
- Experience leading technical engineering teams in a senior or management capacity.
- Strong knowledge of SRE principles, observability, incident management, blameless post-mortems, and automation.
- Experience with AWS, Azure, containerized environments, routing and switching, hybrid cloud connectivity, and large-scale on-premises infrastructure.
- Applicants must be eligible for any required U.S. export authorizations.
Nice to have
- Experience with semiconductor, HPC, or EDA-dependent environments.
- Familiarity with AIOps platforms and AI-assisted incident management tools.
- Knowledge of chaos engineering and resiliency testing frameworks.
Culture & Benefits
- Participation in incentive opportunities based on individual and company performance.
- Collaboration with IT, engineering, integration, cybersecurity, and external service providers.
- Focus on secure, scalable, reliable, and operationally efficient infrastructure.
- Support for 24x7 operations through structured on-call and escalation processes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Staff Engineer (SRE)
110 000 - 230 000$
12 дней назад
Lead Site Reliability Engineer (Cybersecurity)
145 000 - 200 000$
10 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
6 дней назад
Senior Staff Engineer (SRE)
120 000 - 260 000$
6 дней назад
Senior Engineer (SRE/Incident Management)
100 000 - 215 000$
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$