6 дней назад
Site Reliability Engineer
145 000 - 175 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Linux/Cloud Infrastructure): Designing and operating highly available developer and production systems across compute, storage, and networking environments with an accent on Linux administration, infrastructure automation, and observability. Focus on responding to outages, leading incident response, maintaining hardware and datacenter infrastructure, and building Bash and Python tooling for reliable large-scale workloads.
Location: Sunnyvale, California, United States; strictly based locally, preferably in the Immediate Bay Area
Salary: $145,000–$175,000 per year in California
Company
is a semiconductor startup building high-performance, efficient graphics processors for immersive content creation, simulation, and consumption.
What you will do
- Design, implement, and operate highly available and fault-tolerant infrastructure and services.
- Install, maintain, and upgrade server, storage, and networking hardware in office and colocation facilities.
- Monitor developer and production environments and proactively remediate reliability risks.
- Participate in on-call rotations and lead incident response, including triage, mitigation, root cause analysis, and corrective action planning.
- Develop and improve automation and operational tooling with Bash and Python.
- Partner with engineering teams to support development, testing, and production workloads at scale.
Requirements
- 5–7 years of experience managing SRE-related functions.
- Expert-level Linux systems administration in complex production environments.
- Advanced Bash and Python scripting, automation, and diagnostic tooling experience.
- Deep knowledge of server hardware, storage subsystems, and datacenter operations.
- Hands-on experience with Proxmox, VMware vSphere, and/or OpenShift; Docker, containerd, and Kubernetes.
- Experience with AWS and/or Microsoft Azure, plus observability tools such as Prometheus and Grafana. Active government clearance or the ability to obtain one is required.
Nice to have
- Familiarity with C, C++, Rust, Go, and/or Julia.
- CompTIA A+, Azure Engineer, or similar certification.
Culture & Benefits
- First-principles approach to problem solving.
- Values include fearlessness, adaptability, and collaborative learning.
- Medical, dental, and vision premiums fully covered.
- Stock options, 401(k) match, and work-from-home hardware reimbursement.
- Visa sponsorship is not currently available for this role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Sr Staff Site Reliability Engineer (AI)
207 400 - 259 200$
6 дней назад
Site Reliability Engineer - Data, Cloud & Developer Experience (AWS)
140 000 - 225 000$
5 дней назад
Senior Site Reliability Engineer (Kubernetes)
147 600 - 221 400$
7 дней назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
7 дней назад
Site Reliability Engineer, Tech Infra
129 960 - 246 240$
8 дней назад
Site Reliability Engineer (Tech Infrastructure)
129 960 - 246 240$