Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Storage Engineer (System Engineering): Own and maintain the reliability, performance, and capacity of Lambda's production storage fleet across multiple data centers with an accent on software-defined storage and automation. Focus on incident response, monitoring, CI/CD pipelines, and collaboration with hardware and networking teams to ensure high availability and performance.
Location: Hybrid in San Francisco or San Jose office, presence required 4 days per week
Salary: $267,000–$356,000 per year
Company
Lambda is a leading AI cloud infrastructure startup focused on delivering superintelligence compute power to a wide range of customers, including AI researchers and enterprises.
What you will do
- Own reliability, performance, and capacity health of production storage fleet across all data centers.
- Build and maintain monitoring, dashboards, and alerting for storage systems.
- Investigate and resolve storage-related incidents using telemetry and performance profiling.
- Automate ticketing, escalation, and incident-response workflows to reduce manual triage.
- Design and maintain self-healing automation for common failure modes and capacity rebalancing.
- Implement CI/CD pipelines and collaborate with engineering teams to automate deployment and configuration.
- Work with hardware and networking teams to diagnose low-level I/O and network issues.
- Participate in on-call rotation focusing on reducing mean time to recovery and improving automation.
Requirements
- Must be located in or near San Francisco or San Jose with ability to work onsite 4 days per week.
- 5+ years experience operating Linux systems in production or HPC environments with hands-on storage experience at scale.
- Experience with software-defined storage platforms and their APIs.
- Strong incident-response skills owning production storage incidents end to end.
- Experience with monitoring/logging platforms like Prometheus, Grafana, Datadog, and building dashboards and alerts.
- Experience with Kubernetes, CI/CD tooling, containerization, and Infrastructure as Code tools like Ansible and Terraform.
Nice to have
- Experience with advanced software-defined storage solutions such as VAST or Weka.
- Enterprise storage expertise with NetApp, Dell PowerScale, GPFS, or Lustre.
- Experience writing or operating Kubernetes CSI drivers.
- Experience with SR-IOV, virtualization, GPUDirect Storage, RDMA, InfiniBand, or RoCE networking.
- Familiarity with NIC-level diagnostics and fleet-wide operations tooling.
- Contributions to open-source storage projects.
Culture & Benefits
- Generous cash and equity compensation.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401k plan with 2% company match for US employees.
- Flexible paid time off plan actively used by the team.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Systems and Production Technology Engineer (Post-production)
10 000 - 15 833$
4 дня назад
Senior Infrastructure Engineer (Linux)
131 750 - 170 500$
1 день назад
Server Administrator 4 (HPC)
124 400 - 150 138$
1 день назад
High Performance Computing Engineer (HPC)
147 750 - 221 625$
4 дня назад
IT Operations Engineer (Cloud/Network)
115 000 - 155 000$
Anthropic
6 часов назад
AI Infrastructure Operations, Demand Planning
320 000 - 405 000$