обновлено 8 дней назад
Senior Software Engineer, Cloud Reliability (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Engineer, Cloud Reliability (AI): Owning the reliability and production stability of Zilliz Cloud, a multi-cloud distributed database platform, with an accent on Kubernetes, cloud infrastructure, observability, and automation. Focus on debugging complex production failures, building diagnostic and remediation tooling, and improving availability and scalability across large multi-tenant systems.
Location: Redwood City, United States; hybrid workplace. Occasional early morning or evening syncs may be required for collaboration across APAC.
Company
is a fast-growing startup developing vector database technology and cloud infrastructure for enterprise AI applications.
What you will do
- Own the reliability, availability, and production stability of Cloud.
- Debug production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems.
- Build automation and diagnostic tooling for log analysis, alert correlation, incident investigation, runbook automation, and remediation.
- Turn recurring incidents into reusable tools, documentation, automation, and product improvements.
- Improve observability for latency, availability, throughput, and resource efficiency.
- Partner with database and infrastructure engineers to improve reliability, scalability, and automation.
Requirements
- 3+ years of experience building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services.
- Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.
- Hands-on experience with Kubernetes, Docker, and at least one major cloud platform: AWS, GCP, or Azure.
- Strong understanding of distributed systems, availability, scalability, performance, failure recovery, and operational trade-offs.
- Experience with multi-tenant systems or large infrastructure fleets is valuable.
- Familiarity with Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems.
Nice to have
- Experience with distributed databases, storage systems, search systems, or large-scale online systems.
- Experience operating thousands of nodes, clusters, tenants, or customer deployments.
Culture & Benefits
- High ownership of production reliability from end to end.
- High autonomy, trust, and minimal process.
- Fast-paced environment with frequent shipping and a focus on execution.
- Globally distributed collaboration with an on-call setup designed around timezone coverage rather than overnight pages.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
11 дней назад
Sr. Site Reliability Engineer (AI)
11 дней назад
Principal Site Reliability Engineer (Kubernetes)
190 000 - 220 000$
10 дней назад
Site Reliability Engineer I (Azure)
12 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$