Назад
Company hidden
обновлено 8 дней назад

Senior Software Engineer, Cloud Reliability (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer, Cloud Reliability (AI): Owning the reliability and production stability of Zilliz Cloud, a multi-cloud distributed database platform, with an accent on Kubernetes, cloud infrastructure, observability, and automation. Focus on debugging complex production failures, building diagnostic and remediation tooling, and improving availability and scalability across large multi-tenant systems.

Location: Redwood City, United States; hybrid workplace. Occasional early morning or evening syncs may be required for collaboration across APAC.

Company

hirify.global is a fast-growing startup developing vector database technology and cloud infrastructure for enterprise AI applications.

What you will do

  • Own the reliability, availability, and production stability of hirify.global Cloud.
  • Debug production issues across Kubernetes, cloud infrastructure, networking, storage, and distributed database systems.
  • Build automation and diagnostic tooling for log analysis, alert correlation, incident investigation, runbook automation, and remediation.
  • Turn recurring incidents into reusable tools, documentation, automation, and product improvements.
  • Improve observability for latency, availability, throughput, and resource efficiency.
  • Partner with database and infrastructure engineers to improve reliability, scalability, and automation.

Requirements

  • 3+ years of experience building or operating production cloud systems, infrastructure platforms, database systems, or large-scale online services.
  • Bachelor’s degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.
  • Hands-on experience with Kubernetes, Docker, and at least one major cloud platform: AWS, GCP, or Azure.
  • Strong understanding of distributed systems, availability, scalability, performance, failure recovery, and operational trade-offs.
  • Experience with multi-tenant systems or large infrastructure fleets is valuable.
  • Familiarity with Terraform, Helm, Argo CD, Prometheus, Grafana, and CI/CD systems.

Nice to have

  • Experience with distributed databases, storage systems, search systems, or large-scale online systems.
  • Experience operating thousands of nodes, clusters, tenants, or customer deployments.

Culture & Benefits

  • High ownership of production reliability from end to end.
  • High autonomy, trust, and minimal process.
  • Fast-paced environment with frequent shipping and a focus on execution.
  • Globally distributed collaboration with an on-call setup designed around timezone coverage rather than overnight pages.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →