5 часов назад
Senior Staff DevOps Engineer – Orchestration (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Staff DevOps Engineer – Orchestration (AI) (Kubernetes/Terraform): Building and operating the reliability, deployment, and observability infrastructure behind Orchestrator and AI Fabric across on-premises customer datacenters and AWS/GCP with an accent on Kubernetes lifecycle management, CI/CD, infrastructure-as-code, and secure hub-to-edge operations. Focus on debugging distributed production systems, automating certificate and secret rotation, improving incident response, and scaling infrastructure across customer deployments.
Location: US - Headquarters; on-site
Company
builds high-performance infrastructure for demanding artificial intelligence workloads across silicon, systems, and networking.
What you will do
- Own the reliability, deployment, and operational infrastructure for Orchestrator and the AI Fabric environments it manages.
- Provision, upgrade, scale, and harden Kubernetes clusters across bare-metal customer datacenters and AWS/GCP using kubeadm, Rancher, or equivalent tooling.
- Design and operate CI/CD pipelines delivering Go microservices, React UI, Helm charts, and edge appliance images with automated testing gates.
- Manage Terraform-based infrastructure, observability platforms, secrets, certificates, and multi-tenant credential lifecycles.
- Lead incident response, production debugging, root cause analysis, runbook ownership, and post-incident improvements across distributed hub/edge systems.
- Build internal tooling, operational dashboards, capacity plans, and cost-optimization solutions as customer deployments scale.
Requirements
- 8–13 years of experience in SRE, DevOps, or infrastructure engineering for production distributed systems.
- Deep Kubernetes expertise covering cluster administration, networking, storage, RBAC, and pod/node troubleshooting in cloud and bare-metal environments.
- Strong Terraform skills and experience managing multi-environment, multi-provider infrastructure at scale.
- Hands-on ownership of end-to-end CI/CD pipelines and production experience with at least three observability tools such as Prometheus, Grafana, Loki, Splunk, Datadog, or Timestream.
- Strong scripting and automation skills in Python, Bash, or Go, plus Linux systems knowledge covering networking, storage, processes, and performance analysis.
- Production experience with TLS/mTLS certificates, secret rotation, Vault or equivalent systems, and debugging across cloud and on-prem environments.
Nice to have
- Network infrastructure or datacenter automation experience, complex Helm chart authoring, or bare-metal Kubernetes provisioning.
- Experience with eBPF observability tools such as Cilium, Pixie, or Hubble.
- Production experience operating Kafka, ClickHouse, ArangoDB, or Redis.
- Familiarity with SONiC, network switch management, ZTP workflows, PagerDuty, Opsgenie, or infrastructure compliance frameworks such as SOC 2 and FedRAMP.
Culture & Benefits
- Full-time, on-site work at the US headquarters.
- Work on foundational AI infrastructure with high performance and scale requirements.
- Ownership, technical rigor, and fast execution are emphasized.
- Accessibility accommodations are available throughout the hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →