1 месяц назад
Staff Cloud Native Software Engineer (AI Infrastructure)
220 000 - 265 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Cloud Native Software Engineer (AI Infrastructure): Building and operating Kubernetes-native controllers, operators, and networking integrations that connect AI applications with GPU-backed infrastructure with an accent on control-plane extensions, reliability, and observability. Focus on designing reconciliation systems, debugging Kubernetes and networking datapaths, establishing safe rollout patterns, and setting technical direction across teams.
Location: Houston, San Francisco, or Seattle, United States
Salary: $220,000–$265,000 USD per year, plus potential bonus, equity, and/or commission.
Company
Nscale provides high-performance, cost-effective GPU cloud infrastructure for AI startups and enterprise customers.
What you will do
- Design, build, and operate Kubernetes-native controllers, operators, custom resources, admission webhooks, and client-go tooling.
- Extend Kubernetes control-plane capabilities for AI workloads, including network policy controllers, CNI and service-mesh integrations, and scheduling extensions.
- Own major software components end to end and define their design, deployment, operation, and rollout strategy.
- Build observability foundations using metrics, structured events, tracing, and Kubernetes status reporting.
- Develop self-healing systems and operational automation that reduce manual intervention and configuration drift.
- Lead technical design, code reviews, reusable platform patterns, and mentoring across engineering teams.
Requirements
- At least 8 years of production-level software development experience.
- Deep hands-on experience building and operating Kubernetes-native software, including controllers, operators, CRDs, or admission webhooks.
- Strong knowledge of Kubernetes internals, including the API server, informer and lister patterns, reconciliation loops, and the object model.
- Strong networking fundamentals covering CNI, service mesh, kube-proxy or eBPF datapaths, DNS, and load balancing.
- Proficiency in Go or a similar language, with experience shipping well-tested, production-quality code at scale.
- Experience with metrics, tracing, and structured logging, plus the ability to lead technical design and mentor engineers.
Nice to have
- Experience with or strong interest in GPU-backed infrastructure and AI workload patterns.
Culture & Benefits
- Culture centered on innovation, ownership, accountability, openness, and transparency.
- Medical, dental, and vision benefits.
- Flexible paid time off and parental leave.
- Retirement plan participation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →