8 часов назад
Platform Engineer (AI Infrastructure)
225 000 - 290 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Engineer (AI Infrastructure) (Kubernetes/Python/Go/Rust): Building production platform code for large-scale GPU infrastructure, including Kubernetes operators, control planes, compute, storage, networking, and customer-facing services with an accent on distributed systems, API reliability, and infrastructure automation. Focus on designing reconciliation logic, operating production services, integrating security controls, and improving observability across AI workloads.
Location: Hybrid in Palo Alto, California, United States
Salary: $225,000–$290,000 per year in California, plus equity and a discretionary bonus
Company
is a vertically integrated AI infrastructure platform that builds and operates large-scale GPU compute infrastructure, Kubernetes-native services, virtual machines, storage, and networking.
What you will do
- Design and implement Kubernetes operators and controllers for managing platform resource lifecycles.
- Build production platform capabilities across control planes, APIs, compute, storage, confidential computing, networking, and customer-facing services.
- Translate product roadmap requirements and operational learnings into scalable platform features.
- Integrate security requirements into platform controls and remediate platform-level findings.
- Instrument services, define meaningful metrics, and build observability tooling for platform health.
- Own production services through on-call participation, incident response, code reviews, technical design, and cross-team collaboration.
Requirements
- 3–5 years of software engineering experience, including meaningful experience with infrastructure or platform systems.
- Strong backend or systems programming experience in a production environment.
- Experience with Python, Go, Rust, or another compiled or object-oriented language.
- Solid understanding of Kubernetes internals, including control loops, CRDs, controllers, operators, and reconciliation logic.
- Experience with Linux, networking fundamentals, distributed systems, production APIs, testing, version control, code review, and CI/CD.
- Ability to work in a hybrid role based in Palo Alto, California.
Nice to have
- Experience with Go or Rust, confidential computing, security engineering, high-performance networking, distributed storage, or Kubernetes operator frameworks.
- Experience with Prometheus, Grafana, OpenTelemetry, structured logging, SaaS or PaaS platforms, serverless systems, inference serving, GPU infrastructure, or HPC environments.
- Fluency with AI-assisted development tools and experience working across time zones.
Culture & Benefits
- Supportive environment focused on trust, growth, impact, and work-life balance.
- Equity in and compensation based on the work performed.
- Retirement or pension contributions.
- Comprehensive health, wellbeing, and insurance benefits.
- Generous annual vacation allowance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →