2 часа назад
Developer Experience Engineer (AI/HPC)
150 000 - 275 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Developer Experience Engineer (AI/HPC): Building automation, observability, CI/CD, and infrastructure systems that accelerate chip design, simulation, and AI model deployment across cloud and on-premise environments with an accent on developer productivity, high-performance computing, and reproducible workflows. Focus on optimizing Slurm-based GPU workloads, Kubernetes clusters, build systems, and hybrid compute infrastructure.
Location: San Jose, United States; fully in-person
Salary: $150,000–$275,000 per year, plus equity
Company
builds hardware, software, racks, and manufacturing systems for frontier AI inference, with a focus on throughput and latency.
What you will do
- Develop automation tools for development, testing, and deployment workflows.
- Optimize Slurm-based scheduling for AI workloads, simulations, and chip design.
- Build observability solutions with Grafana, Prometheus, and OpenTelemetry.
- Manage Docker and Kubernetes environments, including scalability and reproducibility improvements.
- Enhance CI/CD, build, caching, and artifact management systems.
- Integrate AWS and GCP resources and support secrets management, access control, and developer tooling documentation.
Requirements
- Strong Python skills for automation, scripting, and infrastructure development.
- Experience with Slurm job scheduling in HPC or hybrid environments.
- Hands-on experience with Prometheus, Grafana, and OpenTelemetry.
- Expertise with Docker, Kubernetes, Helm charts, and cluster management.
- Experience managing modern CI/CD pipelines with GitHub Actions, Jenkins, or Buildkite.
- Experience with Terraform or Ansible and cloud compute and storage optimization on AWS or GCP.
Nice to have
- AI/ML data pipelines using Airflow, Prefect, or Dagster.
- Build systems such as Bazel, CMake, or distributed build systems.
- Secrets management with Vault, SOPS, AWS Secrets Manager, or GCP Secret Manager.
- AI/ML training workflows, GPU workload monitoring, or FPGA and ASIC development environments.
Culture & Benefits
- Fully in-person work in San Jose, with collaboration across engineering, research, and technical disciplines.
- Full medical, dental, and vision coverage with generous premium support.
- $2,000 monthly housing subsidy for employees living within walking distance of the office.
- Daily lunch and dinner at the office.
- Relocation support for moves to West San Jose.
- Unlimited compute budget subject to ROI justification.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior AI DevOps Developer (AI)
150 000 - 206 000$
6 часов назад
Forward Deployed Engineers (AI Infrastructure)
180 000 - 240 000$
5 часов назад
Sr. DevOps Engineer (AI)
175 000 - 195 000$
6 часов назад
Senior DevOps Engineer (AI)
170 000 - 185 000$
Windsurf
6 дней назад
Site Reliability Engineer (AI)
4 часа назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$