Назад
Company hidden
2 дня назад

HPC Systems Engineer (AI)

136 300 - 231 700$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
HPC Systems Engineer (AI): Building scalable AI infrastructure, distributed training environments, and model-serving platforms for semiconductor inspection and metrology products with an accent on GPU computing, high-performance storage and networking, and secure orchestration. Focus on optimizing large-scale AI workloads, integrating hardware and software across clusters, and developing reliable infrastructure for model training, inference, RAG pipelines, and agentic AI systems.

Location: Milpitas, California, United States

Salary: $136,300–$231,700 annually

Company

hirify.global develops inspection, measurement, and manufacturing systems for the semiconductor and broader electronics ecosystem, combining physics, data science, and AI.

What you will do

  • Architect, deploy, and operate scalable, secure AI platforms, distributed training environments, and model-serving systems.
  • Evaluate emerging technologies in AI compute, storage, networking, orchestration, and security, and develop reference architectures for enterprise adoption.
  • Integrate AI compute infrastructure and diagnose hardware, firmware, operating system, and platform issues.
  • Design high-performance storage, networking, and data pipelines for large-scale AI workloads.
  • Build observability and monitoring solutions, optimize workload performance, and maintain reliability and availability targets.
  • Implement security-by-design capabilities for AI workloads, intellectual property, and sensitive data while collaborating with research, engineering, IT, and security teams.

Requirements

  • Degree in Computer Science, Computer Engineering, or a related field.
  • 5–8 years of experience in systems engineering, DevOps, or ML infrastructure.
  • Hands-on experience building AI or GPU clusters.
  • Experience with Linux administration, Kubernetes, Docker, Slurm, Ray, TensorFlow, PyTorch, GPU/TPU optimization, containerization, orchestration, and programming.
  • Knowledge of high-speed networking, storage systems, virtualization, server management, model training, fine-tuning, RAG pipelines, and agentic AI systems.
  • Experience with Bash, Python, CI/CD, infrastructure as code, TPM-based encryption, Kubernetes RBAC, OPA/Gatekeeper, and container security.

Culture & Benefits

  • Work in a Central Engineering group focused on AI, physics modeling, high-performance computing, and semiconductor product development.
  • Collaborate with physicists, AI researchers, data scientists, machine learning engineers, software engineers, IT, and security teams.
  • Benefits may include medical, dental, vision, life insurance, 401(k) matching, employee stock purchase, tuition reimbursement, and career development programs.
  • Additional benefits include paid time off, company holidays, family care and bonding leave, wellness programs, financial planning, and student debt assistance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →