Назад
Company hidden
49 минут назад

Senior MLOps & AI Infrastructure Engineer (AI)

149 100 - 215 925$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior MLOps & AI Infrastructure Engineer (AI/ML): Building and operationalizing scalable machine learning pipelines, model lifecycle systems, and AI infrastructure across cloud and on-prem HPC environments with an accent on LLMs, GNNs, reinforcement learning, and production model efficiency. Focus on automating training and deployment, optimizing GPU/TPU inference, managing large-scale data and feature pipelines, and productionizing AI for EDA and chip design.

Location: San Jose, California, United States. Applicants must be eligible for any required U.S. export authorizations.

Salary: $149,100–$215,925 USD per year for the Bay Area, California.

Company

hirify.global is a pure-play FPGA solutions provider developing programmable technologies for AI, cloud, networking, and edge markets.

What you will do

  • Design, build, and maintain scalable ML pipelines for training, evaluation, deployment, and continuous training across cloud and on-prem HPC environments.
  • Build MLOps infrastructure for experiment tracking, model registries, feature stores, retraining, data versioning, and lineage.
  • Develop, fine-tune, and deploy LLMs, GNNs, and reinforcement learning agents for EDA and chip design applications.
  • Containerize and orchestrate ML workloads with Docker, Kubernetes, and GPU node pools while optimizing cloud infrastructure costs and performance.
  • Implement model monitoring, alerting, observability, A/B testing, shadow deployments, and drift detection.
  • Partner with research, software, data, and infrastructure teams; mentor junior engineers and establish ML engineering practices.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Statistics, or a related field, plus 10+ years of industry experience.
  • 10+ years across ML engineering, data science, and MLOps, including PyTorch, TensorFlow, JAX, Hugging Face, and production model deployment at scale.
  • 8+ years of experience with parallelism strategies such as FSDP, DeepSpeed, and data/model parallelism.
  • 10+ years of Python experience and 8+ years with cloud ML platforms, Docker, Kubernetes, and CI/CD pipelines.
  • 5+ years of hands-on experience with MLflow, Weights & Biases, or Neptune.
  • Eligibility for any required U.S. export authorizations is required.

Nice to have

  • Experience applying AI/ML to semiconductor, EDA, or chip design workflows.
  • Experience with HPC schedulers such as LSF or Slurm and GPU cluster management.
  • Knowledge of LLM fine-tuning, RAG, AI agent frameworks, GNNs, geometric deep learning, or reinforcement learning.
  • Experience with DevSecOps, zero-trust security, compliance automation, simulation pipelines, or synthetic data generation.
  • Familiarity with Synopsys, Cadence, or Siemens EDA toolchains; published research or open-source contributions in ML, MLOps, or AI for EDA.

Culture & Benefits

  • Regular full-time employment in a deep-tech and semiconductor environment.
  • Incentive opportunities based on individual and company performance.
  • Work spans cloud, on-prem HPC, EDA, simulation, and chip design environments.
  • AI is used to screen, assess, or select applicants.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →