Назад
Company hidden
9 часов назад

Senior MLOps Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior MLOps Engineer (AI): Building and operating production-grade infrastructure for real-time training, serving, evaluation, and monitoring of specialized Small Language Models with an accent on low-latency inference, multi-GPU orchestration, and enterprise reliability. Focus on designing scalable model pipelines, optimizing GPU and LLM serving systems, and ensuring reproducibility, observability, security, and audit-ready traceability.

Location: Palo Alto, California, United States; on-site

Company

hirify.global is an enterprise AI product and research company building long-running AI agents powered by specialized Small Language Models for financial audit, accounting, and other professional knowledge workflows.

What you will do

  • Design, build, and operate end-to-end ML infrastructure covering training orchestration, experiment tracking, model registries, model CI/CD, and automated evaluation.
  • Own low-latency, high-throughput LLM and SLM serving infrastructure using batching, caching, and autoscaling.
  • Build and manage multi-GPU training and inference clusters across cloud and on-premises environments, including scheduling, utilization, and cost optimization.
  • Implement production observability for latency, throughput, model drift, regressions, and quality with actionable alerting.
  • Apply inference-time optimizations including quantization, distillation support, KV-cache management, and deployment tuning.
  • Harden the platform for enterprise use through reproducibility, versioning, access controls, audit-ready traceability, and MLOps standards.

Requirements

  • 5+ years of experience in MLOps, ML infrastructure, or platform engineering with substantial production ownership.
  • Production experience deploying and scaling LLM inference infrastructure with serving frameworks such as TRT, vLLM, SGLang, or TGI.
  • Strong proficiency with Kubernetes, Docker, and infrastructure as code such as Terraform.
  • Hands-on experience managing GPU clusters and distributed training or serving environments.
  • Proficiency in Python and experience building maintainable production systems.
  • Experience with ML pipelines, orchestration tools, cloud architecture, and a BS degree in computer science or a related technical field.

Nice to have

  • MS degree in computer science or a related technical field.
  • Experience with multi-node GPU training and quantization techniques such as AWQ, GPTQ, FP8, or GGUF.
  • Experience with Spark, Airflow, fine-tuning workflows for LLMs or VLMs, and RLHF or DPO pipelines.
  • Experience in regulated or enterprise environments where reliability, security, and auditability are critical.
  • Contributions to open-source ML infrastructure projects.

Culture & Benefits

  • Work as an early senior infrastructure team member with significant influence over technical foundations and tooling standards.
  • Build infrastructure used in real enterprise deployments rather than demonstrations.
  • Competitive Silicon Valley-standard salary, significant equity, and premium benefits.
  • Collaborate with a team backed by General Catalyst, Walden Catalyst, and Intel.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →