9 часов назад
Senior MLOps Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (AI): Building and operating production-grade infrastructure for real-time training, serving, evaluation, and monitoring of specialized Small Language Models with an accent on low-latency inference, multi-GPU orchestration, and enterprise reliability. Focus on designing scalable model pipelines, optimizing GPU and LLM serving systems, and ensuring reproducibility, observability, security, and audit-ready traceability.
Location: Palo Alto, California, United States; on-site
Company
is an enterprise AI product and research company building long-running AI agents powered by specialized Small Language Models for financial audit, accounting, and other professional knowledge workflows.
What you will do
- Design, build, and operate end-to-end ML infrastructure covering training orchestration, experiment tracking, model registries, model CI/CD, and automated evaluation.
- Own low-latency, high-throughput LLM and SLM serving infrastructure using batching, caching, and autoscaling.
- Build and manage multi-GPU training and inference clusters across cloud and on-premises environments, including scheduling, utilization, and cost optimization.
- Implement production observability for latency, throughput, model drift, regressions, and quality with actionable alerting.
- Apply inference-time optimizations including quantization, distillation support, KV-cache management, and deployment tuning.
- Harden the platform for enterprise use through reproducibility, versioning, access controls, audit-ready traceability, and MLOps standards.
Requirements
- 5+ years of experience in MLOps, ML infrastructure, or platform engineering with substantial production ownership.
- Production experience deploying and scaling LLM inference infrastructure with serving frameworks such as TRT, vLLM, SGLang, or TGI.
- Strong proficiency with Kubernetes, Docker, and infrastructure as code such as Terraform.
- Hands-on experience managing GPU clusters and distributed training or serving environments.
- Proficiency in Python and experience building maintainable production systems.
- Experience with ML pipelines, orchestration tools, cloud architecture, and a BS degree in computer science or a related technical field.
Nice to have
- MS degree in computer science or a related technical field.
- Experience with multi-node GPU training and quantization techniques such as AWQ, GPTQ, FP8, or GGUF.
- Experience with Spark, Airflow, fine-tuning workflows for LLMs or VLMs, and RLHF or DPO pipelines.
- Experience in regulated or enterprise environments where reliability, security, and auditability are critical.
- Contributions to open-source ML infrastructure projects.
Culture & Benefits
- Work as an early senior infrastructure team member with significant influence over technical foundations and tooling standards.
- Build infrastructure used in real enterprise deployments rather than demonstrations.
- Competitive Silicon Valley-standard salary, significant equity, and premium benefits.
- Collaborate with a team backed by General Catalyst, Walden Catalyst, and Intel.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →