Назад
Company hidden
2 дня назад

Manager, Machine Learning Engineering (ML/AI Operations)

219 000 - 329 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Manager, Machine Learning Engineering (ML/AI Operations): Leading the infrastructure, pipelines, and operational practices that keep production machine learning and AI systems reliable, scalable, and safe with an accent on model serving, observability, CI/CD, and team leadership. Focus on building ML operations processes, designing scalable inference and monitoring systems, and balancing generative AI innovation with safety, compliance, cost, and reliability.

Location: San Francisco Bay Area, United States; in-office requirement 3 days per week

Annual salary: $219,000–$329,000 USD, plus equity and benefits.

Company

hirify.global is a community and fundraising platform that helps individuals and nonprofits ask for and provide support.

What you will do

  • Own the reliability, scalability, and operational health of production ML and AI systems, including training pipelines, feature stores, model serving, and observability infrastructure.
  • Lead, hire, and develop ML/AI operations engineers while setting technical direction through architecture decisions, design reviews, and shared engineering practices.
  • Partner with data science and ML engineering teams to improve model deployment through CI/CD, packaging, versioning, rollback, canary, and shadow-release strategies.
  • Establish standards for model monitoring, automated retraining, incident response, SLOs, SLAs, on-call operations, and postmortems.
  • Define the operational strategy for generative AI systems while balancing safety, compliance, cost, innovation speed, and reliability.
  • Manage cloud ML platform and LLM provider relationships, make build-versus-buy decisions, and report reliability and operational risk to engineering leadership.

Requirements

  • 7+ years of hands-on experience building and shipping production machine learning systems, including backend services and ML pipelines in high-availability environments.
  • 1–3+ years of direct engineering management experience, preferably in MLOps, ML platforms, or infrastructure, with experience hiring and developing teams.
  • Strong proficiency in Python and ML frameworks such as PyTorch, TensorFlow, and Scikit-learn, plus software engineering fundamentals including testing, code review, CI/CD, API design, performance, and reliability.
  • Experience operating real-time model serving at scale, including containerization, scalable inference, feature retrieval, and safe rollout strategies.
  • Strong data engineering experience with SQL, Spark/Databricks, Snowflake or similar warehouse technologies, including data quality controls and feature development.
  • Experience with ML monitoring for technical and business metrics such as drift, calibration, segment performance, latency, and error budgets.

Nice to have

  • Experience with generative AI and LLM infrastructure, including latency, cost, and safety guardrails.
  • Master’s or Ph.D. in Computer Science, Statistics, Data Science, or a related technical field.

Culture & Benefits

  • Mission-driven work supporting individuals and nonprofits.
  • Competitive pay with equity, healthcare, dental, vision, life insurance, and a 401(k) savings program.
  • Hybrid-work financial assistance, flexible time off, generous parental leave, and mental health and wellness resources.
  • Learning, development, recognition, volunteering, diversity, equity, and inclusion programs.
  • Collaborative environment guided by values including purpose, trust, excellence, and practical problem-solving.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →