1 час назад
Staff ML Platform Engineer (MLOps)
172 000 - 215 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff ML Platform Engineer (MLOps) (AI/SaaS): Building the platform that powers batch models, real-time recommendations, and LLM-powered products with an accent on reproducible infrastructure, model deployment, evaluation, observability, and cost-aware LLM routing. Focus on designing controlled rollouts, tracing predictions to their inputs, operating production ML systems, and keeping training and inference features consistent.
Location: Remote in the United States or Canada. Toronto-based candidates may optionally work from the Toronto office; approximately 2–3 trips per year are expected.
Salary: USD $172,000–$215,000 in the United States / CAD $172,000–$220,000 in Canada
Company
is a SaaS and AI company building workforce development technology that helps people find better jobs and supports organizations serving job seekers.
What you will do
- Assess existing data pipelines, ML workflows, and architecture, then create and execute a prioritized platform plan.
- Build reproducible compute, training, serving, deployment, and environment-management foundations.
- Develop LLM infrastructure with model routing, prompt and response evaluation, and cost and latency optimization.
- Establish safe model experimentation through A/B tests, shadow deployments, canaries, holdouts, and predefined success criteria.
- Implement monitoring, alerting, regression detection, lineage, and traceability for production recommendations.
- Operate production ML systems, debug incidents, and maintain consistent feature computation between training and inference.
Requirements
- Staff-level hands-on experience in MLOps, ML platforms, ML infrastructure, data platforms, or equivalent platform ownership.
- Experience building model CI/CD, experiment tracking, registries, deployment workflows, and monitoring end to end.
- Production experience with LLM systems, including serving, evaluation, model routing, and cost and latency tradeoffs.
- Experience operating both batch and real-time models, including containers, reproducible environments, and compute provisioning.
- Strong observability, traceability, data engineering, systems design, and production incident-debugging experience.
- Must be located in the United States or Canada.
Nice to have
- Feature store experience and interest in growing into model development.
- Experience evaluating AI/ML observability or LLM evaluation vendors.
- Mission-driven, workforce, or government-adjacent data experience.
- Experience mentoring a small data and engineering team.
Culture & Benefits
- Remote-first work with an optional Toronto office arrangement.
- High-velocity, high-trust environment focused on meaningful workforce impact.
- Approximately 2–3 company trips per year, including an annual August off-site.
- Values include curiosity, outcome ownership, continuous improvement, speed, and collaboration.
- Equal opportunity workplace with reasonable accommodations available during hiring and employment.
Hiring process
- Online application and initial Talent Acquisition screen.
- Hiring manager interview followed by a performance challenge.
- Final one-to-one interviews and decision; the process generally takes about six weeks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →