Назад
Company hidden
23 часа назад

Lead MLOps Engineer (Databricks)

100 000 - 200 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead MLOps Engineer (Databricks) (Databricks/MLflow/Agentic AI): Building and operating production machine-learning workloads and reusable MLOps patterns on Databricks with an accent on medallion data architecture, governance, observability, and FinOps. Focus on optimizing Spark and cluster performance, managing model lifecycles and drift detection, and integrating agentic AI into validation, testing, monitoring, and documentation.

Location: Remote for qualified candidates living outside a 50-mile radius of Bloomington, Illinois; Dunwoody, Georgia; Richardson, Texas; or Tempe, Arizona. Hybrid for qualified candidates within a 50-mile radius of these hub locations. Candidates must be eligible to work in the United States immediately; visa sponsorship is not available.

Salary: $100,000–$200,000 per year, plus potential annual incentive pay of up to 15% of base salary.

Company

hirify.global is an insurance and financial services organization focused on helping customers, investing in communities, and improving everyday life.

What you will do

  • Own the performance, quality, production readiness, governance, lineage, and cost optimization of innovation projects running on Databricks.
  • Design and operate Bronze–Silver–Gold medallion architectures with data quality controls, schema management, and Delta Lake tuning.
  • Optimize Spark queries, partitioning, caching, and Databricks cluster sizing to improve application performance.
  • Manage model lifecycles with MLflow, including experimentation, deployment, retraining, retirement, and drift detection.
  • Define production handoff standards covering runbooks, self-healing, incident targets, governance, FinOps, and observability.
  • Set technical direction, mentor engineers, and integrate agentic AI into validation, pipeline generation, testing, monitoring, and documentation.

Requirements

  • Immediate U.S. work authorization is required; sponsorship is not available.
  • At least one year of production experience running Databricks at the application layer, including Delta Lake, Unity Catalog, PySpark, and medallion architecture.
  • Experience operating production machine-learning systems, handling incidents, writing root-cause analyses, and detecting model drift or degradation.
  • Hands-on experience with MLflow 3.0, model registries, experiment tracking, deployment automation, retraining pipelines, and drift detection.
  • Production-grade Python, PySpark, SQL, AWS, CI/CD for data and ML pipelines, and testing strategies for non-deterministic systems.
  • Experience with FinOps, data lineage, observability, automated incident ticketing, production handoffs, runbooks, and self-healing patterns.

Nice to have

  • Experience with Omnigent, Databricks’ open-source agent meta-harness.

Culture & Benefits

  • Lead individual-contributor role with technical leadership responsibilities and no people management.
  • Work on high-visibility innovation initiatives involving Databricks and agentic AI.
  • Health, dental, vision, telemedicine, mental health, and wellbeing programs.
  • Training, tuition assistance, mentoring, and employee resource groups.
  • Paid time off, parental leave, paid holidays, life leave, bereavement leave, and community service or education support days.
  • 401(k) plan with company contributions of up to 7% of salary, financial coaching, and other financial benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →