Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Backline Engineer (ML/AI) (Databricks Data and AI Platform): Troubleshooting complex distributed ML/AI issues across model training, inference, Model Serving, MLflow, Feature Engineering, Spark, and Delta Lake with an accent on root-cause analysis, performance optimization, and production reliability. Focus on reproducing customer workloads, diagnosing CPU/GPU, memory, concurrency, and deployment failures, and driving resolutions with Engineering and Product teams.
Location: Bellevue, Washington, or San Francisco, California, United States
Salary: $170.40–$255.60 USD
Company
Databricks builds a data and AI infrastructure platform used to develop and scale data, analytics, machine learning, and AI applications.
What you will do
- Serve as a senior escalation point for complex ML/AI issues involving model training, inference, Model Serving, MLflow, Feature Engineering, Spark, and Delta Lake.
- Investigate logs, traces, metrics, profiling data, configurations, source code, and customer workloads to identify root causes.
- Reproduce customer issues through Python and Spark development, workload construction, configuration changes, and performance analysis.
- Troubleshoot training and inference failures, resource utilization, memory and CPU/GPU issues, distributed execution problems, and deployment failures.
- Partner with Engineering and Product to resolve difficult issues, improve products, and develop diagnostics, tooling, automation, and documentation.
- Mentor engineers and strengthen the technical troubleshooting capabilities of the Support organization.
Requirements
- 10+ years of relevant experience with deep expertise in at least one specialized track: Data Engineering, Product Supportability, or AI.
- Deep troubleshooting experience with distributed ML/AI systems across application code, frameworks, infrastructure, and the Databricks platform.
- Strong hands-on Python experience and experience with ML frameworks such as PyTorch, TensorFlow, or Scikit-Learn.
- Strong knowledge of Databricks ML/AI technologies, Apache Spark, distributed computing, memory management, performance analysis, and model lifecycle management.
- Experience with ML deployment infrastructure, including Kubernetes, cloud ML platforms, CI/CD, model monitoring, and production ML systems.
- Strong technical communication skills and the ability to manage customers, technical stakeholders, and ambiguous high-impact problems.
Nice to have
- Expertise in large-scale data engineering and ETL pipelines using Spark, Delta Lake, or Hive.
- Code-level root-cause analysis and profiling using Java, Scala, or Python, including metrics and heap/thread dumps.
- Experience with large-scale machine learning, generative AI, LLM applications, agent-driven workflows, and distributed ML optimization.
Culture & Benefits
- Benefits and perks are offered according to the employee's region.
- Work focuses on customer success, technical excellence, automation, and operational improvement.
- Collaborate across Support, Engineering, Product, and customer organizations.
- Inclusive hiring practices and a commitment to diversity and equal employment opportunity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Associate Field Engineer (AI)
90 000 - 105 000$
6 дней назад
Platform Support Engineer (AI)
CoreWeave
4 дня назад
Technical Support Engineer II (AI/ML)
99 000 - 145 000$
Replit
6 дней назад
Premium Support Engineer (AI)
185 000 - 210 000$
Replit
6 дней назад
Premium Support Engineer (AI)
185 000 - 210 000$
Vapi
22 часа назад
Senior Technical Support Engineer (AI)
154 000 - 169 000$