Назад
Company hidden
обновлено 4 минуты назад

Senior Machine Learning Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Machine Learning Engineer (AI): Develops and productionizes machine learning models for Cloudflare’s serverless inference platform with an accent on inference optimization, benchmarking, model evaluation, and reliable deployment across heterogeneous accelerators. Focus on improving latency, throughput, cost efficiency, and model quality while building scalable tooling for AI applications at Internet scale.

Location: Hybrid; available location is Austin, Texas, United States

Company

hirify.global operates a large global network that protects and accelerates Internet applications and provides infrastructure for customers ranging from individual websites to Fortune 500 companies.

What you will do

  • Develop, optimize, and productionize machine learning models for the serverless inference platform and Workers AI.
  • Build benchmarking and evaluation frameworks for latency, throughput, cost efficiency, and model behavior across language, speech, vision, and other model families.
  • Improve inference through quantization, batching, caching, model compilation, runtime tuning, and accelerator-aware optimization.
  • Integrate models into distributed inference infrastructure across heterogeneous GPUs and next-generation accelerators.
  • Improve deployment workflows through validation, safe rollouts, observability, regression testing, and operational readiness.
  • Mentor engineers, contribute to technical direction, and collaborate with systems, product, hardware, and AI/ML teams.

Requirements

  • Experience building, optimizing, and operating machine learning models in production.
  • Strong proficiency with Python and modern machine learning frameworks such as PyTorch, TensorFlow, or JAX.
  • Hands-on experience with inference optimization for large-scale models, including quantization, batching, caching, compilation, and serving runtime tuning.
  • Experience with inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, or llama.cpp.
  • Understanding of production ML concerns, including evaluation, monitoring, regressions, rollout safety, and reliability.
  • Experience working across ML and systems boundaries, with familiarity with distributed systems, networking, or serverless platforms.

Nice to have

  • Experience contributing to open-source ML tooling, model serving frameworks, or inference runtimes.

Culture & Benefits

  • Work is organized in a hybrid format.
  • Applicants reaching the offer stage may be asked to attend an in-person interview at a hirify.global office or hub.
  • The role may involve access to information controlled under U.S. export control laws; employment may depend on authorization to receive the technology without export-license sponsorship.
  • hirify.global supports initiatives protecting journalism, civil society, election information, and the open Internet.

Hiring process

  • Applicants progressing to the offer stage may be required to attend an in-person interview at a hirify.global office or hub.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →