Назад
Company hidden
3 часа назад

Senior Production Engineer (AI Infrastructure)

170 000 - 205 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Production Engineer (AI Infrastructure): Building and operating reliable managed AI services for large language model workloads with an accent on distributed systems, scalability, observability, and cost-efficient infrastructure. Focus on designing fault-tolerant AI platforms, improving SLI/SLO performance, and solving complex reliability challenges across training and inference clusters.

Location: On-site in San Francisco or Sunnyvale, California, United States

Salary: $170,000–$205,000 per year, plus Restricted Stock Units

Company

hirify.global builds vertically integrated energy and AI infrastructure, operating systems from power generation through cloud services to support large-scale AI workloads.

What you will do

  • Design and operate reliable managed AI services for large language model serving and scaling.
  • Build automation and reliability tooling for distributed AI pipelines and inference services.
  • Define, measure, and improve SLIs and SLOs across AI workloads.
  • Collaborate with AI, platform, and infrastructure teams to optimize large-scale training and inference clusters.
  • Automate observability and develop telemetry and performance-tuning strategies for latency-sensitive services.
  • Investigate and resolve reliability issues using telemetry, logs, and profiling, while contributing to next-generation AI-focused distributed systems.

Requirements

  • Production-grade software engineering experience beyond scripting or Bash.
  • Experience designing and implementing distributed systems.
  • Hands-on experience with large language models or AI/ML infrastructure.
  • SRE experience, including defining SLIs/SLOs, monitoring, observability, reliability improvements, fault tolerance, and automated testing.
  • Proficiency in Python, Go, Java, or C++.
  • Experience with Kubernetes or container orchestration platforms, plus strong collaboration and communication skills.

Nice to have

  • Experience scaling LLM inference or training workloads.

Culture & Benefits

  • Health, vision, and dental insurance options for employees and dependents.
  • HSA contributions, 401(k) matching up to 4% of salary, and paid parental leave.
  • Paid time off, holidays, life insurance, disability coverage, and commuter benefits.
  • Tuition reimbursement, cell phone reimbursement, Teladoc, Calm, and MetLife Legal subscriptions.
  • Restricted Stock Units in a fast-growing technology company.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →