Назад
Company hidden
2 дня назад

Reliability Engineer (AI Infrastructure)

240 000 - 290 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Reliability Engineer (AI Infrastructure): Building reliability standards, observability tooling, and automation for high-throughput AI inference and training infrastructure with an accent on distributed systems, cloud operations, and production telemetry. Focus on defining SLOs, managing incidents, testing failures, and solving reliability issues across GPU scheduling, networking, storage, and customer-facing services.

Location: Hybrid in San Mateo or New York, United States

Salary: $240,000–$290,000 annually plus equity

Company

hirify.global provides an AI platform for building, training, and serving specialized models across text, image, embedding, audio, and multimodal workloads.

What you will do

  • Define SLOs, error budgets, production-readiness criteria, and on-call expectations.
  • Own logging, telemetry, alerting, failure-injection, load-testing, and self-healing automation tooling.
  • Identify reliability issues across service boundaries, dependencies, retries, timeouts, and multi-region infrastructure.
  • Coordinate production incidents, lead blameless postmortems, and track corrective actions.
  • Automate repetitive operational work and reduce on-call toil.
  • Partner with cloud infrastructure, inference, training, performance, product, and control-plane teams.

Requirements

  • 5+ years working with Linux internals, system performance troubleshooting, and TCP/IP, HTTP, and gRPC networking.
  • 5+ years writing production-grade tools or systems code in Python, Go, C++, or Rust.
  • Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
  • Experience with distributed systems, microservices, high-throughput control planes, or multi-region deployments.
  • Knowledge of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent practical experience.

Nice to have

  • Experience with Prometheus, Grafana, OpenTelemetry, and actionable alerting.
  • Exposure to GPUs, inference serving, or distributed training.
  • Experience building AI-assisted operations tools, agents, or LLM-based investigation and automation tooling.
  • Open-source contributions to infrastructure, systems, or ML-serving projects.
  • Startup experience and comfort working pragmatically across teams.

Culture & Benefits

  • Work on challenging AI infrastructure problems, including low-latency inference and scalable model serving.
  • High ownership and direct impact in a fast-growing product company.
  • Collaboration with experienced engineers and AI researchers.
  • Equity included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →