Назад
Company hidden
50 минут назад

Staff Engineer (AI Inference)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Engineer (AI Inference) (Distributed Cloud Systems): Building and operating the Inference Cloud Platform for globally distributed AI inference with an accent on multi-region traffic architecture, availability, latency, and reliability. Focus on designing active-active systems, optimizing high-QPS workloads, implementing graceful degradation and traffic control, and solving complex production bottlenecks.

Location: On-site at the Sunnyvale headquarters office, United States

Company

hirify.global builds large-scale AI hardware and software platforms for high-speed model training and inference.

What you will do

  • Shape the architecture and roadmap of major areas within the Inference Cloud Platform.
  • Design and build service discovery, request routing, load balancing, caching, batching, and traffic-management components.
  • Architect active-active, multi-region systems with rapid failover, graceful degradation, clear SLOs, and high resilience.
  • Develop admission control, quota management, rate limiting, and differentiated quality-of-service mechanisms.
  • Write and review production code, lead architectural and design reviews, and make high-consequence technical decisions.
  • Lead incident response, observability, capacity planning, post-incident improvements, and cross-functional technical alignment.

Requirements

  • 8+ years of software engineering experience, including substantial individual-contributor work on large-scale distributed systems or cloud infrastructure.
  • Deep expertise in distributed systems architecture, networking, compute orchestration, container platforms, and multi-region production services.
  • Experience designing highly available, latency-sensitive systems and improving latency, throughput, and capacity efficiency in high-QPS environments.
  • Strong proficiency in Go, C++, or Python and the ability to contribute production code directly.
  • Experience with metrics, logging, tracing, alerting, incident response, and SLO-driven reliability practices.
  • Ability to influence senior engineers and cross-functional partners through technical communication and judgment.

Nice to have

  • Experience with ML inference infrastructure, model serving systems, or GPU-accelerated workloads.
  • Experience with TTFT optimization and tail-latency reduction.

Culture & Benefits

  • Work on an AI platform designed to overcome GPU limitations.
  • Opportunities to publish and open-source AI research.
  • Work with a high-performance AI supercomputer platform.
  • Startup vitality combined with job stability.
  • Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →