Назад
Company hidden
3 часа назад

Software Engineer (AI)

230 000 - 390 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI): Building and operating inference systems for self-hosted models and third-party providers with an accent on distributed infrastructure, low latency, reliability, and cost efficiency. Focus on designing routing and serving architecture, managing GPU capacity and quotas, and optimizing large-scale production inference.

Location: San Francisco, CA, United States; work arrangement: on-site

Salary: $230,000–$390,000 annually, plus equity

Company

hirify.global develops customer-facing AI agents for major brands and operates primarily in person from San Francisco, with offices across North America, Europe, and Asia.

What you will do

  • Design hirify.global’s inference architecture across self-hosted models and third-party inference providers.
  • Build routing, failover, capacity-management, quota, serving, and proxy systems for large-scale production workloads.
  • Operate self-hosted inference on GPU infrastructure, including containers, inference engines, and compute capacity.
  • Optimize latency, throughput, reliability, and cost using techniques such as speculative decoding and serving-engine improvements.
  • Partner with frontier labs, inference providers, Applied Research, Models, and Agent Runtime teams.
  • Contribute to infrastructure supporting post-training and the broader model lifecycle.

Requirements

  • Strong systems thinking and distributed-systems fundamentals.
  • Experience designing, building, and operating large-scale production systems.
  • Strong judgment regarding tradeoffs between latency, reliability, capacity, and cost.
  • Experience owning complex infrastructure from architecture through production operation.
  • Interest in applying systems expertise to AI infrastructure and learning rapidly as the technology evolves.
  • On-site work in San Francisco, CA is required.

Nice to have

  • Experience with ML infrastructure, MLOps, or production inference systems.
  • Experience serving LLMs or other large models at scale.
  • Experience operating self-hosted inference and GPU infrastructure.
  • Familiarity with inference frameworks such as vLLM or SGLang.
  • Experience with post-training infrastructure or inference-performance optimization.

Culture & Benefits

  • Values include trust, customer obsession, craftsmanship, intensity, and family.
  • Unlimited paid time off, medical, dental, and vision benefits for employees and families.
  • Life insurance, disability benefits, parental leave, and fertility and family-building support.
  • Retirement benefits depend on the country of employment.
  • Lunch, snacks, coffee, a discretionary benefit stipend, and free alphorn lessons.
  • Eligible full-time employees may participate in equity plans.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →