Назад
4 дня назад

Software Engineer, Production Inference (AI)

300 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer, Production Inference (AI): Operating and scaling production inference systems for live, multi-tenant model serving with an accent on reliability, rollout safety, observability, and cost-efficient infrastructure. Focus on building resilient serving systems, productionizing inference techniques, leading incident response, and managing capacity as traffic and model sizes grow.

Location: On-site in San Francisco, CA; San Francisco and New York listed as locations

Annual salary: $300,000–$400,000 USD, depending on background, skills, and experience.

Company

Thinking Machines Lab builds AI systems designed to extend human will and judgment, including frontier models, model customization infrastructure, and human-AI interfaces.

What you will do

  • Operate and scale production inference systems serving live traffic, including a multi-tenant serving platform.
  • Own safe, incremental rollouts for new models, model versions, and inference optimizations.
  • Build observability, alerting, and capacity-planning systems for rapid detection and resolution of production issues.
  • Partner with inference and research teams to productionize new serving techniques.
  • Lead incident response, root-cause analysis, and durable remediation for production inference issues.
  • Design graceful degradation, failover, redundancy, and cost-efficient capacity management as usage grows.

Requirements

  • Experience operating large-scale, latency-sensitive production systems.
  • Proficiency in Python and Go or another systems language.
  • Experience with observability, monitoring, and incident response for production services.
  • Strong understanding of distributed systems and failure modes at scale.
  • Comfort with on-call work and leading incident response for critical production systems.
  • Ability to work on-site in the United States, with the role based in San Francisco, CA.

Nice to have

  • Experience running production inference for large language models or other large-scale machine learning systems.
  • Experience with canarying, blue/green deployments, feature flags, or similar rollout systems.
  • Experience with GPU or TPU capacity planning and cost optimization.
  • Familiarity with batching, caching, quantization, and other inference-specific techniques.
  • Experience working autonomously in a fast-changing, early-stage environment.

Culture & Benefits

  • Generous health, dental, and vision benefits.
  • Unlimited paid time off and paid parental leave.
  • Visa sponsorship is available, with support throughout the visa process.
  • Relocation support is available as needed.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →