Назад
2 дня назад

Software Engineer (AI Infra and Model Foundry)

142 800 - 274 800$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI Infra and Model Foundry): Building scalable infrastructure for frontier AI models, from experimentation and training through deployment, inference, and observability, with an accent on high-throughput serving, accelerator utilization, and operational reliability. Focus on optimizing distributed systems through batching, caching, scheduling, quantization, compilation, and memory optimization while leading complex production improvements across research, product, hardware, and cloud teams.

Location: Mountain View, New York, or Redmond, United States

Salary: USD $142,800–$274,800 base pay per year across the U.S.; USD $188,000–$304,200 in the San Francisco Bay Area and New York City metropolitan area.

Company

Microsoft AI builds frontier models and systems that power Copilot and the broader Microsoft ecosystem.

What you will do

  • Set technical direction and lead multi-quarter initiatives across the model lifecycle, from experimentation and training to evaluation, deployment, and inference.
  • Develop high-throughput, low-latency model-serving systems using GPUs and other accelerators.
  • Build platform capabilities for model onboarding, versioning, packaging, rollout, routing, autoscaling, monitoring, and rollback.
  • Improve inference performance and fleet efficiency through batching, caching, workload placement, scheduling, parallelism, quantization, compilation, and memory optimization.
  • Define service-level objectives and build observability, benchmarking, capacity-management, and regression-detection systems.
  • Lead incident response, resolve distributed-systems and hardware bottlenecks, mentor engineers, and improve architecture and operational practices.

Requirements

  • Bachelor’s degree in Computer Science or a related technical field, plus 6+ years of technical engineering experience, or equivalent experience.
  • Professional coding experience with languages including C, C++, C#, Java, JavaScript, Python, Rust, or Go.
  • Experience designing, building, or operating distributed systems or high-scale production infrastructure.
  • Strong systems fundamentals in areas such as concurrency, networking, storage, resource management, reliability, or performance analysis.
  • Experience leading technically ambitious infrastructure initiatives and communicating system tradeoffs and operational risks.

Nice to have

  • Experience serving large language models or other generative models in production.
  • Experience with containers, schedulers, orchestration systems, cloud infrastructure, or machine learning platforms.
  • Experience with continuous batching, prefix or KV caching, speculative decoding, parallelism, quantization, or model compilation.
  • Experience operating highly available GPU-backed services and improving accelerator utilization or fleet economics.

Culture & Benefits

  • Collaborative work with model researchers, reinforcement learning engineers, product teams, and hardware and cloud partners.
  • Focus on reliability, safety, performance, security, privacy, and infrastructure cost.
  • Opportunities to contribute to an inclusive engineering culture through mentoring, knowledge sharing, and technical excellence.
  • Eligible roles may include benefits and additional compensation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →