Назад
14 дней назад

Member of Technical Staff, Inference Systems Research (AI)

188 000 - 304 200$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Inference Systems Research (AI): Building infrastructure, tools, and techniques for running and optimizing frontier generative AI models across reinforcement learning, synthetic data generation, evaluations, and production with an accent on inference performance, GPU kernels, and distributed systems. Focus on debugging performance bottlenecks and numerical instabilities, profiling PyTorch models, and designing scalable systems with frameworks such as vLLM and SGLang.

Location: Mountain View, United States. Employees living within 50 miles of the designated U.S. office are expected to work from the office at least four days per week.

Salary: USD $188,000–$304,200 per year for the San Francisco Bay Area and New York City metropolitan area; a U.S.-wide range of USD $142,800–$274,800 also applies to other locations.

Company

Microsoft AI is a startup-like organization within Microsoft focused on advancing safe and responsible artificial intelligence and frontier models.

What you will do

  • Work with researchers and engineers to implement frontier AI research ideas.
  • Develop systems, tools, and techniques that improve model inference performance.
  • Build debugging tools for performance bottlenecks, numerical instabilities, and distributed systems issues.
  • Create tools and processes that improve collective engineering productivity.
  • Remove technical roadblocks and deliver improvements to users quickly and iteratively.

Requirements

  • Bachelor's degree in Computer Science or a related technical field and 6+ years of technical engineering experience, or equivalent experience.
  • Experience with generative AI, distributed computing, and large-scale production inference.
  • Expertise in Python and its ecosystem, including tools such as uv, pybind/nanobind, and FastAPI.
  • Experience with GPU kernel programming, benchmarking, profiling, and optimizing PyTorch generative AI models.
  • Familiarity with open-source inference frameworks such as vLLM and SGLang, as well as the JAX scaling book.
  • Working knowledge of programming languages including C, C++, C#, Java, JavaScript, or Python.

Nice to have

  • Master's degree in Computer Science or a related technical field and 8+ years of technical engineering experience, or a bachelor's degree and 12+ years of experience.
  • Experience working in a fast-paced, design-driven product development cycle.

Culture & Benefits

  • Collaboration with researchers in Microsoft AI's research organization.
  • Work across kernels, inference algorithms, model architecture co-design, ASIC co-design, distributed systems, and profiling tools.
  • Focus on respect, integrity, accountability, inclusion, and a growth mindset.
  • Potential eligibility for benefits and additional compensation.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →