Назад
3 дня назад

Performance Engineer, Inference Engine (AI)

350 000 - 850 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Performance Engineer, Inference Engine (AI) (LLM inference and high-performance systems): Building and optimizing Anthropic’s inference engine for Claude across accelerator and cloud platforms with an accent on throughput, cost, reliability, latency, and model-state management. Focus on coordinating host-device systems, profiling compute and memory bottlenecks, and designing scalable distributed infrastructure that preserves model quality and safety.

Location: San Francisco, CA or New York City, NY; hybrid work with staff expected to be in an office at least 25% of the time

Annual salary: $350,000–$850,000 USD

Company

Anthropic builds reliable, interpretable, and steerable AI systems designed to be safe and beneficial for users and society.

What you will do

  • Build and optimize the in-house inference engine that manages batching, model placement across chips, memory for weights and activations, forward passes, and model state across requests.
  • Improve throughput, cost, reliability, and latency across accelerator and cloud platforms serving Claude and research workloads.
  • Analyze hardware and interconnect constraints across FLOPs, HBM, PCIe, RDMA, and network links.
  • Develop observability, profile performance, model improvement impact, deploy changes, and measure results iteratively.
  • Maintain high device utilization through caching, scheduling, and efficient coordination between host and accelerator.
  • Collaborate with safeguards and safety teams to preserve model quality and robustness during inference.

Requirements

  • Working mental model of LLM inference, including prefill and decode across accelerator compute, memory, interconnect, and host operations.
  • Strong systems programming skills in Rust, C++, or a similar language, with attention to code quality and testing.
  • Analytical performance methodology: observe and profile, form hypotheses, test changes, and measure results.
  • Ability to learn unfamiliar deep systems quickly and deliver consequential changes.
  • Bachelor’s degree or equivalent education, training, or professional experience in a relevant field.
  • Collaborative communication, willingness to pair program, and care for the societal impact of AI systems.

Nice to have

  • Experience with an LLM serving engine, GPU or accelerator programming, OS internals, or transformer language modeling.
  • Experience building allocators, caches, schedulers, or high-bandwidth transports.
  • Fluency in Rust and experience with determinism, replay, or property-based testing.

Culture & Benefits

  • Collaborative environment with frequent research discussions and pair programming.
  • Work on large-scale AI research efforts focused on trustworthy and steerable AI.
  • Competitive compensation with optional equity donation matching.
  • Generous vacation and parental leave, flexible working hours, and office collaboration spaces.
  • Visa sponsorship is available, subject to role and candidate eligibility.

Hiring process

  • Candidates may use AI during the application process according to Anthropic’s candidate AI usage policy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →