Назад
Company hidden
1 час назад

Member of Technical Staff, Backend, LLM Applications

200 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff, Backend, LLM Applications (AI/LLM Infrastructure): Building and operating scalable backend services and model-serving infrastructure for diffusion LLMs with an accent on latency, throughput, cost, and reliability. Focus on load balancing, autoscaling, model versioning, canary deployments, observability, and zero-downtime rollouts for billions of inference requests.

Location: Bay Area, United States; in-office

Salary: $200,000–$350,000 annual base salary, plus equity and benefits.

Company

hirify.global develops diffusion-based large language models designed to deliver faster and more efficient inference at high quality.

What you will do

  • Design, build, and operate scalable backend services and model-serving infrastructure for diffusion LLMs.
  • Implement load balancing, autoscaling, and traffic routing for model endpoints.
  • Build model versioning, canary deployment, and zero-downtime rollout systems.
  • Develop monitoring, alerting, and observability tooling for SLA compliance and incident response.
  • Benchmark serving frameworks and hardware configurations to guide infrastructure decisions.

Requirements

  • BS, MS, PhD in Computer Science or a related field, or equivalent experience.
  • 5+ years of experience building production backend systems.
  • Strong Python skills, including asynchronous programming and concurrent systems.
  • Solid understanding of distributed systems, networking, and load balancing at scale.
  • Experience with Kubernetes, CI/CD pipelines, and AWS and/or Azure cloud infrastructure.

Nice to have

  • Experience serving LLMs or other large generative models in production at scale.
  • GPU instance management and cloud cost optimization experience.
  • Experience with Terraform, deployment automation, Prometheus, or Grafana.
  • Familiarity with vLLM, Triton Inference Server, or TensorRT-LLM.

Culture & Benefits

  • Collaboration with AI researchers and engineers who developed diffusion models and related technologies.
  • Opportunity to influence foundational AI infrastructure and product direction.
  • Competitive compensation with equity at a rapidly growing startup.
  • Flexible vacation, paid time off, health, dental, and vision insurance.
  • 401(k) match, catered meals, and commuter subsidies.
  • Collaborative and inclusive work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →