Назад
Company hidden
23 часа назад

Staff Engineer - AI/Inference /Gateway (AI)

171 000 - 257 000CAD
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Serbia +7 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Engineer - AI/Inference /Gateway (AI): Building scalable, fault-tolerant Kubernetes services and high-performance inference infrastructure for enterprise Generative AI and LLM workloads with an accent on distributed systems, multi-tenant platforms, observability, and cloud-native deployment. Focus on designing LLM serving capabilities, optimizing low-latency and high-throughput services, and solving complex production reliability challenges across on-premises, hybrid, and public-cloud environments.

Location: Vancouver, Canada; hybrid with at least 3 days onsite per week

Salary: CAD $171,000–$257,000 annually

Company

hirify.global develops the hirify.global Cloud Platform for AI, enabling organizations to build, fine-tune, and run Generative AI, LLM, and Agentic AI applications across on-premises data centers, edge locations, and public clouds.

What you will do

  • Architect and develop horizontally scalable, containerized, fault-tolerant services on Kubernetes for enterprise AI and LLM workloads.
  • Build and operate low-latency, high-throughput inference and platform services for Generative AI and Agentic AI applications.
  • Design distributed systems across compute, storage, networking, virtualization, and other low-level infrastructure layers.
  • Develop multi-tenant services for on-premises, hybrid, and cloud-based AI deployments.
  • Implement observability, CI/CD, deployment automation, reliability improvements, and production troubleshooting.
  • Collaborate with product management, AI, and software engineering teams while contributing across architecture, development, testing, deployment, and operations.

Requirements

  • 8+ years of experience developing maintainable, resilient software products in a product development organization.
  • Strong fundamentals in data structures, algorithms, operating systems, networking, distributed systems, and datacenter architecture.
  • Hands-on experience with Docker, Kubernetes, cloud-native architectures, and multi-tenant services.
  • Production backend development experience with Go, Python, C++, or Rust, including high-performance and performance-sensitive systems.
  • Experience with CI/CD, release automation, distributed data stores, observability platforms, and deployments across on-premises, cloud, and hybrid environments.
  • Knowledge of LLM serving concepts including request routing, rate limiting, token streaming, scheduling, load balancing, quota management, and usage budgeting; a master’s degree in Computer Science or equivalent practical experience.

Nice to have

  • Experience with PyTorch, TensorFlow, GPU acceleration, vLLM, DeepSpeed, Hugging Face TGI, or Triton.
  • Experience with Retrieval-Augmented Generation, vector databases, AI orchestration frameworks, production LLM APIs, or agentic systems.
  • Open-source contributions or experience working in large distributed codebases.

Culture & Benefits

  • Hybrid work arrangement with at least 3 onsite days per week in Vancouver and other applicable workplace locations.
  • RRSP with dollar-for-dollar matching of up to 7% of base salary.
  • Mental health coverage and paramedical benefits.
  • Fully paid maternity and parental leave, plus bereavement leave.
  • RSUs and an Employee Stock Purchase Plan with a 15% discount.
  • Collaboration with a globally distributed enterprise AI team.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →