Назад
Company hidden
3 дня назад

AI Infrastructure Engineer (AI)

168 000 - 205 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Infrastructure Engineer (AI): Building and optimizing infrastructure for real-time computer vision, LLM, LVM, and multimodal inference across large-scale video and sensor data with an accent on model serving, evaluation, GPU utilization, and production reliability. Focus on designing scalable inference systems, benchmarking model quality, optimizing latency and cost, and creating continuous feedback loops for AI improvement.

Location: Hybrid in Redwood City, United States; engineering teams work from the office three days per week.

Salary: $168,000–$205,000 per year, plus equity.

Company

hirify.global develops AI-powered physical security solutions using computer vision, vision-language models, large language models, and multimodal systems.

What you will do

  • Design, build, and maintain infrastructure for real-time computer vision, LLM, LVM, and multimodal inference.
  • Build scalable systems for processing large volumes of video and sensor data.
  • Optimize inference latency, throughput, GPU utilization, reliability, and cost.
  • Develop evaluation harnesses, benchmarks, regression tests, and continuous model-improvement systems.
  • Improve model-serving architecture through batching, caching, routing, quantization, parallelism, and hardware optimization.
  • Build data engines, observability tools, monitoring, and feedback loops for production AI systems.

Requirements

  • 2+ years of industry experience in infrastructure engineering, distributed systems, machine learning platforms, or production AI systems.
  • Strong Python programming and software engineering fundamentals.
  • Experience building scalable machine learning infrastructure for training, inference, evaluation, and deployment.
  • Hands-on experience running deep learning models in production, including LLMs, LVMs, vision-language models, or multimodal models.
  • Experience with model-serving systems such as vLLM, Triton Inference Server, or similar technologies.
  • Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU-based workloads.

Nice to have

  • Experience with CUDA, NCCL, PyTorch, TensorRT, ONNX, or similar machine learning systems technologies.
  • Experience with video understanding, real-time computer vision, multimodal AI, or physical-world AI systems.
  • Experience with model compression, speculative decoding, distillation, pruning, or low-latency serving.
  • Experience with RAG, vector databases, embedding models, re-rankers, or search infrastructure.
  • Experience building internal ML platforms for researchers and applied ML teams.

Culture & Benefits

  • Full-time employees receive stock options.
  • Medical, dental, vision, life, EAP, legal services, and 401(k) benefits.
  • Flexible time off and a winter break between Christmas and New Year’s for most roles.
  • Regular opportunities for in-person collaboration and connection with colleagues.
  • Work includes the latest technology and equipment delivered to employees.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →