Назад
Company hidden
3 часа назад

Software Engineer (AI Infrastructure)

175 000 - 220 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineer (AI Infrastructure) (Python/C++/Kubernetes): Architecting and building cloud infrastructure and backend services for a multi-cloud generative AI platform with an accent on distributed systems, ML workloads, reliability, and scalability. Focus on designing schedulers, resource managers, autoscalers, and model-serving layers, while optimizing compute, storage, networking, and observability across large-scale infrastructure.

Location: San Mateo or New York, United States

Salary: $175K–$220K annually, plus equity

Company

hirify.global provides a platform for building, training, and serving specialized AI models across text, image, embedding, audio, and multimodal workloads.

What you will do

  • Architect and build scalable backend infrastructure for distributed training, inference, and data-processing pipelines.
  • Design and implement job schedulers, resource managers, autoscalers, and model-serving services.
  • Lead technical design discussions, mentor engineers, and establish infrastructure engineering best practices.
  • Optimize compute costs, storage lifecycle management, network performance, system efficiency, and latency.
  • Collaborate with ML, DevOps, product, and infrastructure stakeholders to translate research and product needs into robust solutions.
  • Own systems from design through deployment and observability, ensuring high availability, fault tolerance, disaster recovery, and operational excellence.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 5+ years of experience designing and building backend infrastructure in cloud environments such as AWS, GCP, or Azure.
  • Experience with ML infrastructure and tooling, including technologies such as PyTorch, TensorFlow, Vertex AI, SageMaker, or Kubernetes.
  • Strong software development skills in Python or C++.
  • Deep understanding of distributed systems fundamentals, including scheduling, orchestration, storage, networking, and compute optimization.
  • Experience with monitoring, alerting, logging, and tracing for system observability.

Nice to have

  • Master’s or PhD in Computer Science or a related field.
  • Experience leading infrastructure projects for large-scale ML/AI workloads or high-throughput systems.
  • Familiarity with Terraform, ArgoCD, GitOps, infrastructure-as-code, and CI/CD tooling.
  • Contributions to open-source cloud or ML infrastructure projects.

Culture & Benefits

  • Work on challenging AI infrastructure problems, including low-latency inference and scalable model serving.
  • Build production technology using advanced cloud-native and open-source tools such as Kubernetes, Kubeflow, and MLFlow.
  • Collaborate with experienced engineers and AI researchers in a fast-growing environment.
  • Compensation includes equity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →