3 часа назад
Software Engineer (AI Infrastructure)
175 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer (AI Infrastructure) (Python/C++/Kubernetes): Architecting and building cloud infrastructure and backend services for a multi-cloud generative AI platform with an accent on distributed systems, ML workloads, reliability, and scalability. Focus on designing schedulers, resource managers, autoscalers, and model-serving layers, while optimizing compute, storage, networking, and observability across large-scale infrastructure.
Location: San Mateo or New York, United States
Salary: $175K–$220K annually, plus equity
Company
provides a platform for building, training, and serving specialized AI models across text, image, embedding, audio, and multimodal workloads.
What you will do
- Architect and build scalable backend infrastructure for distributed training, inference, and data-processing pipelines.
- Design and implement job schedulers, resource managers, autoscalers, and model-serving services.
- Lead technical design discussions, mentor engineers, and establish infrastructure engineering best practices.
- Optimize compute costs, storage lifecycle management, network performance, system efficiency, and latency.
- Collaborate with ML, DevOps, product, and infrastructure stakeholders to translate research and product needs into robust solutions.
- Own systems from design through deployment and observability, ensuring high availability, fault tolerance, disaster recovery, and operational excellence.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 5+ years of experience designing and building backend infrastructure in cloud environments such as AWS, GCP, or Azure.
- Experience with ML infrastructure and tooling, including technologies such as PyTorch, TensorFlow, Vertex AI, SageMaker, or Kubernetes.
- Strong software development skills in Python or C++.
- Deep understanding of distributed systems fundamentals, including scheduling, orchestration, storage, networking, and compute optimization.
- Experience with monitoring, alerting, logging, and tracing for system observability.
Nice to have
- Master’s or PhD in Computer Science or a related field.
- Experience leading infrastructure projects for large-scale ML/AI workloads or high-throughput systems.
- Familiarity with Terraform, ArgoCD, GitOps, infrastructure-as-code, and CI/CD tooling.
- Contributions to open-source cloud or ML infrastructure projects.
Culture & Benefits
- Work on challenging AI infrastructure problems, including low-latency inference and scalable model serving.
- Build production technology using advanced cloud-native and open-source tools such as Kubernetes, Kubeflow, and MLFlow.
- Collaborate with experienced engineers and AI researchers in a fast-growing environment.
- Compensation includes equity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Snowflake
5 дней назад
Senior Software Engineer (Natsec)
200 000 - 287 500$
5 дней назад
Senior Backend Engineer (AI)
180 000 - 250 000$
5 дней назад
Senior Software Engineer, Training & Experimentation (AI)
180 000 - 250 000$
6 часов назад
Staff Software Engineer (Backend/AWS)
175 000 - 210 000$
4 часа назад
Backend Senior and Staff Engineers (AI)
100 000 - 300 000$
6 дней назад
Senior Software Engineer, Backend (AI)
190 000 - 230 000$