5 часов назад
Cloud Platform Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cloud Platform Engineer (AI): Building and operating a reliable, scalable AI inferencing service across public cloud and on-premises infrastructure with an accent on uptime, latency, observability, and efficient resource utilization. Focus on designing autoscaling, CI/CD, infrastructure-as-code, and incident response systems for complex distributed services supporting inference workloads across multiple regions.
Location: San Jose, California, United States
Company
builds integrated AI hardware and software platforms, including DataScale systems, SambaFlow software, and the Suite generative AI platform.
What you will do
- Own the reliability, availability, latency, performance, and efficiency of the production AI inferencing service.
- Support AI infrastructure deployments across regions including Asia, Europe, and Latin America.
- Participate in a balanced follow-the-sun on-call rotation and lead incident response, post-mortems, and corrective actions.
- Build monitoring, alerting, dashboards, and service health measurements using tools such as Prometheus, Grafana, and Datadog.
- Design autoscaling, capacity planning, and cost-optimization strategies for variable inference workloads.
- Manage cloud and on-premises infrastructure, CI/CD pipelines, and automation using infrastructure-as-code practices.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 3–5+ years of experience in SRE, DevOps, or a related role supporting large-scale customer-facing services in AWS, GCP, or Azure.
- Strong programming or scripting skills in Python, Go, or Java.
- Experience with Docker, Kubernetes, monitoring, observability, and distributed-system troubleshooting.
- Experience with infrastructure as code, such as Terraform or CloudFormation, and familiarity with CI/CD tools and principles.
Nice to have
- Experience bridging public cloud and on-premises or data-center infrastructure.
- Production experience with ML/AI inferencing, GPU-accelerated workloads, or model-serving frameworks such as vLLM, SGLang, or Ray.
- Knowledge of MLOps, databases, caching systems, and Linux/Unix administration.
Culture & Benefits
- Work on a high-visibility AI platform with direct product and engineering impact.
- Greenfield development opportunity with autonomy in technical decision-making.
- Competitive compensation including equity and benefits.
- US-based full-time benefits include medical, dental, vision, disability, life insurance, HSA, FSA, wellness, and counseling programs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Forward Deployed Engineer, DevOps (AI)
5 часов назад
Forward Deployed Engineers (AI Infrastructure)
180 000 - 240 000$
5 дней назад
Senior DevOps / Platform Engineer (DevSecOps & Cloud Architecture)
2 дня назад
Middle/Senior DevOps Engineer (SaaS)
Inworld AI
5 дней назад
Staff / Principal Platform Engineer (AI)
23 333 - 29 167$
5 дней назад