Senior Devops Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior DevOps Engineer (AI/Kubernetes): Building and scaling a high-impact weather platform using cloud-native services and HPC clusters with an accent on AI-driven operational efficiency and MLOps. Focus on designing self-service platforms, integrating GPU-based model deployments, and optimizing for security, cost, and SLOs.
Location: Washington, District of Columbia, United States. Restricted to U.S. citizens, permanent residents, and protected individuals due to U.S. export control laws.
Salary: $160,000 – $180,000
Company
is the world's leading Resilience Platform, combining space technology, advanced generative AI, and proprietary weather modeling to empower organizations to manage weather-related risks.
What you will do
- Develop and adopt AI-powered tools to optimize Development and Operations processes.
- Collaborate with weather scientists and engineers to optimize service performance, reliability, scale, security, and cost.
- Build self-service platforms and maintain adaptive cloud infrastructure to support business growth.
- Support scientific computing workloads on HPC clusters (SLURM) alongside Kubernetes platforms.
- Integrate MLOps practices for GPU-based model deployment on Kubernetes.
- Maintain production availability by participating in on-call shifts.
Requirements
- 6+ years of experience as a Platform/DevOps/SRE Engineer in containerized cloud environments (AWS, GCP, or Azure).
- Proficiency with Infrastructure as Code (IaC) using Terraform or Crossplane.
- Experience with CI/CD tools and deployment methodologies in Kubernetes.
- Experience implementing monitoring systems such as Datadog, Prometheus, Grafana, or ELK Stack.
- Proficiency with scripting languages including Python, Node.js, and Go.
- Must be a U.S. citizen, permanent resident, or protected individual as required by U.S. export control laws.
Nice to have
- Experience with HPC/scientific computing (e.g., Slurm, AWS ParallelCluster, Azure CycleCloud).
- Familiarity with parallel filesystems like Lustre or NFS.
- Experience building agentic DevOps workflows.
Culture & Benefits
- Comprehensive health benefits and unlimited paid time off.
- Flexible hours and a culture that values impact over hours worked.
- Opportunity to work with a small global team using cutting-edge space and AI technology.
- Defined career growth paths emphasizing expertise and professional development.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →