DevOps Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
DevOps Engineer (AI): Scaling and hardening the infrastructure for a physical AI platform designed for energy field workers with an accent on data pipelines and ML training platforms. Focus on building observability, reliability foundations, and CI/CD guardrails for high-volume processing across AWS and Kubernetes.
Location: Hybrid in New York City (2 days per week in person)
Salary: $160,000 – $220,000 USD
Company
develops physical AI and robotics software to modernize field work for energy companies, increasing productivity for engineers and crews.
What you will do
- Lead the improvement and scaling of infrastructure for data pipelines, ML training platforms, and web applications.
- Build observability and reliability foundations, including monitoring, alerting, and SLO/SLI definitions for Apache Airflow on Astronomer.
- Design and implement CI/CD guardrails for production changes and safe rollout mechanics for Astronomer deployments.
- Improve the reliability and operational visibility of machine learning inference operations and output correctness.
- Create operational tooling, runbooks, and engineering standards to automate toil and improve debugging at scale.
- Collaborate with data platform and engineering teams to manage complex data movement across AWS (S3, SQS, Lambda) and Kubernetes.
Requirements
- 7-10 years of experience in observability, systems/infrastructure engineering, SRE, or DevOps.
- Must be based in the New York City area to work from the Lower Manhattan office (hybrid).
- Strong hands-on experience with Infrastructure-as-Code (Terraform).
- Proficiency in container orchestration and debugging (Kubernetes and/or ECS).
- Deep Linux debugging skills and ability to investigate production issues using logs and metrics.
- Ability to reason about end-to-end architecture with product impact and failure handling in mind.
Nice to have
- Experience with Apache Airflow and/or Astronomer.
- Experience with AWS and DuploCloud.
- Background in geospatial, imagery, LiDAR, or point-cloud domains.
- ML Ops skills related to model deployment, inference reliability, and CI/CD for model artifacts.
- Experience working in early-stage, fast-moving startup environments.
Culture & Benefits
- Opportunity to be the first full-time DevOps Engineer in a fast-growing, mission-driven startup.
- Collaborate with experts from top institutions (Penn, Caltech, CMU) and leading tech companies (Palantir, Stripe).
- Work on high-impact technology reducing wildfire risk and accelerating storm recovery for major US utilities.
- Hybrid work model in midtown Manhattan.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →