SRE/Infrastructure Engineer (AI)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
TL;DR
SRE/Infrastructure Engineer (AI): Building and scaling the production infrastructure for a physical AI platform designed for field workers with an accent on observability, reliability, and ML pipeline orchestration. Focus on hardening Apache Airflow on Astronomer, optimizing Kubernetes and AWS services, and ensuring the stability of high-volume data processing.
Location: Hybrid (Lower Manhattan, NYC office - 2 days per week in person)
Salary: $160,000 - 220,000 USD
Company
empowers energy companies to modernize field work using physical AI, sensors, and advanced software to increase productivity and safety.
What you will do
- Lead the improvement and scaling of infrastructure for the data pipeline, ML training platform, and web app.
- Implement observability and reliability foundations including monitoring, alerting, and SLO/SLI definitions for Apache Airflow and Kubernetes.
- Design and own CI/CD guardrails for production changes and Astronomer deployments.
- Increase the reliability and operational visibility of machine learning inference operations.
- Develop operational tooling, runbooks, and engineering standards to automate away toil.
Requirements
- 7-10 years of experience in observability, systems/infrastructure engineering, SRE, or DevOps.
- Proficiency with Infrastructure-as-Code, specifically Terraform.
- Hands-on experience with Kubernetes or AWS ECS.
- Strong Linux debugging skills and ability to investigate production issues using logs and metrics.
- Must be based in or able to work from the New York City office (Hybrid).
Nice to have
- Experience with Apache Airflow and/or Astronomer.
- Background in early-stage startups or fast-moving environments.
- Familiarity with geospatial, imagery, lidar, or point-cloud domains.
- ML Ops skills regarding model deployment and inference reliability.
Culture & Benefits
- Collaborate with a mission-driven team of experts from top institutions (Penn, Caltech, CMU) and elite tech companies.
- Work on high-impact technology that reduces wildfire risk and accelerates storm recovery.
- Early-stage environment where ownership is highly valued and processes evolve quickly.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β