Infrastructure Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Infrastructure Engineer (AI): Securing, scaling, and operating the cloud infrastructure, databases, storage, search clusters, microservices, and machine learning pipelines powering computer vision products with an accent on reliability, security, and cost-effective delivery. Focus on designing Kubernetes and infrastructure-as-code solutions, operating high-availability ML services, improving observability, and responding to complex security and scalability challenges.
Location: Remote within the United States, with access to hubs in New York City and San Francisco; daytime US hours and weekly synchronous meetings are expected.
Salary: USD $165,000–$200,000 base
Company
is an AI company building computer vision tools, open-source resources, and hosted machine learning infrastructure used by more than one million developers.
What you will do
- Secure, scale, and maintain cloud architecture, databases, file storage, search clusters, microservices, and ML pipelines.
- Run and optimize high-availability machine learning inference services across AWS and/or GCP.
- Build infrastructure-as-code solutions with Terraform, Helm, Bash, and Python.
- Define SLOs and SLAs, participate in incident response, and improve observability and alerting.
- Identify infrastructure cost-optimization opportunities and contribute code to product features.
- Address vulnerabilities and harden systems for SOC 2, HIPAA, and GDPR readiness while participating in on-call rotations.
Requirements
- Production experience building and managing Kubernetes-based containerized applications at scale.
- Experience operating and scaling large applications, particularly ML or AI systems, in AWS and/or GCP.
- Proficiency in Python and Node, with the ability to collaborate with full-stack engineers.
- Hands-on experience with GPUs, Docker, Kubernetes, and machine learning infrastructure; familiarity with PyTorch or TensorFlow.
- Experience with CI/CD tools such as GitHub Actions or Spacelift and practical cloud security practices.
- Ability to work during daytime US hours and attend recurring synchronous meetings.
Nice to have
- Experience using or contributing to open-source devtools, infrastructure, or security projects.
- Experience working with customer security teams and supporting secure product integrations.
- Experience building computer vision projects.
Culture & Benefits
- High-autonomy environment focused on ownership, accountability, practical execution, and continuous improvement.
- Remote work supported alongside optional work from New York City or San Francisco hubs.
- $4,000 annual travel stipend for working with colleagues in person.
- $350 monthly productivity stipend for home-office or coworking expenses.
- Health insurance coverage of up to 100% for employees and their partners or families.
- Company equity and relocation support for employees joining a hub.
Hiring process
- Application review, including LinkedIn, GitHub, and relevant open-source work; a technical screen may be provided.
- 45-minute hiring manager introduction, followed by a 45-minute CTO interview and a 90-minute hands-on infrastructure interview.
- Final culture and leadership discussions, reference checks, and a background check.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →