23 часа назад
Cloud Support Engineer
145 000 - 175 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Cloud Support Engineer (Cloud/HPC): Supporting customers using sustainable GPU cloud infrastructure with an accent on Linux troubleshooting, incident management, and high-performance computing technologies. Focus on diagnosing VM and hardware issues, coordinating 24/7 incident response, and collaborating with SRE, networking, and storage teams on root-cause analysis.
Location: New York, NY, United States; on-site Tuesday through Friday and remote on Saturday
Salary: $145,000–$175,000 annually plus bonus and restricted stock units
Company
is an AI infrastructure company providing sustainable, low-cost GPU cloud computing for demanding AI workloads, scientific simulations, and computational biology.
What you will do
- Provide technical customer support through Zendesk while meeting SLAs and maintaining 95%+ CSAT.
- Participate in a 24/7 on-call rotation and lead initial incident triage, communication, and response coordination.
- Diagnose and resolve VM, hardware, scaling-test, and cloud infrastructure issues using CLI and internal tools.
- Manage alert triage, maintenance-window preparation, and node delivery testing.
- Collaborate with SRE, networking, and storage teams through root-cause analysis and RCA delivery.
- Create onboarding materials, knowledge-base documentation, and standard operating procedures.
Requirements
- Bachelor’s degree in IT, Computer Science, Engineering, or a related field, or 4+ years of equivalent technical experience.
- Strong Linux command-line and Git skills.
- 5+ years of customer support experience, preferably in cloud, storage, or networking environments.
- Experience with Kubernetes, Slurm, Terraform, Grafana, and public cloud platforms such as AWS, Azure, or GCP.
- Understanding of HPC technologies including InfiniBand, RDMA, RoCE, and software-defined networking.
- Ability to work on-site in New York Tuesday through Friday and participate in the 24/7 on-call rotation.
Nice to have
- Cloud, Kubernetes, AWS, NVIDIA, Linux Foundation, or InfiniBand certifications.
- Experience with automation tools and scripting languages.
- Experience mentoring, training, and onboarding colleagues.
- Interest in using technology to support a more sustainable future.
Culture & Benefits
- Equity packages and restricted stock units.
- Paid time off, holidays, leave programs, and parental leave.
- Health, dental, vision, life, disability, and mental health coverage.
- 401(k) plan with company matching up to 4% of salary and HSA contributions.
- Professional development, tuition reimbursement, commuter benefits, cell phone stipend, meals allowance, and volunteer time off.
- Global travel insurance, emergency assistance, and location-specific programs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
1 день назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
Lambda
6 дней назад
Software Engineer (Cloud Infrastructure)
206 000 - 275 000$
6 дней назад
IT Operations Engineer (Cloud/Network)
115 000 - 155 000$
22 часа назад
HPC Operations Engineer
175 000 - 225 000$
1 день назад
Systems Administrator
120 000 - 148 000$
1 день назад
Linux Device Management Engineer
160 000 - 200 000$