Staff Engineer (High Performance Computing)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff Engineer (High Performance Computing): Designing and operating cloud-native HPC infrastructure for drug discovery, research, modeling, and large-scale data processing with an accent on multi-cloud architecture, high-throughput computing, and platform reliability. Focus on building scalable AWS/GCP environments, automating infrastructure with IaC, optimizing GPU-enabled workloads, and solving complex performance, observability, security, and cost-efficiency challenges.
Location: Hybrid in New York City, United States. Candidates must have permanent U.S. work authorization; U.S. work visa sponsorship is not available. Occasional travel may be required. Relocation assistance may be available based on business needs and eligibility.
Salary: $124,400–$207,400 annual base salary, plus a 17.5% bonus target and eligibility for a share-based long-term incentive program.
Company
develops medicines and applies computational science and scientific computing to drug discovery and development.
What you will do
- Define the technical vision, roadmap, standards, and architecture for cloud-based HPC services across AWS and GCP.
- Design and operate high-throughput, parallel, low-latency infrastructure supporting HPC and ML/AI workloads.
- Own OS image development, job scheduler configuration, high-performance storage, and core cloud services.
- Automate provisioning and lifecycle management with Terraform, CloudFormation, and infrastructure-as-code practices.
- Establish monitoring, logging, alerting, dashboards, and KPIs for reliability, workload optimization, and cost efficiency.
- Lead technical discussions with cloud providers, troubleshoot complex issues, and mentor HPC engineers and scientific computing specialists.
Requirements
- Bachelor’s degree in computer science, life science, data science, or a similar field, with 6+ years of cloud infrastructure engineering experience.
- Proven experience developing and supporting robust HPC frameworks in cloud environments.
- Expertise with AWS or GCP compute and storage services relevant to HPC.
- Strong knowledge of CI/CD, observability, distributed systems, production reliability, cloud networking, identity, and security controls.
- Experience with monitoring frameworks such as CloudWatch, Prometheus, or Grafana.
- Permanent U.S. work authorization required; visa sponsorship is not available.
Nice to have
- Master’s degree and 10–15 years of HPC or cloud engineering experience.
- Expertise with EKS, GKE, Kubernetes, distributed computing, HPC job schedulers, and NVIDIA GPU compute.
- Experience with AWS ParallelCluster, Parallel Computing Services, or Google Cloud Cluster Toolkit.
- Knowledge of Linux administration, cloud financial models, cost optimization, application delivery, user support, and resource optimization.
Culture & Benefits
- Hybrid work environment with occasional travel.
- Comprehensive medical, prescription drug, dental, and vision coverage.
- 401(k) matching and an additional retirement savings contribution.
- Paid vacation, holidays, personal days, caregiver/parental leave, and medical leave.
- Performance bonus target of 17.5% and eligibility for a share-based long-term incentive program.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →