12 часов назад
Site Rel Engineer Sr. - Model Integration Platform (Linux/GPFS/Slurm)
86 250 - 158 125$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Rel Engineer Sr. - Model Integration Platform (Linux/GPFS/Slurm): Stabilizing production environments and managing a model integration platform with an accent on site reliability, monitoring, capacity planning, and incident resolution. Focus on designing SLO and SLA dashboards, tuning complex systems, resolving priority incidents, and driving root-cause analysis and disaster recovery improvements.
Location: In-office at an Information Technology Hub in Cleveland, Ohio; Pittsburgh, Pennsylvania; Dallas, Texas; Birmingham, Alabama; Denver, Colorado; or Phoenix, Arizona.
Base salary: $86,250–$158,125 per year, plus incentive eligibility.
Company
A financial services organization focused on delivering customer solutions while managing enterprise risk.
What you will do
- Stabilize production environments and sites through reliability engineering, analytics, metrics, site design consulting, platform management, and capacity planning.
- Design and implement monitoring systems and dashboards for applications, service sites, and platforms.
- Establish and implement service-level agreements and objectives using operational monitoring.
- Analyze production metrics, perform performance tuning, troubleshoot priority incidents, and participate in blameless post-mortems.
- Lead complex incident response, testing strategy, root-cause analysis, and continuous improvement to reduce mean time to resolution.
- Mentor and train junior team members on infrastructure management and disaster recovery best practices.
Requirements
- Linux administration, particularly RHEL, Linux networking and security, storage and file systems including LVM and NAS.
- IBM GPFS / Spectrum Scale and Slurm Scheduler administration.
- Shell/Bash scripting and basic Python.
- Spark, Jupyter, production support, incident resolution, performance monitoring, and troubleshooting.
- Bachelor’s degree and 3+ years of relevant industry experience, or a comparable combination of education, certifications, and experience.
- Ability to work in the office at one of the listed U.S. Information Technology Hubs; employment visa sponsorship and STEM OPT participation are not available.
Nice to have
- IBM Spectrum Conductor, AWS or Azure, Ansible/Python automation, VMware or KVM, Docker or Kubernetes.
- Oracle, MySQL, or PostgreSQL databases; Tableau or Power BI reporting.
- Apache Tomcat installation, tuning, and administration.
Culture & Benefits
- Inclusive, supportive in-office workplace with an emphasis on customer focus and risk management.
- Medical, prescription, dental, and vision coverage, with a Health Savings Account option.
- 401(k) matching, pension, stock purchase plans, insurance, disability protection, and dependent care support.
- Educational assistance, wellness programs, paid holidays, parental leave, and 15–25 vacation days depending on career level and tenure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Sr. Site Reliability Engineer
130 000 - 140 000$
3 дня назад
Site Reliability Engineer (Fintech)
120 000 - 155 000$
4 дня назад
Site Reliability Engineer/L3 Support (AWS/Kubernetes)
110 000 - 130 000$
6 дней назад
Staff Site Reliability Engineer (Fintech)
160 000 - 210 000$
4 дня назад
Site Reliability Engineer (Cloud)
100 000 - 140 000$
5 дней назад
Site Reliability Engineer, Storage - Enterprise Technology
200 000 - 250 000$