Назад
Company hidden
8 дней назад

Senior Systems Reliability Engineer (Java)

109 000 - 150 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Systems Reliability Engineer (Java): Building and operating production reliability systems for the Texas power grid with an accent on SLO governance, observability, Java application reliability, and NERC/CIP compliance. Focus on designing self-healing automation, diagnosing complex JVM and distributed-system failures, implementing chaos engineering, and validating dual-datacenter failover.

Location: Taylor, TX, United States; hybrid work required two days per week

Salary: $109,000–$150,000 per year

Company

hirify.global develops and operates solutions supporting the Texas power grid and wholesale electricity market.

What you will do

  • Design, build, and operate reliability tooling, automation frameworks, self-healing systems, and production software.
  • Own SLI/SLO governance, error budgets, alert quality, and observability architecture for assigned systems.
  • Lead high-severity incident response, dual-datacenter failover, root-cause analysis, and actionable post-mortems.
  • Implement chaos engineering programs, CI/CD reliability gates, capacity planning, and production-readiness standards.
  • Diagnose and improve Java/Spring Boot and JVM performance across PostgreSQL, Oracle, Kafka, ActiveMQ, and REST/SOAP integrations.
  • Design reliable Kubernetes/OpenShift platforms, infrastructure-as-code, environment promotion pipelines, and automated patch compliance workflows.

Requirements

  • At least 5 years of progressive experience in systems reliability, software engineering with an SRE focus, or a closely related field.
  • Production experience with SLO frameworks, observability platforms, chaos engineering, automated remediation, and incident response.
  • Strong Python and Java skills, including deep Java/Spring Boot and JVM performance expertise.
  • Experience with Linux, Bash, CI pipelines, Kubernetes, OpenTelemetry, basic networking, and Kubernetes/OpenShift workload reliability.
  • Expertise with the Grafana LGTM stack, Dynatrace, Splunk, and infrastructure-as-code using Terraform, Ansible/AAP, or equivalent.
  • Bachelor’s degree in Computer Science, Software Engineering, MIS, or a related field is required.

Nice to have

  • Experience with dual-datacenter or hybrid cloud reliability architecture.
  • NERC/CIP compliance, control implementation, audit preparation, and regulatory engagement experience.
  • Master’s degree in a related field.
  • Azure, Certified Kubernetes Administrator, or ITIL certification.

Culture & Benefits

  • Collaborative, inclusive environment focused on accountability, leadership, innovation, trust, and expertise.
  • Participation in a 24/7 on-call rotation.
  • Formal mentoring and opportunities to contribute technical standards, talks, documentation, and shared tooling.
  • Work supports the development of solutions for current and future energy challenges.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →