Назад
Company hidden
обновлено 1 день назад

Senior Platform Software Engineer (Cache, Query and Compute Platforms)

138 000 - 205 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Platform Software Engineer (Cache, Query and Compute Platforms): Operating and scaling Redis, Kvrocks, Trino, and Spark platforms for reliable, high-performance production services with an accent on availability, fault tolerance, and operational readiness. Focus on leading incident response, designing failover and disaster recovery procedures, tuning distributed workloads, and building automation and observability for 24/7 systems.

Location: McLean, Virginia; hybrid with 3 days per week onsite. Candidates based in the Tysons vicinity will be prioritized.

Salary: $138,000–$205,000 annual base salary.

Company

hirify.global provides the hirify.global Experience Cloud, a SaaS platform for managing customer, employee, patient, candidate, and resident experiences.

What you will do

  • Operate and maintain Redis, Kvrocks, Trino, and Spark platforms in production.
  • Manage cluster deployments, replication, persistence, failover, upgrades, capacity, and performance.
  • Troubleshoot Trino query execution, catalogs, S3 access, Hive Metastore integrations, and Spark cluster health.
  • Lead production incident triage, root-cause analysis, performance tuning, and recovery validation.
  • Review application architectures and production-readiness requirements.
  • Build automation, monitoring, alerts, documentation, and troubleshooting playbooks while participating in a periodic 24/7 on-call rotation.

Requirements

  • At least 5 years of experience in software, systems, platform, DevOps, or reliability engineering.
  • At least 3 years of experience operating distributed platforms in production.
  • Hands-on production experience managing and scaling at least two of Redis, Kvrocks, Trino, or Spark.
  • Experience with incident triage, root-cause analysis, performance optimization, disaster recovery, automated failover, cluster re-sharding, and zero-downtime upgrades.
  • Programming experience in Java, Go, Python, or a similar language.
  • Experience conducting architecture reviews, writing technical design documents, and establishing operational runbooks.

Nice to have

  • Experience operating distributed platforms on Kubernetes or cloud infrastructure.
  • Experience integrating Trino with object storage, Hive Metastore, or external data catalogs.
  • Experience with observability, capacity planning, platform automation, and failure-mode testing.
  • Experience supporting large-scale, business-critical cache, query, or compute platforms.

Culture & Benefits

  • Medical, dental, and vision coverage.
  • 401(k), disability, life, and AD&D insurance.
  • Statutory leaves, paid parental leave, and paid holidays.
  • Equal opportunity workplace committed to diversity and inclusion.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →