Назад
Company hidden
обновлено 4 дня назад

Senior Platform Software Engineer (Streaming and Event-Driven Platforms)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Argentina
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Platform Software Engineer (Streaming and Event-Driven Platforms) (Kafka/Flink/Debezium): Operating and scaling production streaming and event-driven platforms with an accent on availability, reliability, fault tolerance, and schema evolution. Focus on managing Kafka clusters, Flink jobs, Debezium connectors, incident response, observability, and automation for 24/7 business-critical services.

Location: Palermo, Ciudad Autónoma de Buenos Aires, Argentina; hybrid with 3 days per week onsite

Company

hirify.global provides hirify.global Experience Cloud, a SaaS platform for managing experiences, insights, and actions across customer, employee, patient, candidate, and resident journeys.

What you will do

  • Operate and maintain Kafka, Flink, Debezium, and Apicurio platforms in production.
  • Manage Kafka brokers, partitions, replication, ISR health, consumer groups, upgrades, recovery, and lag mitigation.
  • Operate Flink runtimes and jobs, including checkpoints, savepoints, state management, scaling, and recovery.
  • Manage Debezium connector lifecycles, offsets, snapshots, failures, and schema changes, while supporting schema evolution through Apicurio Registry.
  • Lead production incident triage, root-cause analysis, performance tuning, architecture reviews, and reliability improvements.
  • Build automation, observability, alerts, documentation, operational playbooks, and participate in a periodic 24/7 on-call rotation.

Requirements

  • At least 5 years of experience in software, systems, platform, DevOps, or reliability engineering, including 3 years operating distributed platforms in production.
  • Hands-on experience configuring, scaling, and operating production Apache Kafka clusters, including partition reassignment, replication, ISR, consumer group rebalances, and consumer lag mitigation.
  • Experience diagnosing and resolving distributed-systems incidents involving network partitions, data loss, service degradation, or high latency.
  • Experience designing or implementing high-availability and fault-tolerance mechanisms.
  • Professional working English proficiency, spoken and written.
  • Experience with Java, Go, Python, or a similar programming language, plus formal architecture reviews and system design documentation.

Nice to have

  • Production experience with Apache Flink, Debezium, Kafka Connect, or a schema registry.
  • Knowledge of Avro, JSON Schema, Protobuf, and schema compatibility practices.
  • Experience operating streaming platforms on Kubernetes or cloud infrastructure.
  • Experience with platform automation, observability, major upgrades, disaster recovery, and failure-mode validation.

Culture & Benefits

  • Hands-on work on large-scale, business-critical streaming platforms.
  • Participation in a periodic on-call rotation supporting production reliability and performance.
  • Commitment to diversity and equal employment opportunity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →