обновлено 4 дня назад
Senior Platform Software Engineer (Streaming and Event-Driven Platforms)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Platform Software Engineer (Streaming and Event-Driven Platforms) (Kafka/Flink/Debezium): Operating and scaling production streaming and event-driven platforms with an accent on availability, reliability, fault tolerance, and schema evolution. Focus on managing Kafka clusters, Flink jobs, Debezium connectors, incident response, observability, and automation for 24/7 business-critical services.
Location: Palermo, Ciudad Autónoma de Buenos Aires, Argentina; hybrid with 3 days per week onsite
Company
provides Experience Cloud, a SaaS platform for managing experiences, insights, and actions across customer, employee, patient, candidate, and resident journeys.
What you will do
- Operate and maintain Kafka, Flink, Debezium, and Apicurio platforms in production.
- Manage Kafka brokers, partitions, replication, ISR health, consumer groups, upgrades, recovery, and lag mitigation.
- Operate Flink runtimes and jobs, including checkpoints, savepoints, state management, scaling, and recovery.
- Manage Debezium connector lifecycles, offsets, snapshots, failures, and schema changes, while supporting schema evolution through Apicurio Registry.
- Lead production incident triage, root-cause analysis, performance tuning, architecture reviews, and reliability improvements.
- Build automation, observability, alerts, documentation, operational playbooks, and participate in a periodic 24/7 on-call rotation.
Requirements
- At least 5 years of experience in software, systems, platform, DevOps, or reliability engineering, including 3 years operating distributed platforms in production.
- Hands-on experience configuring, scaling, and operating production Apache Kafka clusters, including partition reassignment, replication, ISR, consumer group rebalances, and consumer lag mitigation.
- Experience diagnosing and resolving distributed-systems incidents involving network partitions, data loss, service degradation, or high latency.
- Experience designing or implementing high-availability and fault-tolerance mechanisms.
- Professional working English proficiency, spoken and written.
- Experience with Java, Go, Python, or a similar programming language, plus formal architecture reviews and system design documentation.
Nice to have
- Production experience with Apache Flink, Debezium, Kafka Connect, or a schema registry.
- Knowledge of Avro, JSON Schema, Protobuf, and schema compatibility practices.
- Experience operating streaming platforms on Kubernetes or cloud infrastructure.
- Experience with platform automation, observability, major upgrades, disaster recovery, and failure-mode validation.
Culture & Benefits
- Hands-on work on large-scale, business-critical streaming platforms.
- Participation in a periodic on-call rotation supporting production reliability and performance.
- Commitment to diversity and equal employment opportunity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →