Назад
4 дня назад

Senior Principal Engineer II (Kafka)

192 000 - 358 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Principal Engineer II (Kafka) (Distributed Systems/Kafka): Defining and building the next-generation Kafka Background Plane for data and metadata compaction, retention, segment lifecycle management, object garbage collection, and multi-tenant task orchestration with an accent on correctness, scalability, and fault tolerance. Focus on designing distributed scheduling and coordination, resolving concurrency and partial-failure challenges, and driving multi-year architecture across storage, cloud infrastructure, control planes, and SRE.

Location: Remote within the United States; eligible locations include Poughkeepsie, Lowell, Rochester, Tucson, Research Triangle Park, Armonk, Boston, Bellevue, Atlanta, San Jose, Dallas, Austin, San Francisco, Seattle, Georgia, Minnesota, Texas, New York, North Carolina, Arizona, Washington, Massachusetts, and California. Some travel may be required based on business demand.

Salary: $192,000–$358,000 per year.

Company

IBM Software develops hybrid cloud, AI, and data platforms that help organizations operate across cloud, on-premises, private datacenter, and edge environments.

What you will do

  • Define the multi-year technical strategy and architecture for the next-generation Kafka Background Plane.
  • Design and optimize distributed scheduling, task coordination, isolation, autoscaling, backpressure, and dependency-execution workflows.
  • Build correctness and safety models for compaction, log rewriting, retention, segment lifecycle, metadata, and object garbage-collection workflows.
  • Scale background processing across partitions, throughput, storage, and lag while preserving sub-millisecond foreground latency.
  • Lead cross-organization architecture decisions and global production rollouts across storage, serving engines, consensus layers, cloud infrastructure, control planes, and SRE.
  • Establish standards for fault tolerance, chaos testing, observability, canary safety, incident mitigation, and technical mentorship.

Requirements

  • Bachelor's degree required; a master's degree is preferred.
  • Extensive experience architecting, building, and operating mission-critical distributed systems at planetary scale.
  • Deep expertise in storage engines, distributed log stores, stream processing, replication, or consensus protocols.
  • Strong knowledge of Apache Kafka internals, including storage semantics, log compaction, tiered storage, and segment indexing, or equivalent streaming systems.
  • Experience delivering multi-year architectural roadmaps, defining SLA/SLO and durability targets, and establishing automated recovery practices.
  • Outstanding technical communication, problem-solving, cross-team leadership, and mentorship skills.

Nice to have

  • Experience with log-structured storage, custom storage engines, cloud object stores, or large-scale data-rewriting pipelines.
  • Experience designing fault-tolerant object lifecycle strategies and preventing data leakage, orphaned resources, or premature deletion.
  • Experience with chaos engineering, traffic replay, deterministic simulation testing, or large-scale load generation.
  • Experience with multi-cloud, multi-region deployments across AWS, GCP, and Azure.

Culture & Benefits

  • Regular employment with healthcare, dental, vision, mental health, life insurance, disability coverage, and retirement programs including 401(k).
  • Paid time off includes holidays, sick time, vacation, parental bonding leave, and other paid care leave programs.
  • Access to an AI-driven learning platform, training resources, and industry-recognized certifications.
  • Inclusive employee resource groups, volunteer opportunities, and discounts on products and services.
  • Environment focused on continuous learning, experimentation, feedback, collaboration, and personal responsibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →