Назад
Company hidden
обновлено 5 дней назад

Senior Data Engineer (Cybersecurity)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Argentina
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior Data Engineer (Spark/Python): Building and maintaining the ML/AI data infrastructure for an email security platform with an accent on scalable data pipelines and feature engineering frameworks. Focus on designing Iceberg-based data lakes, optimizing petabyte-scale datasets, and supporting LLM-powered agents for threat detection.

Location: Cordoba, Argentina

Company

Global leader in human- and agent-centric cybersecurity protecting Fortune 100 companies and millions of organizations worldwide.

What you will do

  • Develop and maintain scalable data pipelines on AWS/Azure using Spark, Airflow, and Kubernetes to process email data at scale.
  • Design and optimize Iceberg-based data lake tables and schemas for petabyte-scale datasets distributed globally.
  • Build feature engineering frameworks supporting offline batch processing and online real-time feature serving for ML models.
  • Develop training data pipelines optimized for distributed ML model training, ensuring data lineage and reproducibility.
  • Collaborate with data scientists and security researchers to translate requirements into production-grade data solutions.
  • Mentor junior engineers and foster a culture of engineering excellence and knowledge sharing.

Requirements

  • Industry experience building distributed data systems and high-scale pipelines in cloud environments (AWS, Azure, or GCP) using Spark, Flink, or similar.
  • Deep proficiency in Python for developing production-grade data processing code.
  • Strong experience with Infrastructure-as-Code frameworks, particularly Terraform.
  • Hands-on experience with open table formats for data lakes (Apache Iceberg, Hudi, DeltaLake).
  • Experience with AWS Athena, Glue, or similar data query and cataloging services.
  • Proficiency with Apache Airflow or similar workflow orchestration tools for pipeline management.

Nice to have

  • Experience with feature stores such as Feast or Tecton.
  • Familiarity with Kubernetes for containerized data workloads.
  • Background in building data infrastructure specifically for machine learning and AI applications.
  • Experience with data quality frameworks and observability tools for pipelines.

Culture & Benefits

  • Competitive compensation and comprehensive benefits package.
  • Flexible work environment.
  • Annual wellness and community outreach days.
  • Global collaboration and networking opportunities.
  • Recognition culture and support for career success on your own terms.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →