Senior Data Engineer (Cybersecurity)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Data Engineer (Spark/Python): Building and maintaining the ML/AI data infrastructure for an email security platform with an accent on scalable data pipelines and feature engineering frameworks. Focus on designing Iceberg-based data lakes, optimizing petabyte-scale datasets, and supporting LLM-powered agents for threat detection.
Location: Cordoba, Argentina
Company
Global leader in human- and agent-centric cybersecurity protecting Fortune 100 companies and millions of organizations worldwide.
What you will do
- Develop and maintain scalable data pipelines on AWS/Azure using Spark, Airflow, and Kubernetes to process email data at scale.
- Design and optimize Iceberg-based data lake tables and schemas for petabyte-scale datasets distributed globally.
- Build feature engineering frameworks supporting offline batch processing and online real-time feature serving for ML models.
- Develop training data pipelines optimized for distributed ML model training, ensuring data lineage and reproducibility.
- Collaborate with data scientists and security researchers to translate requirements into production-grade data solutions.
- Mentor junior engineers and foster a culture of engineering excellence and knowledge sharing.
Requirements
- Industry experience building distributed data systems and high-scale pipelines in cloud environments (AWS, Azure, or GCP) using Spark, Flink, or similar.
- Deep proficiency in Python for developing production-grade data processing code.
- Strong experience with Infrastructure-as-Code frameworks, particularly Terraform.
- Hands-on experience with open table formats for data lakes (Apache Iceberg, Hudi, DeltaLake).
- Experience with AWS Athena, Glue, or similar data query and cataloging services.
- Proficiency with Apache Airflow or similar workflow orchestration tools for pipeline management.
Nice to have
- Experience with feature stores such as Feast or Tecton.
- Familiarity with Kubernetes for containerized data workloads.
- Background in building data infrastructure specifically for machine learning and AI applications.
- Experience with data quality frameworks and observability tools for pipelines.
Culture & Benefits
- Competitive compensation and comprehensive benefits package.
- Flexible work environment.
- Annual wellness and community outreach days.
- Global collaboration and networking opportunities.
- Recognition culture and support for career success on your own terms.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →