8 часов назад
Data Platform Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Platform Engineer (Scala/Spark): Building and operating TB-scale batch and streaming data pipelines and platform components with an accent on reliability, performance, security, and cost efficiency. Focus on designing fault-tolerant ETL/ELT systems, owning SLOs and incident response, and developing governance, observability, and self-service capabilities for analytics teams.
Location: Cape Town, hybrid, with two days per week in the office
Company
operates a commerce partnership marketing platform that helps brands manage affiliate, influencer, content publisher, and customer referral partnerships.
What you will do
- Own data platform components and pipeline domains end to end, including design, implementation, reliability, performance, and cost.
- Build scalable batch and streaming ETL/ELT pipelines with Spark, Dataproc, Kafka, Pub/Sub, BigQuery, and other data stores.
- Maintain platform infrastructure including Dataproc, Kafka, Airflow/Astronomer, BigQuery, and Cloud Storage, while optimizing queries, partitioning, throughput, and resource usage.
- Define and monitor SLOs for freshness, success, and latency; implement monitoring, alerting, data quality checks, and Tier 2 incident response.
- Implement security and governance practices including access control, secrets management, encryption, audit logging, lineage, retention, and data contracts.
- Develop self-service platform capabilities, document architectures and runbooks, participate in code reviews, and mentor Associate Data Platform Engineers.
Requirements
- 3–5 years of experience in data platform, data engineering, or backend/distributed systems engineering, including ownership of production systems at TB scale or systems with stringent SLAs.
- Strong Python and production experience with a JVM language; Scala is the primary pipeline language and requires a willingness to become proficient.
- Hands-on experience with distributed batch processing using Spark or streaming with Kafka or Pub/Sub, plus working familiarity with the other area.
- Advanced SQL, workflow orchestration with Airflow, and experience with a major cloud platform, ideally GCP.
- Experience with Git, CI/CD, automated testing, monitoring, alerting, incident response, on-call, and secure data-system design.
- Clear written and verbal communication, with experience or aptitude for mentoring less experienced engineers.
Nice to have
- Infrastructure as code with Terraform or an equivalent tool and automated deployment experience.
- Experience migrating legacy big-data platforms to cloud-native alternatives.
- Familiarity with dbt, Looker, semantic layers, Astronomer Cosmos, Scala functional programming, or JVM tuning.
- Experience with Grafana, Datadog, Cloud Monitoring, Dataflow, SingleStore, or BigTable.
- A bachelor's degree in Computer Science, Engineering, Mathematics, or a related field, or experience in digital marketing technology.
Culture & Benefits
- Responsible PTO policy and a flexible environment supporting work-life balance.
- Up to 12 fully covered therapy or coaching sessions per year, with additional dependent coverage.
- Monthly gym reimbursement, a technology stipend for the home office, and a monthly internet allowance.
- Restricted Stock Units with a three-year vesting schedule, pending Board approval.
- Free Coursera subscription, PXA courses, and paid parental leave of up to 26 weeks for the primary caregiver and 13 weeks for the secondary caregiver.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →