Назад
Company hidden
4 часа назад

Data Engineer (AWS/PySpark)

192 000 - 281 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Engineer (AWS/PySpark): Building and operating production data pipelines that retrieve, transform, and integrate data sources into a mission-critical cyber operations platform with an accent on PySpark, AWS Glue, schema design, and restricted environments. Focus on diagnosing production failures, improving observability and data quality, and packaging integrations for classified deployments.

Location: On-site full-time at hirify.global's New York City office; initial onboarding of approximately two weeks at the Arlington, VA office, followed by occasional travel to Arlington. U.S. citizenship and eligibility to obtain a U.S. Government security clearance are required.

Salary: $192,000–$281,000 base; estimated total compensation of $281,000–$455,000.

Company

hirify.global develops software for offensive cyber operations supporting the United States and its allies.

What you will do

  • Design and build AWS Glue data pipelines in PySpark to retrieve, parse, transform, and load new data sources.
  • Investigate unfamiliar source systems, define schemas and analytical data models, and integrate data into applications using ClickHouse and PostgreSQL.
  • Deploy, operate, and package integrations for office-accessible and restricted environments.
  • Monitor pipeline health, data quality, throughput, and freshness using the LGTM observability stack, dashboards, and alerts.
  • Diagnose production incidents, perform root cause analysis, and improve runbooks, fixes, and upstream designs.
  • Work with forward deployed engineers to support field deployments and turn operational feedback into platform capabilities.

Requirements

  • 5+ years of experience in data engineering or software engineering with substantial data infrastructure work.
  • Expert Python skills and experience maintaining production-grade codebases.
  • Hands-on experience building and operating ETL/ELT pipelines with Spark, preferably AWS Glue jobs using PySpark.
  • Strong SQL and schema design skills, including partitioning and indexing decisions, plus experience with column-oriented analytical databases.
  • Experience debugging production data systems, tracing failures, and resolving incidents under time pressure.
  • Ability to work on-site full-time in New York City; U.S. citizenship and eligibility to obtain a U.S. Government security clearance are required.

Nice to have

  • Experience with data lakes, large-scale query technologies, streaming systems, or message queues.
  • Experience with observability tooling such as Grafana, Datadog, or Splunk.
  • Experience operating software in air-gapped, classified, or restricted environments.
  • Familiarity with Docker, container orchestration, and CI/CD.

Culture & Benefits

  • Hands-on ownership of production systems used by cyber operations teams.
  • Work on data integrations for restricted and classified environments with strong testing, documentation, and operational discipline.
  • Medical, dental, and vision plan options, plus life, disability, HSA, and FSA options.
  • Paid parental leave, paid holidays, and flexible PTO.
  • 401(k) with pre-tax and Roth options, plus dependent care FSA.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →