Data Engineer (Python)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Data Engineer (Python/Postgres): Building a data platform for regulatory indices by acquiring fragmented public data and transforming it into clean, auditable datasets with an accent on web scraping, medallion architecture, and index methodology. Focus on implementing resilient ingestion flows, designing scalable Postgres schemas, and ensuring mathematical transparency in published rankings.
Location: Remote (LATAM)
Company
builds AI-powered platforms that streamline regulatory complexity for clients in heavily regulated industries such as energy and government.
What you will do
- Source data from fragmented portals, APIs, PDFs, and commercial sites using advanced scraping techniques and anti-bot evasion.
- Design and manage Postgres schemas and migrations end-to-end using SQLAlchemy and Alembic.
- Build idempotent medallion (bronze → silver → gold) data pipelines with content hashing and lineage tracking.
- Implement and validate index methodologies, including winsorization, normalization, and weighting.
- Convert ambiguous index requirements into executable technical plans through product discovery.
- Orchestrate pipelines using Prefect and AWS ECS Fargate with comprehensive observability.
Requirements
- Must be based in LATAM.
- Strong proficiency in Python and SQL with production experience in Postgres schema design.
- Experience with lakehouse/medallion architecture, specifically idempotent ingestion and lineage.
- Advanced web scraping skills, including browser automation and resilience against hostile sources.
- Statistics literacy for index construction (normalization, weighting, and sensitivity testing).
- Comfort with modern Python tooling, typing, linting, and CI/CD discipline.
Nice to have
- Experience with Prefect, Airflow, or Dagster.
- AWS (ECS, S3) and Terraform.
- Proficiency with polars, pyarrow, or Jupyter/marimo.
- LLM-in-pipeline experience (pydantic-ai, AWS Bedrock, evals).
- Background in quantitative research, actuarial science, or government open data.
Culture & Benefits
- High-impact work at the intersection of AI and critical infrastructure regulation.
- Full end-to-end ownership of indices, from raw source to published methodology.
- Modern AI-native development environment utilizing Claude Code and Cursor.
- Remote-first culture with a small, high-velocity team.
- Competitive compensation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →