Senior Site Reliability Engineer - Workflow Automation (Apache Airflow)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Site Reliability Engineer - Workflow Automation (Apache Airflow): Owning and improving the reliability, scalability, and operational excellence of workflow orchestration platforms with an accent on Apache Airflow, Automic/UC4, observability, and infrastructure automation. Focus on designing platform tooling, modernizing deployments, resolving production incidents, and reducing operational toil across complex scheduling environments.
Location: Hybrid role based in Austin, United States
Company
is an investment management firm offering financial products and services.
What you will do
- Provide primary production support and incident escalation for Apache Airflow and Automic/UC4 platforms.
- Define and improve SLOs, SLIs, error budgets, monitoring, capacity, and platform performance.
- Design automation and tooling that reduce operational toil and improve developer experience.
- Modernize platforms through workload migrations, deployment improvements, containerization, and managed services.
- Develop infrastructure as code and observability solutions using tools such as Terraform, Helm, Ansible, ELK, Grafana, and Prometheus.
- Partner with data engineering, application, and security teams while maintaining runbooks and on-call procedures.
Requirements
- Bachelor’s degree in a technical field or equivalent practical experience.
- 5+ years of experience in SRE, DevOps, or platform engineering.
- Deep hands-on experience with Apache Airflow, including distributed executors, DAG best practices, and multi-environment deployments.
- Experience with enterprise job scheduling platforms, strong Linux and Windows knowledge, and cloud environments, preferably AWS.
- Proficiency in Python and shell scripting, plus experience with Kubernetes, Docker, and CI/CD pipelines.
- Strong observability knowledge and the ability to resolve incidents and communicate clearly under pressure.
Nice to have
- Experience with Automic/UC4 and managed Airflow services such as AWS MWAA, Cloud Composer, or Astronomer.
- Familiarity with dbt, Kafka, Snowflake, or other data platform technologies.
- Experience migrating workloads from legacy schedulers and managing error budget or SLO frameworks.
Culture & Benefits
- Hybrid work environment with participation in on-call rotations.
- Comprehensive benefits supporting employees and their families.
- Educational initiatives and career development programs.
- Programs and celebrations focused on company history, culture, and growth.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →