6 дней назад
Principal Infrastructure Engineer (AI)
12 500 - 20 800$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Infrastructure Engineer (AWS/Kubernetes/AI): Building and scaling Sezzle’s production infrastructure across AWS, Kubernetes, Aurora RDS, observability, and infrastructure-as-code with an accent on throughput, latency, resilience, and cost efficiency. Focus on designing failure-tolerant platforms, recovering from major incidents, implementing disaster recovery, and developing AI-assisted SRE tooling with controlled access and auditability.
Location: Remote in Argentina
Salary: $12,500–$20,800 USD gross per month, based on location and experience.
Company
is a fintech and retail technology company building interest-free installment payment and shopping experiences for consumers and merchants.
What you will do
- Own the architecture and evolution of core infrastructure supporting increasing traffic, data volume, and workload complexity.
- Design resilient AWS account, IAM, networking, multi-AZ, and multi-region architectures.
- Build and operate Kubernetes platforms, including cluster lifecycle, workload isolation, resource allocation, autoscaling, upgrades, and deployment reliability.
- Scale and optimize Aurora RDS for MySQL and Postgres through query tuning, indexing, capacity planning, replication, failover, and safe migrations.
- Improve reliability through observability, SLOs, disaster recovery, incident response, infrastructure-as-code, safe migrations, and cloud cost optimization.
- Build AI-assisted tooling for incident investigation, capacity analysis, runbook automation, anomaly analysis, and toil reduction with bounded permissions and auditable actions.
Requirements
- Bachelor’s degree in Computer Science or a related technical field.
- 12+ years of experience designing and operating production infrastructure, platforms, SRE systems, or related engineering systems at scale.
- Deep production expertise with AWS, Kubernetes, and RDS/Aurora for MySQL and/or Postgres; EKS experience is strongly preferred.
- Strong coding and automation skills in Golang, Python, or similar languages, plus Terraform or equivalent infrastructure-as-code experience.
- Strong systems fundamentals across Linux, networking, DNS, TLS, storage, concurrency, and distributed-system failure modes.
- Experience with 24/7 high-availability platforms, incident response, on-call rotations, disaster recovery, observability, load testing, capacity planning, safe CI/CD, and active use of AI tooling.
Nice to have
- Experience in fintech, payments, or banking environments with demanding reliability, security, and audit requirements.
- Experience with multi-region architectures, chaos engineering, failure testing, and distributed data recovery tradeoffs.
- Proficiency with Prometheus, Grafana, Loki, Tempo, or comparable observability systems.
- Experience building internal platforms, self-service tooling, progressive delivery, reusable infrastructure components, or AI-assisted incident automation.
Culture & Benefits
- Work remotely from Argentina in a full-time role.
- Participate in an on-call rotation and take ownership of production recovery and postmortem improvements.
- Work closely with application engineering, Security, Compliance, and engineering leadership.
- Use open-source technologies and build internal capabilities where appropriate.
- Operate in a culture emphasizing high standards, direct communication, accountability, simplicity, measurable outcomes, and practical decision-making.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
7 дней назад
Engineering Manager, SRE (AI)
7 дней назад
Senior Infrastructure SRE
139 000 - 155 000CAD
6 дней назад
Senior Staff Site Reliability Engineer
232 338 - 290 422$
6 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
12 дней назад
Site Reliability Engineer II (AWS)
100 000 - 110 000$