4 часа назад
Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Cloud/Observability): Operating and improving monitoring platforms, cloud infrastructure, CI/CD pipelines, and customer-facing services with an accent on observability, incident response, and cybersecurity remediation. Focus on detecting and resolving production anomalies, defining service level indicators, automating operational work, and strengthening infrastructure security across cloud environments.
Location: Brazil; remote
Company
helps large enterprises transform through AI deployment, application modernization, data and AI, marketing technology, and business strategy.
What you will do
- Operate and maintain monitoring platforms, dashboards, monitors, log pipelines, APM instrumentation, synthetic tests, and Real User Monitoring.
- Investigate logs, traces, metrics, error tracking data, failures, regressions, and anomalies; lead triage, root-cause analysis, and resolution.
- Tune alerting and notification routing, define service level indicators, and improve reliability, scalability, and cost efficiency.
- Instrument services with meaningful metrics, structured logs, and distributed traces while partnering with engineering teams on observability.
- Improve cloud infrastructure, CI/CD pipelines, release processes, runbooks, escalation paths, and operational documentation.
- Partner with InfoSec and engineering teams to remediate infrastructure security issues, including TLS, security headers, DNS configuration, and exposed technology information.
Requirements
- 3+ years of experience in site reliability engineering, DevOps, platform engineering, or production-focused software engineering.
- Hands-on experience with New Relic, Grafana, Splunk, or Dynatrace, including dashboards, monitoring, log management, and APM.
- Strong troubleshooting and root-cause analysis skills using logs, traces, and metrics.
- Working proficiency in Python, TypeScript/JavaScript, Go, Bash, or another scripting or programming language.
- Experience with AWS, GCP, or Azure and solid knowledge of networking, containers, and Linux fundamentals.
- Familiarity with Terraform, CloudFormation, Pulumi, GitHub Actions, on-call work, incident response, and post-incident reviews.
Culture & Benefits
- Remote home-office work in Brazil.
- Health and dental insurance, life insurance, meal and food allowance, and childcare assistance.
- Extended paternity leave and responsible parenting support.
- Profit sharing and results participation.
- Access to learning platforms, CI&T University, language learning, wellness programs, and discount partnerships.
- Inclusion specialists, affinity groups, and workplace support for people with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Site Reliability Engineer Engineer
Valletta.Software | AI-Care
11 часов назад
Senior DevOps / SRE Support Engineer (LATAM)
5 000 - 5 500$
4 дня назад
Site Reliability Engineering, Engineer II (Kubernetes)
5 дней назад
Senior Site Reliability Engineer (AWS)
Вкусно - и точка
7 дней назад
SRE-инженер
7 часов назад