обновлено 8 часов назад
Site Reliability Operations Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Operations Engineer (AWS/Cloud Infrastructure): Maintaining 24/7 stability of internal IT infrastructure and mission-critical backend systems with an accent on incident response, monitoring, automation, and cloud operations. Focus on troubleshooting complex Linux and Windows environments, coordinating zero-downtime deployments, and leading infrastructure migrations and performance improvements.
Location: Remote from Latin America; operations swing shift from 2:00 PM to 10:30 PM PST
Company
provides nearshore staff augmentation and technology solutions, connecting Latin American technology professionals with U.S. companies and digital transformation projects.
What you will do
- Monitor multi-platform IT infrastructure with AWS CloudWatch, New Relic, Nagios, and Sumo Logic, refining alert thresholds and enabling proactive remediation.
- Act as an escalation point for complex incidents across Linux/UNIX, Windows, virtual servers, and virtual desktop environments.
- Coordinate and automate code deployments through Jenkins, GitLab, or similar CI/CD tools, supporting non-disruptive releases and zero-downtime updates.
- Lead infrastructure projects involving migrations, cloud upgrades, and performance tuning.
- Collaborate with application developers, third-party vendors, and incident management teams.
- Maintain standard operating procedures and oversee enterprise backup operations with Commvault, Veeam, and AWS Backup.
Requirements
- 5+ years of experience in an Operations Center, SRO/NOC, or cloud infrastructure environment.
- Hands-on experience with full-stack application deployments, Windows and UNIX/Linux administration, virtualization, scripting, log analysis, and performance troubleshooting.
- Practical AWS experience across storage, virtual machines, and networking, plus enterprise monitoring tools.
- Programming or scripting skills in PowerShell, Python, or Bash.
- Experience with Jenkins or GitLab, ServiceNow or Jira, and enterprise backup solutions.
- Strong verbal and written communication skills for working with technical teams, executives, and external vendors.
Nice to have
- Advanced AWS certifications.
- Experience with AI/ML tools for infrastructure monitoring and predictive analytics.
- ITIL-aligned or enterprise Change and Incident Management experience.
- Bachelor’s degree in Computer Science, Information Technology, or a related field.
Culture & Benefits
- 100% remote work from Latin America with a reliable internet connection.
- Competitive compensation paid in USD.
- Paid time off and an emphasis on well-being and work-life balance.
- Autonomy to manage working time based on results.
- Opportunities to work with leading U.S. companies and collaborate with a multicultural network of more than 600 professionals across 25+ countries.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Site Reliability Engineer Engineer
4 дня назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
4 дня назад
Senior Site Reliability Engineer (AWS)
GlobalLogic
3 дня назад
Site Reliability Engineer (Senior SRE / Systems DevOps)
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
1 час назад