1 месяц назад
Site Reliability Engineer Intern
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer Intern (Python/Ansible): Monitoring and improving the reliability of critical applications and global data center infrastructure with an accent on automation, observability, and incident management. Focus on building automated remediation tools, troubleshooting production incidents, and coordinating failover testing for network, systems, and application environments.
Location: Dallas, Texas, United States — headquarters
Company
operates a global online vehicle auction platform connecting vehicle sellers with more than 750,000 buyers in over 190 countries.
What you will do
- Monitor global data centers and application infrastructure, identifying and resolving issues before they affect operations.
- Design, build, and optimize monitoring and automation tools with Python and Ansible for metrics collection and automated remediation.
- Maintain observability and collaboration tooling, including Datadog and Kubernetes-related tools.
- Support incident management through triage, impact assessment, severity classification, tracking, root cause analysis, and post-incident reviews.
- Coordinate periodic failover testing across network, systems infrastructure, and application environments.
- Partner with Product Development, DevOps, Network, Systems, Database, and Infrastructure teams while maintaining SOPs, diagrams, and training materials.
Requirements
- Core knowledge of Linux and Windows systems, virtual environments, basic networking, scripting, automation, and observability tools.
- Experience triaging, tracking, and resolving incidents in a production environment.
- Ability to assess impact, classify severity, communicate incident updates, and document timelines and resolutions.
- Strong troubleshooting, analytical, written, oral, and interpersonal communication skills.
- Ability to work independently, manage competing priorities, and perform effectively in a flexible work schedule.
- Intermediate programming and scripting proficiency.
Nice to have
- Experience with VMware vSphere and virtual machine management.
- Familiarity with Datadog and other monitoring or observability tools.
- Experience with AWS or GCP.
- Basic understanding of AI tools, prompt-based workflows, and automation-enhanced operational support.
Culture & Benefits
- Full-time role supporting 24/7 operations and stringent SLA commitments.
- Collaborative environment focused on diversity, inclusion, and cross-functional cooperation.
- Exposure to global data center infrastructure and critical application operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →