4 дня назад
Production Support Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Production Support Engineer (Incident Management/API Troubleshooting): Monitoring and supporting a revenue-critical sales platform by triaging incidents, investigating APIs and logs, and coordinating communication with partners, call centers, and engineering with an accent on incident ownership, stakeholder communication, and SLA management. Focus on leading post-mortems, tracking corrective actions, prioritizing defects, and responding to alerts during U.S. Eastern business hours.
Location: Remote from Argentina, Brazil, or Colombia. Availability during U.S. Eastern business hours, 9 AM–6 PM ET is required; Eastern timezone is strongly preferred for onboarding and on-call coordination.
Company
is an AI-native consulting and technology services firm delivering cloud, data, software engineering, and artificial intelligence solutions for enterprise transformation.
What you will do
- Monitor platform health, triage incidents, and distinguish critical outages or service degradation from non-critical bugs and defects.
- Investigate incidents using API calls, responses, logs, and observability tools.
- Own incident response from initial alert through post-mortems, corrective actions, and closure.
- Notify stakeholders of critical issues, communicate SLA risks, and translate technical findings for partners and call center teams.
- Manage Jira incident queues, cross-reference tickets, and prioritize defects within engineering sprint cycles.
- Participate in cross-functional meetings and join on-call rotations after ramp-up, responding to OpsGenie alerts within defined SLA windows.
Requirements
- 2+ years of troubleshooting and resolving issues in application, server, or infrastructure environments.
- 2+ years of providing clear status updates and resolution information to stakeholders at multiple levels.
- Working knowledge of APIs, including the ability to interpret API calls and responses; experience with Postman or a similar tool.
- Experience with logging and observability tools such as Splunk, Datadog, or Sumo Logic, plus SQL troubleshooting and ad hoc reporting.
- Basic ability to read HTML and JSON and use browser developer tools; able to participate in technical bridge calls and follow incidents to resolution.
- Availability during U.S. Eastern business hours, 9 AM–6 PM ET, with strong communication, prioritization, organization, and customer service skills.
Nice to have
- AWS experience at Cloud Practitioner level or above.
- Familiarity with Git, Terraform or similar infrastructure-as-code tools, and OpsGenie or similar alerting platforms.
- Experience with AI tooling, deployment and release processes, on-call rotations, and incident severity frameworks.
- Experience communicating with partners or call centers during live incidents.
- Background in travel, hospitality, or high-volume transactional platforms.
Culture & Benefits
- Comprehensive onboarding documentation and a structured six-month ramp to full self-sufficiency.
- On-call rotations begin after readiness is established, with manager support during early rotations.
- Active alert coverage runs from 8 AM to 1 AM Eastern, with overnight suppression windows.
- SEV-1 incidents are rare, occurring roughly once per quarter or less; most work takes place during business hours.
- Global collaboration, continuous learning, innovation, and emphasis on ethical AI standards.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →