16 часов назад
Principal Software Engineer - Site Reliability Platform (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Software Engineer - Site Reliability Platform (AI) (SRE platforms and distributed services): Designing and building internal reliability platforms for cloud access policies, cost governance, observability, operational excellence, and availability with an accent on AI-powered production systems, scalability, and performance. Focus on closing reliability gaps through hands-on integrations, improving livesite mitigation and postmortems, driving adoption across engineering teams, and mentoring software engineers.
Location: Bangalore, India
Company
creates enterprise automation software and software robots designed to accelerate human achievement.
What you will do
- Design, build, and ship SRE platform systems and AI-powered capabilities used by engineering teams in their critical path.
- Participate in livesite monitoring rotations, handle escalations, drive mitigations, and lead detailed postmortems.
- Improve availability, scalability, performance, reliability, and operational excellence based on production learnings.
- Drive platform adoption by writing integrations, pairing with engineers, removing friction, and measuring delivered outcomes.
- Plan and estimate work, coordinate scheduling and staffing, and influence engineering processes and best practices.
- Mentor software engineers through hands-on coaching, advice, and training.
Requirements
- 10+ years of experience architecting and engineering large-scale distributed commercial applications and services.
- Experience building complex internal platforms adopted by 10+ teams in their critical path.
- Experience building and maintaining complex AI-powered applications in production.
- Proficiency in one or more object-oriented languages, such as C#, C++, Go, or Python, with strong computer science fundamentals.
- Deep knowledge of data structures, algorithms, multithreading, synchronization, asynchronous patterns, cloud programming, service-oriented and microservice architectures, HTTP applications, and web services.
- Experience with Azure, AWS, or GCP and managed services such as AKS or GKE is required; experience with databases including Azure SQL, CosmosDB, Azure Data Lake, Power BI, MongoDB, MySQL, or DynamoDB is expected.
Nice to have
- Experience working with or managing production Kubernetes infrastructure.
Culture & Benefits
- Work with globally distributed engineering teams.
- Many roles offer flexibility in when and where work gets done, depending on business and team needs.
- Applications are assessed on a rolling basis with no fixed deadline; the requisition may close when a qualified candidate is selected.
- Inclusive workplace with equal-opportunity practices and reasonable accommodations available on request.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Site Reliability Engineer - AI Enablement
1 день назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
2 дня назад
Senior SRE Engineer (AI)
1 день назад
Platform Engineer (AI)
187 000 - 250 000$
2 дня назад
Staff / Principal Platform Engineer (AI)
280 000 - 350 000$
2 дня назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$