8 часов назад
Senior SRE (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior SRE (AI) (Python/Kubernetes): Designing and building scalable infrastructure, CI/CD pipelines, and production systems that support AI data operations with an accent on infrastructure automation, reliability, and observability. Focus on developing core infrastructure components, optimizing deployment and batch-processing systems, and applying site-reliability practices to maintain availability and performance.
Location: Canada; remote
Salary: Market competitive salary with quarterly variable compensation
Company
A mission-driven organization combining advanced technology with a global network of people to make unusable data usable and create real-world impact at scale.
What you will do
- Design and implement scalable core infrastructure components with a high degree of autonomy.
- Optimize deployment pipelines, environment provisioning, production operations, and high-throughput batch jobs.
- Use Infrastructure as Code tools such as Terraform to manage and scale complex infrastructure.
- Develop multi-step CI/CD pipelines covering build, testing, deployment, monitoring, environment setup, and artifact handling.
- Improve production reliability, availability, and performance through monitoring, alerting, observability, and site-reliability practices.
- Collaborate with software engineering, product, and business stakeholders, explaining complex technical issues to technical and non-technical audiences.
Requirements
- 5+ years of experience building and operating infrastructure in production environments.
- Fluent Python and strong experience writing production-ready code.
- Experience with Docker, Kubernetes, cloud platforms such as GCP or AWS, and Infrastructure as Code tools such as Terraform.
- Experience with CI/CD platforms and automated build, test, and deployment pipelines.
- Experience applying site-reliability principles including availability, observability, and automation across production systems.
- Degree in Computer Science, Engineering, or a quantitative or computational field, or equivalent practical experience.
Nice to have
- Familiarity with Prometheus or Grafana.
- Experience with Ansible, Chef, Puppet, or similar configuration management tools.
- Exposure to multi-cloud or hybrid-cloud environments.
Culture & Benefits
- Mission-driven, people-centric, innovative, and globally connected working environment.
- Full-time fixed-term employee position with an expected duration of 6 months.
- Remote workplace with a hybrid working model mentioned among the benefits.
- Market competitive salary and quarterly variable compensation.
- Comprehensive medical cover and group life insurance.
- Personal development and professional growth opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →