обновлено 1 день назад
Principal Applied AI Developer (Foundation Models Infrastructure)
153 000 - 224 400CAD
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Applied AI Developer (Foundation Models Infrastructure): Building and scaling cloud-native platform services for training, inference, evaluation, deployment, and serving of machine learning models with an accent on Kubernetes, distributed computing, reliability, security, and developer self-service. Focus on defining platform architecture, leading cross-team technical initiatives, improving production operations, and enabling safe, cost-effective AI workflows at scale.
Location: Remote from Canada; the role is based in Quebec, Canada
Salary: CAD 153,000–224,400 base salary, with potential additional bonuses, equity, and benefits.
Company
develops cloud-scale software, data platforms, and AI-enabled tools for designing, building, and operating the world around us.
What you will do
- Define and drive the technical strategy for foundation model and machine learning infrastructure.
- Design and operate resilient, secure, observable, scalable, and cost-effective platform services across the model lifecycle.
- Build developer-facing APIs, tools, workflows, and self-service capabilities for researchers and ML developers.
- Work hands-on with Kubernetes, Ray, SageMaker, AWS, and cloud-native technologies supporting distributed training and scalable inference.
- Lead complex cross-team initiatives, influence technical direction, and translate ambiguous research and business requirements into executable engineering plans.
- Establish platform standards for reliability, observability, governance, deployment, versioning, incident response, and production readiness.
Requirements
- Bachelor’s or master’s degree in computer science, computer engineering, machine learning, or equivalent practical experience.
- At least 8 years of professional software engineering experience with large-scale, cloud-native, distributed, platform, or machine learning infrastructure systems.
- Deep experience designing, building, and operating production-grade services for model training, inference, serving, evaluation, deployment, or observability.
- Strong experience with Kubernetes, cloud-native infrastructure, distributed computing, and technologies such as Ray or SageMaker.
- Experience with CI/CD, automated testing, infrastructure as code, monitoring, alerting, production operations, and resilient system design.
- Proven ability to lead cross-team technical initiatives, mentor developers, influence stakeholders without direct authority, and communicate effectively.
Nice to have
- Experience with AI-assisted development, coding-agent orchestration, MCPs, prompt and context design, evaluation loops, and reusable AI-enabled workflows.
- Experience building internal developer platforms, ML platforms, self-service infrastructure, or platform APIs.
- Experience with large-scale training, fine-tuning, batch or real-time inference, model serving, and foundation model workflows.
- Knowledge of model governance, versioning, lineage, Trusted AI, security, privacy, and responsible AI practices.
Culture & Benefits
- Work remotely from Canada within a global software and AI organization.
- Collaborate with AI researchers, ML developers, product teams, architects, security, privacy, and platform groups.
- Participate in Agile, Kanban, and other modern development methodologies.
- Competitive base salary with potential annual bonuses, equity grants, and a comprehensive benefits package.
- Work in a culture focused on belonging, ownership, quality, accountability, and meaningful product impact.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →