Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Building and maintaining scalable infrastructure for deploying machine learning models with an accent on reliability, performance, and automation. Focus on establishing best practices, mentoring team members, and collaborating with cross-functional teams.
Location: Hybrid in San Francisco, Montreal, New York, Seattle, or Toronto; remote work is also listed.
Salary: $165,000–$330,000 per year, plus equity.
Company
Baseten provides inference infrastructure and developer tooling that help AI companies bring machine learning models into production.
What you will do
- Build and maintain scalable infrastructure for deploying and operating machine learning models.
- Establish infrastructure standards and best practices for reliability and performance.
- Automate deployment processes, including CI/CD pipeline management.
- Own products and projects end-to-end, from specification through execution.
- Collaborate with cross-functional teams to translate requirements into technical solutions.
- Mentor junior engineers and contribute to organizational knowledge sharing.
Requirements
- Bachelor’s, master’s, or Ph.D. degree in computer science, engineering, mathematics, or a related field.
- Extensive experience with Kubernetes and scalable infrastructure.
- Experience with infrastructure-as-code tools such as Terraform, CloudFormation, or Pulumi.
- Experience with CI/CD tools such as GitHub Actions, GitLab CI, CircleCI, or Jenkins.
- Ability to own projects from specification through execution and make sound technical trade-offs.
- Ability to work in the listed hybrid or remote locations.
Nice to have
- Open-source observability experience with Prometheus, ELK, Grafana, or OpenTelemetry.
- Interest in learning about machine learning; prior machine learning experience is not required.
Culture & Benefits
- Competitive compensation with meaningful equity.
- Flexible paid time off and a company-wide winter break.
- Paid parental leave and a fertility and family-building stipend.
- U.S. employees receive company-paid medical, dental, and vision insurance for themselves and dependents.
- U.S. employees have access to a company-facilitated 401(k).
- Exposure to a range of machine learning startups and related learning opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
9 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
5 дней назад
Staff Site Reliability Engineer (AI)
220 000 - 260 000$
9 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
10 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
8 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$