14 дней назад
Senior Site Reliability Engineer (Applied Machine Learning)
206 086 - 341 734$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (Applied Machine Learning): Designing and operating large-scale distributed machine learning, recommendation, and LLM inference infrastructure with an accent on MLOps automation, fault tolerance, and system reliability. Focus on architecting ML platform components, building multi-stage CI/CD pipelines, analyzing complex distributed failures, and leading incident response and mentoring.
Location: Bellevue, Washington, United States; Seattle is also listed under conditions. Full-time, onsite role, 40 hours per week.
Base salary: $206,086–$341,734 annually, with potential additional bonuses, incentives, and restricted stock units.
Company
A technology joint venture focused on data privacy, cybersecurity, national security, and protecting U.S. user data and applications.
What you will do
- Design and architect large-scale distributed machine learning, recommendation, and LLM inference systems.
- Make performance, cost, and maintainability decisions for ML platform components and workloads.
- Build sophisticated MLOps automation and multi-stage CI/CD pipelines for ML and LLM model deployment.
- Perform root cause analysis for distributed system failures and develop fault-tolerant, self-healing components.
- Provide user support, respond to incidents, conduct postmortems, and participate in technical operations rotations.
- Mentor junior site reliability engineers and interns.
Requirements
- Master’s degree or foreign equivalent in Computer Science, Engineering, Information Technology, or a related field plus 3 years of related experience; or bachelor’s degree or foreign equivalent plus 5 years of progressive post-bachelor’s experience.
- At least 3 years of experience provisioning servers with Linux operating systems.
- At least 3 years of experience managing system resources and deployments with configuration management tools.
- At least 3 years of experience building tools with programming and scripting languages.
- At least 3 years of experience processing large volumes of logs and data, creating reports, dashboards, and alerts with analytical and search-processing tools.
- At least 3 years of experience deploying and managing applications with containerized tools.
Culture & Benefits
- Inclusive workplace with reasonable accommodations available during recruitment.
- Medical, dental, and vision insurance from the first day of employment.
- 401(k) savings plan with company match, life insurance, disability coverage, and wellbeing benefits.
- Paid parental leave, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →