Machine Learning Engineer, Infra, AI for Drug Discovery
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: South San Francisco, California, United States. The role is available in multiple locations, including California and New York. Relocation benefits are not available.
Salary: $147,600–$274,000 annually for the California location; $141,100–$262,100 annually for New York. A discretionary annual bonus may also be available.
Company
, a member of the Roche group and a biotechnology industry pioneer, develops medicines and applies computational science to serious and life-threatening diseases.
What you will do
- Design, implement, deploy, and operate scalable model-serving infrastructure for machine learning, scientific, LLM, and agentic workloads.
- Evolve the internal model deployment platform into a reliable self-service platform with reusable configurations, APIs, command-line tools, and documentation.
- Improve scalability and reliability through autoscaling, scale-to-zero, workload isolation, traffic management, faster startup, and reduced latency and request failures.
- Build observability and operational tooling for model usage, latency, reliability, resource consumption, inference cost, and service-level indicators.
- Develop model lifecycle infrastructure covering registration, versioning, evaluation, promotion, release gates, monitoring, rollback, and retraining integrations.
- Partner with machine learning, data, scientific, and platform teams while owning workstreams from design through production support.
Requirements
- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- At least 3 years of relevant experience in software, infrastructure, platform engineering, DevOps, MLOps, or a related area.
- Strong Python skills and experience shipping maintainable production software, services, automation, or developer tooling.
- Experience operating cloud systems, preferably AWS, including services such as EKS, EC2, S3, IAM, SQS, SNS, and CloudWatch.
- Experience with containers, Kubernetes, Helm, Terraform or Pulumi, CI/CD, automated testing, and Git-based release practices.
- Knowledge of distributed systems, observability, troubleshooting with metrics and logs, and communicating technical tradeoffs.
Nice to have
- Experience with KServe, Triton, vLLM, Ray Serve, Prefect, Dagster, or similar serving and workflow-orchestration frameworks.
- Experience optimizing model startup, throughput, batching, autoscaling, or GPU utilization.
- Familiarity with model registries, experiment tracking, model evaluation, data drift, regression analysis, or MLOps platforms.
- Experience with event-driven systems, scientific computing, high-performance computing, distributed training, or large-scale data processing.
- Strong interest in life sciences and drug discovery.
Culture & Benefits
- Work within Roche's AI for Drug Discovery group and Computational Sciences Center of Excellence.
- Contribute to scientific and production workflows supporting drug discovery and medicine development.
- Benefits are available according to the company's benefits program.
- A discretionary annual bonus may be available based on individual and company performance.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →