15 дней назад
Senior MLOps Engineer (GCP)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (GCP): Architecting and maintaining GCP infrastructure, ML deployment pipelines, and production tooling for complex AI models with an accent on scalable serving, automated training and deployment, and ML observability. Focus on deploying low-latency inference services, building drift-aware monitoring and retraining workflows, and scaling AI prototypes into secure, auto-scaling microservices.
Location: Remote Estonia
Company
develops cybersecurity solutions that help customers monitor, manage, and protect risks related to digital identities and personal information.
What you will do
- Architect and manage scalable GCP-based machine learning infrastructure with Vertex AI, GKE, GCS, Cloud Run, and GPU/TPU compute.
- Own the end-to-end deployment and serving lifecycle for machine learning models, including high-throughput, low-latency inference services.
- Build automated and reproducible pipelines for model training, testing, evaluation, and deployment.
- Implement monitoring for system health and ML metrics, including latency, throughput, prediction accuracy, feature drift, and data distribution shifts.
- Provide standardized training environments, runtime infrastructure, and deployment templates for AI and research engineers.
- Partner with data engineering teams on feature stores, dataset versioning, and stream or batch processing, and lead the transition of prototypes into secure, auto-scaling microservices.
Requirements
- At least 5 years of hands-on experience designing, deploying, and maintaining production ML workloads in cloud environments.
- Deep practical experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, IAM, and VPC configurations.
- Expertise with Docker, Kubernetes/GKE, Triton Inference Server, vLLM, and MLflow.
- Experience with Airflow, Vertex AI Pipelines, GitHub Actions, ArgoCD, and Terraform.
- Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
- Hands-on experience with logging, telemetry, and ML observability tools such as Grafana, Prometheus, and GCP Cloud Monitoring.
Nice to have
- Experience with large-scale LLM or deep learning inference and training workloads.
- GCP Professional Machine Learning Engineer or Professional Cloud Architect certification.
- Familiarity with feature stores such as Feast or Vertex AI Feature Store.
Culture & Benefits
- Work on cybersecurity products addressing evolving customer digital security needs.
- Operate in a nimble, growth-oriented organization where individual contributions have visible impact.
- Opportunity to learn new technologies, products, and markets as the company expands.
- Inclusive workplace committed to preventing discrimination and harassment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →