3 дня назад
Senior MLOps Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior MLOps Engineer (AI) (GCP/Vertex AI): Building and operating production infrastructure, deployment pipelines, and monitoring systems for complex AI models with an accent on scalable GCP architecture, low-latency inference, and ML observability. Focus on converting research prototypes into secure auto-scaling services, automating training and deployment, and supporting reliable enterprise-grade AI workloads.
Location: Remote Bulgaria
Company
develops cybersecurity solutions that help customers monitor, manage, and protect risks associated with digital identities and personal information.
What you will do
- Architect and manage scalable machine learning infrastructure on Google Cloud Platform using Vertex AI, GKE, Cloud Storage, Cloud Run, and GPU/TPU instances.
- Own end-to-end model deployment and build high-throughput, low-latency inference services with Docker, Triton Inference Server, vLLM, and MLflow.
- Build reproducible training, testing, evaluation, and deployment pipelines with Airflow, Vertex AI Pipelines, and GitHub Actions.
- Implement system and ML observability, including latency, throughput, uptime, feature drift, prediction accuracy, and data distribution monitoring.
- Provide scalable training environments and standardized deployment templates for AI and research engineers.
- Lead the transition of AI prototypes and notebooks into secure, resilient, auto-scaling microservices while integrating feature stores, dataset versioning, and stream or batch processing.
Requirements
- At least 5 years of hands-on experience designing, deploying, and maintaining production machine learning workloads in cloud environments.
- Deep experience with GCP, including Vertex AI, Cloud Storage, GKE, Cloud Run, IAM, and VPC configurations.
- Expertise with Docker, Kubernetes/GKE, Triton Inference Server, vLLM, and MLflow.
- Experience with Airflow, Vertex AI Pipelines, GitHub Actions, ArgoCD, and Terraform.
- Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
- Hands-on experience with logging, telemetry, and ML drift detection using Grafana, Prometheus, GCP Cloud Monitoring, or similar tools.
Nice to have
- Experience running large-scale LLM or deep learning inference and training workloads.
- GCP Professional Machine Learning Engineer or Professional Cloud Architect certification.
- Familiarity with feature stores such as Feast or Vertex AI Feature Store.
Culture & Benefits
- Work on cybersecurity products addressing real customer protection needs.
- Operate in a nimble, growth-oriented organization where individual contributions have visible impact.
- Opportunity to learn new technologies, products, and markets as the company expands.
- Inclusive workplace committed to preventing discrimination and harassment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →