Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior ML Engineer (AI/LLM infrastructure): Designing and operating infrastructure for machine learning and LLM systems, from training and evaluation workflows to production model serving, with an accent on Kubernetes, cloud platforms, CI/CD, and observability. Focus on building GPU-based workloads, evolving LLM platforms, improving production reliability, and moving ML solutions from experimentation into scalable services.
Location: Tel Aviv, Israel
Company
Cyera builds a unified security control plane for protecting data, access, and AI systems.
What you will do
- Design and build workflows for model training, large-scale evaluation, batch prediction, and experimentation.
- Develop ML infrastructure and developer tooling, including automated processes, CI/CD improvements, and model build flows.
- Operate production ML and LLM services, improving monitoring, observability, reliability, and performance.
- Deploy and operate GPU-based training, evaluation, and inference workloads on Kubernetes and cloud infrastructure.
- Contribute to self-hosted model serving, LLM gateways, internal SDKs, evaluation infrastructure, and integrations with model providers.
- Own engineering solutions end-to-end and collaborate with research, data science, backend, and DevOps teams.
Requirements
- 4+ years of experience in ML engineering, software engineering, backend engineering, or a similar hands-on engineering role.
- B.Sc. in Computer Science, Software Engineering, or a related technical field, or equivalent practical experience.
- Strong Python and software engineering skills for building maintainable, production-quality systems.
- Hands-on experience with Kubernetes, Docker, cloud environments such as AWS, GCP, or Azure, and distributed production services.
- Experience with ML infrastructure and lifecycle areas including orchestration, model serving, evaluation, training, or production inference.
- Experience with CI/CD, observability, production operations, end-to-end problem solving, and cross-functional technical collaboration.
Nice to have
- GPU workloads or high-scale ML inference, including vLLM, BentoML, Argo Workflows, or Kubeflow.
- LLM platforms and infrastructure, including gateways, evaluation and observability tooling, self-hosted models, or internal SDKs.
- Agentic systems, memory layers, agent evaluation, or related infrastructure.
Culture & Benefits
- Ownership of technical solutions and working practices.
- Encouragement to take initiative, move quickly, and turn ideas into impact.
- Collaborative environment focused on shared success, learning, and continuous improvement.
- Inclusive workplace welcoming diverse backgrounds, perspectives, and experiences.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →