4 дня назад
Sr. Member of Technical Staff (AI Inference)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. Member of Technical Staff (AI Inference): Designing resilient, high-availability software and cloud deployment workflows for scalable AI inference services with an accent on AWS infrastructure, automation, container orchestration, and distributed observability. Focus on building fault-tolerant recovery mechanisms, optimizing parallel inference workloads, and debugging model deployment and networking issues across Kubernetes-based environments.
Location: Headquarters/Sunnyvale Office, United States; on-site
Company
Systems develops large-scale AI computing hardware and software for high-speed model training and inference.
What you will do
- Design and develop resilient, highly available software features with automated recovery and fault-tolerant architecture.
- Build and maintain AWS-based deployment workflows for low-latency, scalable AI inference services.
- Develop Python scripts and APIs for data preprocessing, inference execution, and post-processing.
- Use multithreading and asynchronous processing to improve resource efficiency on AWS compute instances.
- Deploy inference software in Docker containers and define Kubernetes orchestration strategies for reliable scaling.
- Monitor, debug, and resolve service defects using logs, metrics, distributed traces, CloudWatch, Grafana, and related tools; document infrastructure, APIs, defects, and releases.
Requirements
- Master's degree or foreign equivalent in Computer Science or a related field.
- At least 18 months of experience as an Information Security Analyst, Software Engineer, Senior Member of Technical Staff, IT Senior Applications Engineer, or in a related role.
- Infrastructure-as-Code and deployment automation experience with Terraform, AWS CloudFormation, AWS CDK, and Ansible.
- Containerization and orchestration experience with Docker, Kubernetes, AWS EKS, AWS ECS, AWS Fargate, and Helm.
- Experience with AWS EC2, AWS Lambda, Auto Scaling Groups, monitoring and tracing tools, PostgreSQL, Redis, NFS, Jenkins, and Git.
- Programming experience with Python, Node, JavaScript, and Flask.
Culture & Benefits
- Work on an AI platform designed beyond the constraints of GPUs.
- Opportunities to publish and open-source AI research.
- Work with one of the world's fastest AI supercomputers.
- Startup vitality combined with job stability.
- Non-corporate culture focused on individual beliefs, learning, growth, and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Developer Experience Engineer (AI/HPC)
150 000 - 275 000$
2 дня назад
Senior Staff/Principal Deployment Automation Engineer (AI)
250 000 - 300 000$
4 дня назад
Sr. DevOps Engineer (AI)
175 000 - 195 000$
4 дня назад
Forward Deployed Engineers (AI Infrastructure)
180 000 - 240 000$
2 дня назад
Forward Deployed Engineer (AI)
4 дня назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$