1 день назад
Platform Ops Lead (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Ops Lead (Kubernetes): Leading platform operations support for a modern GitLab and Kubernetes DevOps platform serving biomedical applications, with an accent on operational reliability, developer enablement, and infrastructure troubleshooting. Focus on resolving deployment and runtime issues, maintaining error budgets and SLOs, improving incident response through postmortems, and guiding a team of support engineers.
Location: Bethesda, Maryland, United States; remote options available
Company
provides professional services and technology support for NCBI, a biomedical information organization within the National Library of Medicine at the National Institutes of Health.
What you will do
- Lead the platform operations support team assisting internal developers with migration from legacy development and deployment workflows.
- Operate and improve a modern DevOps platform based on GitLab and Kubernetes across cloud and on-premises environments.
- Identify and resolve operational, deployment, and runtime problems in microservice environments.
- Prioritize incidents against error budgets and SLOs, provide technical solutions, and support on-call operations.
- Create SOPs and process documentation, compile postmortems, and track action items to reduce future outages.
- Interview, recommend, train, and support new team members.
Requirements
- Bachelor’s degree in STEM or equivalent experience.
- Strong systems debugging skills and comfort with Linux or the UNIX command line.
- Experience with a programming or scripting language and with creating process, procedure, and SOP documentation.
- General understanding of TCP/IP, HTTP, and related protocols.
- Customer-focused, team-oriented communication skills, with the ability to work with users at varying levels of IT knowledge.
- Ownership, initiative, sound judgment, integrity, and responsibility.
Nice to have
- Experience with Kubernetes, OpenShift, cloud platforms, Linux systems administration, and service reliability engineering.
- Automation with Bash, Ruby, Python, Go, Java, Scala, Rust, C++, or Perl; Puppet is preferred for configuration management.
- GitLab or TeamCity CI, Git, automated CI/CD pipelines, GitOps tools such as ArgoCD, and service mesh technologies such as Linkerd or Istio.
- Monitoring and alerting with Grafana, Prometheus, OpsGenie, or the TIGK stack.
- Docker, Linux internals and networking, attached network storage, AWS, GCP, Azure, Google Anthos, and distributed systems design.
Culture & Benefits
- Flexible working hours and remote work options.
- Medical, dental, and vision coverage.
- 401(k) plan with employer contribution.
- Paid holidays, vacation, tuition reimbursement, and access to on-site and off-site training courses.
- Conference attendance support in a technical and scientific environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →