5 дней назад
Site Reliability Engineer II (GCP)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer II (GCP/BigQuery): Maintaining reliable, observable, and performant cloud and network systems across GCP-based data platforms with an accent on automation, monitoring, troubleshooting, and continuous optimization. Focus on building access-monitoring and log-recording tooling, improving incident response and on-call efficiency, and optimizing BigQuery workloads, CI/CD pipelines, and production capacity.
Location: Dearborn, Michigan, United States; remote work from home
Company
provides technology and consulting services, supporting clients with cloud platforms, enterprise systems, and emerging technologies.
What you will do
- Automate routine infrastructure tasks and implement reliability solutions with infrastructure teams.
- Monitor and manage production environments, identify issues proactively, and lead troubleshooting.
- Build tooling for system access monitoring, log session recording, and reliability administration across distributed data centers.
- Improve on-call efficiency, incident management, and post-mortem analysis with engineering teams.
- Perform capacity planning and optimize system performance, stability, security, and traffic handling.
- Maintain monitoring and alerting systems, documentation, diagrams, cloud infrastructure, BigQuery workloads, and CI/CD pipelines.
Requirements
- Bachelor’s degree and at least 4 years of IT experience.
- At least 3 years of development experience and practitioner-level experience with a coding language or framework.
- Hands-on experience with Google Cloud Platform and BigQuery.
- Experience with Dynatrace or comparable monitoring and observability tools such as Datadog or New Relic.
- Familiarity with ServiceNow and incident, problem, and change management.
- Remote work from home in the United States.
Nice to have
- Experience with GCP Cloud Run and Python.
- Experience defining and tracking SLAs, SLOs, and SLIs.
- Familiarity with AI tools, including agents, skills, LLMs, and copilots.
- Strong troubleshooting and problem-solving skills.
Culture & Benefits
- Opportunities to collaborate with technical experts and clients.
- Long-term career development and access to emerging technologies.
- Health, dental, and vision coverage.
- Paid time off, paid holidays, 401(k) matching, life and disability insurance.
- Professional development opportunities and wellness programs.
- Inclusive workplace focused on diversity, dignity, and respect.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
AI Ops Engineer
220 000 - 240 000$
6 дней назад
SW Engineer - Developer Systems Reliability Engineering (AI-AIOps)
88 000 - 136 900$
7 дней назад
Sr. Manager, Site Reliability
6 дней назад
Staff Site Reliability Engineer (GCP/Kubernetes)
3 дня назад
Senior Data Platform Engineer (Web3)
126 000 - 180 000$
6 дней назад