19 часов назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure): Building and operating large-scale, massively distributed, and fault-tolerant systems with an accent on automation, scalability, monitoring, and incident response. Focus on designing reliable cloud infrastructure, diagnosing production issues, defining SLOs, SLIs, and SLAs, and preventing recurring failures through performance testing and blameless postmortems.
Location: San Jose R&D; fully in-person schedule up to 5 days a week
Salary: $136,800–$259,200 annually
Company
Joint Venture operates TikTok-related apps with a focus on data privacy, cybersecurity, national security, and protection of U.S. user data.
What you will do
- Develop and maintain automation procedures to improve system efficiency and reduce manual intervention.
- Design, deploy, and operate robust systems in collaboration with software engineering teams.
- Build for scalability across web traffic, data growth, and large-scale distributed systems.
- Implement monitoring, metrics, and performance tests to identify bottlenecks and track system health.
- Participate in on-call rotations, incident management, diagnosis, resolution, and prevention of production issues.
- Define SLOs, SLIs, and SLAs while conducting sustainable user support and blameless postmortems.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field and 3+ years of experience.
- Professional experience as a Site Reliability Engineer, Systems Engineer, or similar software engineering professional.
- Proficiency in Python, Go, Java, or Shell scripting.
- Experience with network architecture, database modeling, cloud systems, and large-scale distributed systems.
- Strong understanding of Linux operating systems and open-source technologies.
- Availability to work in person in San Jose up to 5 days per week.
Nice to have
- Experience with Docker, Kubernetes, or equivalent container orchestration platforms.
- Knowledge of monitoring tools and methodologies such as Prometheus and Grafana.
- Strong debugging, problem-solving, strategic thinking, and cross-functional communication skills.
Culture & Benefits
- Medical, dental, and vision insurance from day one.
- 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
- Wellbeing benefits, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
- Collaborative, diverse, intellectually curious, and problem-solving-oriented work environment.
- On-call participation and blameless incident response practices.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
24 часа назад
Site Reliability Engineer (Tech Infrastructure)
129 960 - 246 240$
14 часов назад
Site Reliability Engineer, Tech Infra
129 960 - 246 240$
13 часов назад
Site Reliability Engineer, Compute - USDS
136 800 - 359 720$
17 часов назад
Senior Site Reliability Engineer (Compute)
177 688 - 341 734$
18 часов назад
Site Reliability Engineer, Reliability Team - USDS
122 574 - 259 200$
1 день назад
Site Reliability Engineer, AI Infrastructure (AI)
129 960 - 246 240$