14 часов назад
Site Reliability Engineer, Compute - USDS
136 800 - 359 720$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Compute - USDS (Go/Python, Kubernetes, distributed systems): Building and operating large-scale, fault-tolerant computing systems with an accent on automation, scalability, monitoring, and incident response. Focus on designing robust infrastructure, diagnosing production issues, defining SLOs, SLIs, and SLAs, and resolving performance bottlenecks in distributed systems.
Location: San Jose, United States; fully in-person schedule up to 5 days a week
Salary: $136,800–$359,720 annually, plus potential discretionary bonuses, incentives, and restricted stock units.
Company
USDS is a TikTok joint venture focused on data privacy, cybersecurity, national security, and the protection of U.S. user data and applications.
What you will do
- Develop and maintain automation procedures that improve system efficiency and reduce manual intervention.
- Design, deploy, and operate robust systems in collaboration with software engineering teams.
- Ensure scalability for growing web traffic and data across large-scale distributed systems.
- Implement monitoring tools, metrics, and performance tests to identify system health issues and bottlenecks.
- Participate in on-call rotations, incident management, production diagnosis, resolution, and prevention.
- Define SLOs, SLIs, and SLAs while conducting sustainable user support and blameless postmortems.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field, plus 3+ years of experience.
- Professional experience as a Site Reliability Engineer, Systems Engineer, or similar software engineering professional.
- Programming experience with Go, Python, or other languages, with a focus on automation and operational excellence.
- Experience with network architecture, database modeling, cloud systems, and large-scale distributed systems.
- Strong knowledge of Linux operating systems and open-source technologies.
- Excellent debugging, problem-solving, strategic thinking, communication, and cross-functional collaboration skills.
Nice to have
- Knowledge of monitoring tools and methodologies such as Prometheus and Grafana.
- Experience with Docker, Kubernetes, or equivalent container orchestration platforms.
Culture & Benefits
- Fully in-person work schedule supporting close collaboration, alignment, and execution.
- Medical, dental, and vision insurance from day one.
- 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
- Wellbeing benefits, 10 paid holidays, 10 paid sick days, and 17 days of Paid Personal Time.
- Work in a culture emphasizing intellectual curiosity, openness, self-direction, diversity, and blameless problem-solving.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
19 часов назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
9 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
7 часов назад
Site Reliability Engineer, Edge Services - USDS
136 800 - 359 720$
1 день назад
Staff Site Reliability Engineer (GCP/Kubernetes)
112 500 - 187 500$
Reddit
5 дней назад
Staff Software Engineer (Observability)
217 000 - 303 900$
6 дней назад
Site Reliability Engineer (AI Infrastructure)
200 000 - 240 000$