20 ΡΠ°ΡΠΎΠ² Π½Π°Π·Π°Π΄
Site Reliability Engineer (SRE)
100Β 000 - 180Β 000$
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Π’Π΅ΠΊΡΡ:
TL;DR
Site Reliability Engineer (SRE) (Kubernetes/Cloud): Operating and improving large-scale distributed systems in production with an accent on observability, incident response, and infrastructure automation. Focus on designing reliable Kubernetes platforms, building CI/CD and monitoring systems, conducting capacity and performance engineering, and strengthening resilience through chaos engineering and failure-mode analysis.
Location: 100% remote within the United States
Salary: $100,000β$180,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Define and refine SLOs, SLIs, and error budgets for critical production services.
- Lead incident response, act as an incident commander when needed, and conduct post-incident reviews.
- Design monitoring, logging, and tracing solutions using Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar tools.
- Automate operational workflows with Python, Go, Bash, or similar languages.
- Architect and operate Kubernetes clusters, container workloads, autoscaling, capacity planning, network policies, and service-mesh integrations.
- Build CI/CD pipelines and improve reliability through canary releases, chaos engineering, failover design, security improvements, and developer enablement.
Requirements
- 10+ years of SRE, DevOps, or production engineering experience with large-scale distributed systems.
- Bachelorβs degree in Computer Science, Engineering, or a related technical discipline.
- Strong programming skills in Python, Go, or Java.
- Deep hands-on Linux experience, including networking, performance tuning, and systems troubleshooting.
- Production experience with Kubernetes, containers, observability tooling, and CI/CD pipelines.
- Experience with distributed-system design, incident response, documentation, and technical communication.
Nice to have
- Experience with production SLOs, error budgets, chaos engineering, capacity planning, performance engineering, or large-scale load testing.
- Experience with AWS, Azure, or GCP.
- Familiarity with Istio, Linkerd, Consul, Chaos Monkey, Gremlin, or Litmus.
Culture & Benefits
- Full-time direct W2 employment.
- Career growth opportunities within an established organization.
- Blameless culture focused on operational excellence.
- Mentorship and collaboration with application development and security teams.
- New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates are encouraged to apply.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β
ΠΠΎΡ ΠΎΠΆΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Staff Production Engineer (SRE)
140Β 400 - 372Β 300$
1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Staff Production Engineer (SRE) (Federal)
119Β 000 - 170Β 000$
2 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Systems Engineer (SRE) (AWS/Kubernetes)
4Β 750 - 7Β 670β¬
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Senior Site Reliability Engineer (Fintech)
160Β 000 - 200Β 000$
3 Π΄Π½Ρ Π½Π°Π·Π°Π΄
Staff Site Reliability Engineer (AI/Blockchain)
195Β 000 - 257Β 500$
1 Π΄Π΅Π½Ρ Π½Π°Π·Π°Π΄
Site Reliability Engineer (SRE), Data Products
135Β 000 - 160Β 000$