Назад
1 дСнь назад

Staff Software Engineer, MetalDev (Go)

207Β 000 - 275Β 000$
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Π“Ρ€Π΅ΠΉΠ΄
senior
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Staff Software Engineer, MetalDev (Go): Building and operating Go-based distributed services for bringing GPU data center infrastructure online, managing hardware lifecycles, and automating fleet operations with an accent on reliability, observability, and hardware-aware APIs. Focus on scaling Kubernetes infrastructure, translating hardware failures into resilient software improvements, and leading incident response across large GPU server fleets.

Location: New York, NY or Sunnyvale, CA, United States

Salary: $207,000–$275,000 per year, plus discretionary bonus, equity awards, and benefits.

Company

CoreWeave provides cloud infrastructure, tools, and services that help AI labs, startups, and enterprises build and scale AI applications.

What you will do

  • Design, build, and operate Go-based distributed services for large-scale GPU data center infrastructure.
  • Automate data center bring-up, hardware discovery, health monitoring, remediation, and production operations.
  • Develop REST and gRPC APIs and workflows for BMCs, firmware, server health, and rack-level infrastructure.
  • Improve observability, alerting, and operational tooling for rapid issue detection and resolution.
  • Translate incidents and hardware failure modes into reliability and resilience improvements.
  • Provide technical leadership through architecture, design and code reviews, project leadership, and mentoring.

Requirements

  • 8+ years of software engineering experience focused on infrastructure, cloud engineering, and distributed databases in large-scale data center or cloud environments.
  • Expertise in Go and experience building REST/gRPC APIs for mission-critical platforms.
  • Experience architecting and scaling cloud-native Kubernetes infrastructure and distributed services.
  • Hands-on experience with Prometheus, Grafana, PromQL, CI/CD pipelines, and large GPU server fleets.
  • Experience leading incident response and postmortems, with a strong focus on service reliability.
  • Bachelor’s, master’s, or doctoral degree in computer science or a related field, or equivalent experience.

Nice to have

  • Working knowledge of Kafka, ClickHouse, CRDB, DMTF, Redfish APIs, and GPU servers.
  • Experience contributing to and collaborating with open source communities.

Culture & Benefits

  • Medical, dental, and vision insurance fully paid by CoreWeave.
  • Life, disability, flexible spending, and health savings insurance benefits.
  • Tuition reimbursement, employee stock purchase program, and 401(k) with employer match.
  • Flexible PTO, paid parental leave, childcare support, and mental wellness benefits.
  • Catered lunch at office and data center locations, with a casual work environment.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’