Назад
3 дня назад

Staff Software Engineer, RL Environments

252 000 - 315 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Software Engineer, RL Environments (Python/Distributed Systems): Building the platform for creating, running, verifying, and delivering reinforcement learning environments at scale with an accent on sandboxed execution, environment packaging, rollout orchestration, trajectory capture, and trustworthy reward signals. Focus on designing high-throughput systems, building adversarially robust graders, and translating research goals into production infrastructure.

Location: San Francisco, CA or New York, NY

Base salary: $252,000–$315,000 USD per year for eligible full-time positions in San Francisco, New York, and Seattle.

Company

Scale AI develops data, tooling, and full-stack technologies for training, evaluating, and deploying reliable AI systems.

What you will do

  • Own the technical foundation for building, running, verifying, and delivering reinforcement learning environments at scale.
  • Design sandboxed execution, environment packaging and versioning, rollout orchestration, trajectory capture, and verifier frameworks.
  • Build authoring tools that enable engineers and domain experts to create environments without repeatedly rebuilding infrastructure.
  • Instrument real applications and design task suites that expose specific model capability gaps.
  • Develop graders and reward signals that remain reliable under adversarial optimization.
  • Set technical direction across multiple teams while implementing complex systems hands-on.

Requirements

  • 8+ years of software engineering experience with strong distributed systems, system design, data structures, and algorithms fundamentals.
  • Strong Python skills and production software experience, plus familiarity with TypeScript/React, Go, Rust, or a similar technology.
  • Deep experience with containerization and sandboxed execution using Docker, virtual machines, gVisor, Firecracker, Kubernetes, or equivalent technologies.
  • Experience operating high-throughput backend systems involving orchestration, job scheduling, queuing, and large-scale data pipelines.
  • Hands-on experience with LLMs, including agent loops, tool calling, MCP, or evaluation harnesses.
  • Ability to own ambiguous problems end to end, establish technical direction, and communicate with engineers, researchers, and non-engineering partners.

Nice to have

  • Experience building reinforcement learning environments, agentic benchmarks, or evaluation harnesses.
  • Familiarity with RLHF, RLAIF, RLVR, GRPO/PPO-family algorithms, rejection sampling, reward modeling, and reward hacking prevention.
  • Experience with RL training or serving stacks such as verl, TRL, Ray, vLLM, or SGLang.
  • Experience with cloud infrastructure, Infrastructure as Code, CI/CD, observability, and high-scale sandbox or code-execution systems.
  • Experience building internal tools for daily use by non-engineers and working in research-adjacent engineering roles.

Culture & Benefits

  • Comprehensive health, dental, and vision coverage.
  • Retirement benefits, generous paid time off, and a learning and development stipend.
  • Equity compensation may be available for eligible roles, subject to approval.
  • A commuter stipend may be available for this role.
  • Inclusive and equal opportunity workplace with reasonable accommodation support.

Hiring process

  • Compensation is determined during the interview process based on work location, skills, experience, qualifications, and interview performance.
  • The same role may not be reconsidered until 90 days after a previous application.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →