15 часов назад
Systems Development Engineer
190 000 - 230 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Systems Development Engineer (AI Infrastructure): Improving the reliability, operability, and customer experience of a production AI orchestration platform with an accent on production debugging, automation, and observability. Focus on building diagnostics and durable platform fixes across cloud infrastructure, Kubernetes, distributed systems, networking, storage, IAM, and deployment systems.
Location: Hybrid, based out of the Seattle office
Salary: $190K–$230K annually
Company
develops Flyte, an open-source data and AI orchestration platform for production workloads at scale.
What you will do
- Investigate and resolve customer-impacting production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability.
- Convert recurring customer issues into automation, product improvements, runbooks, tests, and design changes.
- Build internal tools and diagnostics that improve issue detection, investigation, and resolution.
- Improve logs, metrics, dashboards, alerts, and customer-visible debugging information.
- Participate in system design and development, promoting operationally reliable and easier-to-debug systems.
- Define production-readiness, alert-quality, runbook, observability, regression-prevention, and code-quality practices while reducing on-call load and time to resolution.
Requirements
- Strong software engineering skills in Python, Go, Java, or a similar language.
- Experience debugging production systems across multiple layers of the stack.
- Practical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM.
- Experience with infrastructure as code, deployment systems, CI/CD, observability, and operational automation.
- Ability to turn ambiguous customer symptoms into technical diagnoses and durable remediation.
- Clear written and verbal communication, including root-cause analyses, runbooks, technical recommendations, and design feedback.
Nice to have
- Experience operating customer-facing SaaS, cloud infrastructure, self-hosted or on-premises deployments, or workflow orchestration systems.
- Experience with batch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability.
- Experience improving on-call health, reducing ticket volume, or building production diagnostics.
- Experience working across support, customer success, product, and engineering teams.
Culture & Benefits
- Full engineering-team support for on-call responsibilities.
- Medical coverage with 100% of employee premiums and 90% of dependent premiums paid.
- Dental and vision coverage with 90% of premiums paid for employees and dependents.
- Meaningful equity options, unlimited time off, and 12 company holidays.
- 401(k) matching, paid parental leave, flexible scheduling, and onsite meals and snacks for office employees.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer (AI Infrastructure)
200 000 - 240 000$
5 часов назад
Infrastructure Engineer (AI)
200 000 - 300 000$
3 дня назад
Sr. DevOps Engineer II (AI)
119 000 - 221 000$
5 дней назад
Software Development Engineer (US Federal)
137 000 - 205 400$
22 часа назад
Software Development Engineer (AWS Federal Cloud)
137 000 - 205 400$
6 дней назад
Software Engineer, Development Tools (DevOps)
150 000 - 250 000$