9 дней назад
Senior Software Engineer, Platform (AI Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Engineer, Platform (AI Infrastructure): Building observability capabilities, internal tooling, and monitoring for a large-scale AI cloud platform with an accent on AI/ML application monitoring, low-level infrastructure metrics, and developer productivity. Focus on developing Prometheus exporters, integrating automated testing into CI/CD pipelines, and driving platform experiments from concept through production deployment.
Location: Singapore
Company
Technologies develops and operates sustainable AI infrastructure and a large-scale GPU cloud platform for AI model training and deployment across Asia Pacific.
What you will do
- Develop and integrate AI/ML application-level monitoring, including model accuracy tracking and performance observability.
- Build purpose-specific Prometheus exporters for low-level component and interconnect fabric monitoring.
- Enhance internal dashboards, CLI tools, automation scripts, and self-service portals to improve developer and operations productivity.
- Expand automated unit, integration, security, load, and end-to-end test coverage.
- Drive new product experiments from concept through production deployment and adoption.
- Contribute to AI-augmented development tools and workflows.
Requirements
- 7+ years of software engineering experience, including at least 3 years focused on platform or observability engineering.
- Bachelor’s degree in computer science or a related technical field.
- Strong experience with Go, Python, or Node; SQL, PromQL, LogQL, and GraphQL; and observability tools such as Loki, Grafana, Tempo, Prometheus, Thanos, or ClickHouse.
- Experience with Kafka or Pulsar, cloud platforms, Docker, configuration management, and CI/CD pipelines.
- Experience with automated testing frameworks and AI-augmented development tools and workflows.
- Clear and effective written and spoken English communication.
Nice to have
- Knowledge of Linux internals, networking stacks, distributed storage, or high-performance computing.
- Experience in high-growth startups or regulated industries with SOC 2 Type 2 and ISO 27001 requirements.
Culture & Benefits
- Full-time employment within an engineering and technology team.
- Work focused on sustainable, energy-efficient AI infrastructure.
- Inclusive workplace welcoming candidates from diverse backgrounds.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →