2 дня назад
Software Engineer, AI Systems
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineer, AI Systems (AI/LLM): Building and operating production reasoning systems that connect incident evidence, knowledge graphs, retrieval, and specialized AI agents with an accent on traceability, evaluation, and human oversight. Focus on orchestrating multi-step tool-calling workflows, grounding model outputs, diagnosing production regressions, and monitoring quality, cost, latency, and failure modes.
Location: United States; remote. Relocation will not be considered.
Company
is a VC-backed pre-seed venture building an enterprise learning intelligence layer for safety-critical organizations in partnership with The AES Corporation and AI Fund.
What you will do
- Build and operate multi-step LLM reasoning pipelines with model calls, tool calls, graph queries, retrieval, quality gates, and specialist-agent handoffs.
- Extend the coordinated agent orchestration layer supporting incident investigation, review, causal analysis, and enterprise learning.
- Design grounding and retrieval across Neo4j graph traversal, vector search, hybrid retrieval, company knowledge, and historical cases.
- Build evaluation datasets, scoring systems, regression suites, model comparisons, human-label loops, and per-stage quality attribution.
- Implement tracing, tool-call audits, cost and latency monitoring, failure handling, and quality dashboards for production AI systems.
- Select and compare models from OpenAI, Anthropic, and Google while shaping the AI roadmap with product and knowledge engineering.
Requirements
- Production AI or ML engineering experience, including shipping LLM systems used by real users.
- Hands-on experience building and debugging multi-step, tool-calling workflows with LangGraph, LangChain, or an equivalent framework.
- Experience with repeatable LLM evaluation using representative datasets, regression testing, LLM-as-judge methods, or human review loops.
- Experience assembling retrieval context for LLMs and making informed decisions about evidence selection and context size.
- Ownership of deployed systems through monitoring and incident response, including diagnosing and fixing failures or regressions.
- Strong Python production skills, including FastAPI, asynchronous services, testing, observability, and maintainable interfaces; Neo4j and Cypher or a comparable graph store are strongly preferred.
Nice to have
- Deep graph experience with Cypher, schema evolution, MERGE patterns, embeddings, and live knowledge graph operations.
- Experience with enterprise AI security, including prompt-injection awareness, context-leak prevention, tenant isolation, role-based access, and policy-layer separation.
- Experience with Azure, Azure AI Search, Pinecone, MongoDB Atlas, pgvector, Elasticsearch, or similar search platforms.
- Prior experience operating B2B enterprise SaaS systems.
Culture & Benefits
- Work on evidence, causal pathways, controls, organizational history, and corrective actions in a safety-critical domain.
- Evaluation-first culture focused on measurable, observable, and improvable quality.
- Direct customer impact through products used by safety teams in energy, utilities, infrastructure, construction, and manufacturing.
- Small founding team with high ownership and close collaboration with the CTO, product, and knowledge engineering.
Hiring process
- The role reports to the CTO and involves close collaboration with product and knowledge engineering.
- No agency inquiries; sponsorship is not provided.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →