Applied AI Developer (Agent Evaluation)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Applied AI Developer (Agent Evaluation) (Python/LLM): Building and evaluating multi-agent AI systems for automated test generation, execution, and developer workflow optimization with an accent on agentic workflows, benchmark design, and MCP-based tooling. Focus on measuring agent accuracy, context fidelity, latency, test quality, and end-to-end task completion across integrated development environments and testing pipelines.
Location: Toronto, Ontario, Canada
Salary: CAD 99,000–145,200 base salary annually, with potential bonuses, stock grants, and benefits.
Company
Develops software that helps innovators design and manufacture buildings, vehicles, factories, and entertainment products.
What you will do
- Develop and orchestrate multi-agent AI systems for automated test generation, execution, and development workflow optimization.
- Design agentic workflows covering UI, API, integration, and system-level testing.
- Build evaluation frameworks and benchmark suites to compare AI agents with commercial and domain-specific solutions.
- Develop MCP-based tooling and integrate agent workflows with IDEs and compatible developer services.
- Measure and optimize latency, accuracy, context fidelity, and end-to-end task completion.
Requirements
- Bachelor’s or master’s degree in computer science, machine learning, or a related applied AI field.
- Expertise in Python and machine learning frameworks including PyTorch, Transformers, and scikit-learn.
- Experience applying large language models to software understanding, test generation, or agentic AI systems.
- Knowledge of AI evaluation methodologies, agentic task-completion metrics, and test-quality assessment.
- Strong foundations in statistical analysis, experimental design, and developer productivity measurement.
Nice to have
- Software engineering or QA experience with close collaboration with development teams.
- Experience with Playwright, Selenium, Pytest, Appium, and CI/CD pipelines.
- Hands-on experience building, evaluating, and optimizing MCP servers and tool integrations.
- Experience with LangSmith, Langfuse, AgentBench, RAGAS, DeepEval, or similar AI evaluation and observability tools.
- Experience with Azure AI Foundry/ML or AWS cloud ML platforms.
Culture & Benefits
- Work on AI and generative AI solutions focused on improving developer productivity and experience.
- Collaborate with AI engineers, software architects, and product engineering teams.
- Compensation may include annual cash bonuses, stock grants, and a comprehensive benefits package.
- Inclusive culture focused on belonging and meaningful work.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →