Dayforce is a global human capital management (HCM) company headquartered in Toronto, Ontario, and Minneapolis, Minnesota, with operations across North America, Europe, Middle East, Africa (EMEA), and the Asia Pacific Japan (APJ) region.
Our award-winning Cloud HCM platform offers a unified solution database and continuous calculation engine, driving efficiency, productivity and compliance for the global workforce.
Our brand promise - Makes Work Life Better™ - Reflects our commitment to employees, customers, partners and communities globally.
About the opportunity
We are looking for an AI Engineer - Agentic Systems Evaluation to help define how we measure, test, and improve the quality of enterprise AI systems.
As AI evolves from conversational assistants and Retrieval-Augmented Generation (RAG) into tool using agents, multi-agent systems, and autonomous business workflows, evaluating only the final response is no longer enough.
An agent may reach the right answer while choosing the wrong tool, taking unnecessary steps, retrieving incorrect context, failing to escalate to a human, or violating business rules.
This role will build the evaluation frameworks needed to understand not only whether an AI system succeeded, but how it succeeded, how reliably it can repeat that outcome, and whether the architecture is appropriate for the problem.
What you’ll get to do
Build Agentic Evaluation Frameworks
Design evaluation methodologies covering the complete AI execution lifecycle:
Intent → Planning → Retrieval → Tool Use → Reasoning → Action → Business Outcome
Evaluate systems across dimensions including:
- task and business outcome accuracy
- planning and decision quality
- retrieval quality and groundedness
- tool selection and execution
- agent routing and delegation
- human escalation decisions
- reliability and failure recovery
- safety and policy compliance
- latency, token consumption, and cost
Evaluate Agentic Design Patterns
Design experiments and benchmarks that help engineering teams determine which architecture works best for a given problem.
Evaluate patterns such as:
- Single-agent vs. multi-agent systems
- Supervisor/router architectures
- Planner–executor patterns
- Sequential and parallel workflows
- Tool-using agents
- Human-in-the-loop workflows
- Long-running agents with state and memory
Measure whether additional agent complexity actually improves task success, reliability, and business outcomes enough to justify increased latency, cost, and operational complexity.
Advance RAG & Knowledge Evaluation
Build rigorous evaluation approaches for enterprise retrieval and knowledge systems, including:
- retrieval precision and relevance
- groundedness and citation accuracy
- source authority and freshness
- chunking, metadata, and indexing strategies
- semantic vs. hybrid search and reranking
- permission-aware retrieval
Evaluate how retrieval decisions ultimately impact downstream agent performance, rather than treating RAG evaluation as an isolated problem.
Build Automated Evaluation & Regression Testing
Develop scalable evaluation infrastructure including:
- golden and synthetic datasets
- scenario and adversarial test suites
- deterministic graders
- LLM-as-a-Judge evaluation
- human evaluation workflows
- trace-based evaluation
- automated regression testing
Integrate evaluations into AI development and release pipelines so changes to models, prompts, retrieval, tools, or agent architectures can be measured before reaching production.
Build Agent Trace & Failure Analysis
Analyze complete agent execution traces including planning, retrieved context, tool calls, handoffs, retries, exceptions, latency, and cost.
Develop failure taxonomies that distinguish between:
Model | Retrieval | Planning | Tool | Routing | Memory | Integration | Policy | Orchestration failures
Turn production failures and user feedback into measurable regression tests and engineering improvements.
Define Production AI Quality
Establish measurable quality standards and release criteria for AI systems.
Metrics may include Task Success Rate, First-Pass Success Rate, Tool Selection Accuracy, Agent Routing Accuracy, Plan Execution Fidelity, Failure Recovery Rate, Human Escalation Accuracy, Business Outcome Accuracy, Cost / Latency per Successful Task
Help teams answer a fundamental question:
Is this AI system reliable, safe, efficient, and valuable enough to operate in production?
Skills and experience we value
Strong experience with:
- Python and software engineering
- LLM application development
- AI/ML evaluation and experimentation
- automated testing and data analysis
- RAG, embeddings, hybrid search, and reranking
- tool/function calling and structured outputs
- agent orchestration and multi-agent workflows
- AI observability and tracing
- APIs and enterprise integrations
Experience with platforms or frameworks such as Agents SDK, LangGraph/LangChain, Microsoft AI Foundry, Amazon Bedrock, or similar agent platforms is valuable.
You should be comfortable working with evaluation techniques such as offline/online evals, deterministic graders, model-based graders, human evaluation, synthetic datasets, adversarial testing.
What would make you stand out
You don't stop when an agent successfully completes a task. You ask:
- Did it choose the right approach?
- Did it use the right tools and information?
- Can it succeed consistently?
- Can it recover when something fails?
- Did adding more agents improve the outcome?
- Could a simpler architecture achieve the same result?
- What did the successful outcome cost?
You turn those questions into measurable experiments that help engineering teams build better AI systems.
Why This Role Matters
Enterprise AI is moving from systems that answer questions to systems that make decisions and perform work.
That changes how quality must be measured.
The next generation of AI evaluation must measure:
What the system understood → what it retrieved → what it decided → what it did → and whether the business outcome was correct.
This role will help establish the engineering discipline required to make those systems measurable, reliable, and production-ready.
What’s in it for you
Dayforce is fueled by the diversity of our talented employees. We are an equal opportunity employer and consider and embrace ALL individuals and what makes them unique. We believe our employees should be happy and healthy, with peace of mind and a sense of fulfillment.
We encourage individuals to apply based on their passions.
Dayforce encourages personal and professional growth. We offer excellent time away from work programs, comprehensive wellness initiatives and recognition through competitive pay and benefits.
With a commitment to community impact, including volunteer days and our charity, Dayforce Cares we provide opportunities for you to thrive both in your career and personal life. Our focus is not just on your job but on supporting you to be the best version of yourself.
This job posting is for an existing vacancy
Artificial intelligence may be used in the screening, assessment, or selection of applicants for this position.
About the Salary Ranges
Please note that the salary range mentioned in this job description should serve simply as a guide. The final compensation offered may vary based on a variety of factors, including bonuses and/or incentives, or a candidate’s experience, skills, budget and location. Our company is committed to providing a fair, equitable, and competitive package that reflects the value an individual brings to the organization.
Proficiency in English is required for this position as this role will regularly interact with English-speaking stakeholders, co-workers, managers and/or clients across the world. Further, our back office support teams, including but not limited to Human Resources, are primarily English speaking. Employees need to be able to communicate with these departments in English to appropriately administer their business relationship. Due to the significant high volume of interactions with these English-speaking co-workers, managers, stakeholders and/or clients, which is inherent to this position, it is not possible to reorganize the company's activities to avoid this requirement.
Fraudulent Recruiting
Beware of fraudulent recruiting. Legitimate Dayforce contacts will use an @dayforce.com email address. We do not request money, checks, equipment orders, or sensitive personal data during the recruitment process. If you have been asked for any of the above, or believe you have been contacted by someone posing as a Dayforce employee, please refer to our fraudulent recruiting statement found here: https://www.dayforce.com/be-aware-of-recruiting-fraud
Dayforce actively monitors all job applications to ensure authenticity. Submissions determined to be fraudulent or misleading will be declined from the recruitment process
Pour consulter cette offre d'emploi en français, veuillez utiliser le lien: https://jobs.dayforcehcm.com/fr-CA/mydayforce/alljobs