Applied AI Scientist
We are looking for exceptional quantitative thinkers to build agentic systems that help people make better decisions.
This is not primarily a conventional data science role, an LLM application-development role, or a prompt-engineering role. We are interested in people who want to understand how intelligent systems can reason over complex problems, use tools and models effectively, learn from evidence, and become systematically better over time.
What You Will Do
- Design and build agentic decision-support systems that can reason over data, interact with tools and models, maintain state, decompose complex problems and support consequential human decisions.
- Develop mechanisms for systematic improvement. Design evaluation frameworks, feedback loops, benchmarks, experiments and instrumentation that allow us to determine where a system fails, why it fails and whether a proposed change genuinely improves it.
- Turn “make the agent better” into a quantitative problem. Define measurable objectives and failure modes; build representative test sets; analyze performance distributions rather than anecdotes; and distinguish real improvement from movement on a convenient metric.
- Explore self-improving agentic architectures. This may involve reflection and critique, adaptive tool or model selection, memory, search, planning, generated training/evaluation data, policy improvement, human feedback, simulation or other mechanisms. We are interested in what works - not in allegiance to a particular architecture.
- Bring statistical and machine-learning judgment to agentic systems. Use probabilistic modelling, statistical inference, ML, optimization or causal reasoning when they make the system more useful, more reliable or easier to improve.
- Generate high-quality software using AI coding systems. We expect you to work AI-natively. The objective is not to demonstrate how quickly you can type Python; it is to produce excellent software.
- Own difficult problems end to end. Work from an initially ambiguous business or decision problem through problem formulation, system design, implementation, evaluation and deployment.
- Challenge the problem definition. A technically sophisticated solution to the wrong problem is still the wrong solution. You should be comfortable questioning assumptions and reframing what is being optimized.
- Work directly with senior technical and business stakeholders. Explain complex ideas precisely without hiding behind jargon, and make uncertainty, assumptions and limitations explicit.
What We Are Looking For
Exceptional quantitative foundations
You have rigorous training in a highly quantitative discipline such as mathematics, physics, statistics, engineering, econometrics, operations research or another discipline involving substantial mathematical and quantitative reasoning.
You should be comfortable reasoning mathematically about unfamiliar problems rather than relying solely on methods you have used before.
Experience with agentic systems
You have meaningful hands-on experience designing or developing agentic systems.
We are particularly interested in experience involving some combination of:
- multi-step reasoning and planning
- structured outputs and stateful workflows
- memory and context management
- multi-agent or decomposed-agent architectures
- agent evaluation
- automated critique or refinement
- feedback-driven adaptation
- decision-support applications
- mechanisms intended to improve system performance over repeated iterations.
A master's or PhD is welcome, but neither is a substitute for strong analytical thinking.
The role is product-oriented: you will shape the capabilities, workflows and evaluation systems that make our agentic products more useful, reliable and improvable.
You will turn difficult enterprise problems, user needs and client lessons into reusable product capabilities. Working with technical, product, domain and client-facing teams, you will identify high-leverage problems, formulate improvement hypotheses, and design experiments that distinguish genuine progress from superficial gains.
We expect you to move quickly from ideas to prototypes and measurable evidence without compromising reasoning, quality or technical integrity. The goal is not to build the most elaborate AI system, but products that make better decisions, whose performance we can understand, and which we know how to improve.
You do not need to be a specialist in every branch of machine learning, but you should have enough statistical maturity to reason properly about evidence.
We expect practical familiarity with statistical modelling and machine learning and, more importantly, sound instincts around:
- uncertainty
- validation
- sampling and selection effects
- overfitting
- experiment design
- measurement
- predictive versus causal claims
- distribution shift
- determining whether a system actually works.
What matters is whether you can make those systems produce excellent code.
You should be capable of:
- designing a coherent software architecture before or while implementation emerges;
- decomposing work effectively for coding agents;
- inspecting and interrogating unfamiliar generated code;
- identifying bad abstractions, subtle bugs and unnecessary complexity;
- designing meaningful tests rather than merely achieving test coverage;
- maintaining reproducible environments;
- understanding APIs, data pipelines and production system boundaries;
- debugging systems whose implementation you did not personally type; and
- leaving behind code that another strong engineer can understand, extend and trust.
You are accountable for every line you ship even if you personally typed none of them.