- Experienced Application Manager (Site Reliability Engineer) who is responsible for day-to-day reliability
- Implement and operationalize Service Level Availability (SLAs) and related reliability metrics
- 6+ years of experience in SRE, Production Engineering, Platform Engineering, Application Support, or Technology Service Delivery.
Experience with observability, automation, Python, and PowerShell.
- Strong automation and scripting capabilities.
- Experience with PowerShell and other automation tools.
- Python
- GitHub Actions
- Strong technical knowledge of cloud platforms, enterprise systems, and application, data, and platform architectures.
- Capital Markets
- Total Fund Market Investment
- Proven experience managing high-availability production environments, incident response, troubleshooting, and root cause analysis.
- AI-powered engineering tools, including IDE assistants and MCP-enabled integrations, to support troubleshooting and service restoration.
- Experience working with APIs, messaging/queueing platforms, distributed systems, CI/CD, Infrastructure as Code, DevOps, and DataOps.
- Working knowledge of Agile, Waterfall, DevOps, ITIL, and COBIT.
- Experience with Jira, Confluence, and Git.
- Excellent communication and collaboration skills, with the ability to partner effectively with architects, business analysts, DBAs, QA, and cross-functional technology teams.
Contract/Hybrid
Toronto