Responsibilities
Operate and support a large-scale enterprise platform, using monitoring, observability, and alerting tools to maintain reliability. Respond to incidents, contribute to root cause analysis, drive remediation, and collaborate with technical, product, and risk teams.
Requirements
Requires hands-on experience with enterprise platform operations, monitoring and alerting tools, incident response, root cause analysis, and reliability engineering practices. Strong communication and collaboration skills are needed; experience in financial services, automation scripting, cloud platforms, containers, CI/CD, and SLOs, SLIs, and error budgets is an advantage.