Position: Systems Reliability Engineer (SRE) โ Embedded Finance
Locations: Jacksonville, FL; Berkeley Heights, NJ; Alpharetta, GA; Toronto, ON
Role Overview & Key Responsibilities:
- Own the overall reliability, resiliency, and availability of an enterprise-scale Embedded Finance (EmFi) platform.
- Design, implement, and maintain end-to-end monitoring and alerting frameworks using Splunk, Dynatrace, Grafana, and Datadog.
- Define, track, and optimize platform health metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
- Serve as the primary technical driver during incident response and high-severity outage remediation calls.
- Lead the Root Cause Analysis (RCA) process - investigating incidents, documenting event sequences, and driving corrective actions to prevent recurrence.
- Manage incident ticket lifecycles, tracking, aging, and reporting using ServiceNow.
- Develop automation scripts and tools to reduce operational toil, MTTD, and MTTR.
Qualifications & Required Skills:
- Proven experience as a Systems Reliability Engineer (SRE) supporting large-scale enterprise platforms.
- Hands-on proficiency with monitoring and observability tools: Splunk, Dynatrace, Grafana, and Datadog.
- Demonstrated experience leading incident response calls and executing Root Cause Analysis (RCA).
- Strong understanding of reliability engineering concepts (resiliency, fault tolerance, monitoring best practices).
- Experience with ServiceNow or equivalent enterprise incident management ticketing systems.
- Strong written and verbal technical communication skills.
Nice to Have:
- Background in Financial Services, Payments, or Embedded Finance.
- Scripting proficiency in Python, Go, or Bash for operational automation.
- Knowledge of cloud platforms, containers, and CI/CD deployment pipelines.