DESCRIPTION
Company Profile
At Brokerage, we advise, originate, trade, manage, and distribute capital for governments, institutions, and individuals, always operating with a standard of excellence. We are a leading global financial services firm conducting business through three principal business segments: Institutional Securities, Wealth Management, and Investment Management. The Firm’s employees serve clients worldwide from more than 1,200 offices across 43 countries.
As a market leader, the talent and passion of our people are critical to our success. We share a common set of values rooted in integrity, excellence, and a strong team ethic. Brokerage provides a strong foundation for building a professional career, offering people opportunities to learn, achieve, and grow. A philosophy that respects different lifestyles, perspectives, and individual needs is an important part of our culture.
Department Profile
Operations Technology, or OpsTech, is an innovation-driven organization comprising functionally aligned engineering teams that support critical Firm operations. We develop strategic technology solutions for key business processes, including settlements, confirmations, regulatory reporting, position keeping, and client reference data management.
If you are interested in solving complex technical problems, improving developer productivity, and building reliable engineering solutions in a dynamic global environment, OpsTech is the place for you.
Position Description
We are looking for an experienced DevOps, Kubernetes, and Site Reliability Engineer with a minimum of five years of relevant industry experience. Experience within financial services or another highly regulated technology environment is preferred.
The successful candidate will join the Operations Technology team and help design, automate, deploy, and support reliable application platforms and CI/CD capabilities. The role will work closely with application development, infrastructure, network, security, and production support teams to improve software delivery, platform stability, operational resilience, and engineering efficiency.
The position requires strong hands-on experience with Kubernetes, Docker/Podman, Linux, networking, CI/CD pipelines, GitHub, Python, and shell scripting. The candidate should understand modern DevOps and SRE practices and be comfortable supporting production systems, troubleshooting complex application and infrastructure issues, and automating repetitive operational tasks.
The ideal candidate will be passionate about automation, production reliability, and continuous improvement. The candidate should be organized, disciplined, detail-oriented, self-motivated, collaborative, and focused on delivering measurable engineering outcomes.
QUALIFICATIONS
Responsibilities
- Design, build, maintain, and enhance CI/CD pipelines and supporting build, test, release, and deployment infrastructure.
- Develop and maintain automated deployment solutions for applications running in Kubernetes and containerized environments.
- Partner with application development teams to improve build, test, deployment, and release processes.
- Support the deployment of applications, configuration changes, patches, and platform upgrades across development, testing, and production environments.
- Build and maintain Kubernetes deployment artifacts, including YAML configuration, Helm charts, or equivalent packaging and configuration mechanisms.
- Develop automation using Python, Linux shell scripting, and related tools.
- Manage and improve GitHub repositories, branching strategies, pull-request workflows, access controls, and automated repository processes.
- Integrate automated testing, code-quality validation, security scanning, dependency checks, and linting into CI/CD pipelines.
- Troubleshoot application, infrastructure, container, Kubernetes, network, and deployment issues in complex environments.
- Investigate production incidents, identify root causes, and implement preventive or corrective engineering solutions.
- Apply SRE principles to improve system reliability, availability, scalability, observability, and operational readiness.
- Define and improve monitoring, alerting, dashboards, operational metrics, and production support procedures.
- Automate routine operational activities to reduce manual effort and operational risk.
- Collaborate with infrastructure, network, database, cybersecurity, release management, and application teams.
- Create and maintain technical documentation, operational runbooks, deployment procedures, and troubleshooting guides.
- Participate in design reviews, production-readiness reviews, incident reviews, and continuous-improvement initiatives.
- Participate in an after-hours production support and on-call rotation when required.
Required Skills
- Minimum of 5 years of relevant experience in DevOps, SRE, production engineering, platform engineering, infrastructure engineering, or a related discipline.
- Strong hands-on experience with Kubernetes, including application deployment, configuration, troubleshooting, scaling, services, ingress, secrets, and operational support.
- Strong hands-on experience with Docker/Podman and containerized application environments.
- Strong Linux and UNIX system administration and troubleshooting skills.
- Strong experience with Linux shell scripting, such as Bash or KornShell.
- Hands-on programming and automation experience using Python or a comparable language.
- Strong understanding of CI/CD concepts, software delivery lifecycles, release automation, and deployment strategies.
- Hands-on experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms.
- Hands-on experience with Git and GitHub, including repository management, pull requests, branching strategies, release workflows, and automated checks.
- Experience developing automation using YAML, Ansible, or an equivalent automation framework.
- Strong understanding of networking concepts, including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, load balancing, and network troubleshooting.
- Experience supporting applications and resolving production issues in a fast-paced environment.
- Understanding of SRE practices, including monitoring, incident response, root-cause analysis, service reliability, operational readiness, and continuous improvement.
- Experience integrating code-quality tools, security scanning, automated testing, and policy controls into CI/CD pipelines.
- Ability to troubleshoot issues across application, operating system, container, network, infrastructure, and database layers.
- Experience working in an Agile development environment and using tools such as Jira.
- Strong written and verbal communication skills.
- Ability to collaborate effectively with globally distributed engineering and support teams.
- Ability to prioritize work, manage multiple tasks, and deliver results with limited supervision.
- Ability to participate in an after-hours on-call support rotation.
Desired Skills
- Experience with enterprise Kubernetes platforms such as OpenShift or another managed Kubernetes environment.
- Experience provisioning on-demand environments using virtual machines and containers.
- Experience working with Azure, AWS, GCP, or another cloud platform.
- Experience with infrastructure-as-code and configuration-management technologies.
- Experience with Kubernetes package-management and deployment tools such as Helm.
- Experience with GitOps deployment models and tools.
- Knowledge of service mesh technologies, container networking, ingress controllers, and API gateways.
- Experience with observability and telemetry platforms, including metrics, logs, traces, dashboards, and alerting.
- Experience developing or maintaining Grafana dashboards.
- Familiarity with production incident management, problem management, change management, service operations, and release management.
- Understanding of high availability, disaster recovery, capacity management, and production resiliency.
- Experience with relational database technologies such as DB2, Sybase, or Oracle.
- Experience working in financial services or another regulated enterprise environment.
- Familiarity with secure software supply-chain practices, secrets management, certificate management, and vulnerability remediation.
- Education: Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical industry experience.
Candidate Profile
The successful candidate will be:
- Passionate about automation, reliability, and engineering excellence.
- Comfortable owning technical problems from initial investigation through resolution.
- Able to balance delivery speed with production stability, security, and operational risk.
- Proactive in identifying opportunities to simplify processes and remove manual work.
- Capable of communicating complex technical issues clearly to both technical and non-technical stakeholders.
- A collaborative team member who shares knowledge and contributes to the continuous improvement of engineering practices.