At a glance
Automatically prepared from the listing. Check the original description for the full requirements.
Jump to the original descriptionResponsibilities
The role involves implementing monitoring systems and responding to incidents to minimize downtime and restore services. Responsibilities also include automating repetitive tasks, managing cloud infrastructure, and collaborating with developers to ensure system reliability.
Requirements
Candidates need over 8 years of experience in production support with expert knowledge of OCP and Windows environments. Proficiency in scripting languages, container services like Kubernetes, and database management is required.
Working hours
40 hours per week
Skills
- OCP
- Windows Environment
- Dynatrace
- Linux
- Shell Scripting
- Python
- Power Shell
- SQL
- MongoDb
- Kubernetes
- Disaster Recovery
- Production Support
- Incident Response
- Automation
- Infrastructure Management
- Release Engineering
Visa sponsorship
Not detected in the job text
Categories
- Technology
- Software
- Engineering
- Consulting
- Customer Service & Support
Keywords
- OCP
- OpenShift
- Windows
- Linux
- Dynatrace
- Shell Scripting
- Python
- Power Shell
- SQL
- MongoDb
- Kubernetes
- Disaster Recovery
- Production Support
- Incident Response
- Automation
- Infrastructure Management
- Capacity Planning
- Release Engineering
- Cloud Platform
- IT Services
- IT Consulting
- Monitoring
- Alerting
- System Reliability
- Availability
Original job description
We are looking for a Production support Engineer to work for the Commercial Line of Business.
Monitoring And Alerting
Below is the detailed Job Description:
Implement and maintain monitoring systems to proactively identify potential issues and alert engineers to problems before they impact users.
Incident Response
Respond to incidents and outages, diagnose problems, and implement solutions to minimize downtime and restore service.
Automation
Automate repetitive tasks and processes to improve efficiency and reduce manual effort.
Infrastructure Management
Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources.
Capacity Planning
Plan for future capacity needs to ensure systems can handle anticipated workloads.
Release Engineering
Develop and maintain processes for deploying software updates and releases.
Collaboration
Work closely with developers, operations teams, and other stakeholders to ensure system reliability and availability.
Documentation
Maintain clear and concise documentation of systems, processes, and procedures.
Continuous Improvement
Identify areas for improvement and implement changes to enhance system reliability and performance.
Skills And Qualifications
Cloud Platform (OCP)
8+ Years experience in production support handling Prod incidents.
Excellent knowledge of OCP and windows environment.
Monitoring tools ( Dynatrace )
Operating System (Windows, Linux)
Scripting (Shell Scripting, Python, Power Shell)
Database (SQL database management, MongoDb)
Container Services (Kubernetes)
Disaster Recovery Planning and execution