Applicant Portal
:
Job Details: Infrastructure Engineer
Full details of the job.
Vacancy Name Infrastructure Engineer Vacancy No VN968 Function Syn - Delivery & Operations Work Location One St Peter's Square, Manchester, M2 3DE Basis Permanent Full Time/Part Time Full time Employment Duration - Hours Per Week 35.00 Drivers Licence Required No Benefits Competitive Salary and Benefits package will be provided to the successful candidate. About the Role
Position Overview
As a Infrastructure Engineer at Synapse360, you will provide advanced first-line operational support across multiple strategic managed service environments. This is not a traditional helpdesk role and an exciting opportunity for someone to progress their career.
You will be expected to independently manage monitoring platforms, triage and resolve incidents, validate backup operations across both cloud and on-premises environments and provide meaningful first-response support for critical P1 and P2 incidents.
You will work within a dedicated five-person team participating in a weekly on-call rota to provide 24×7 P1 incident response coverage. The role demands strong technical foundations across both Microsoft Azure and VMware virtualization platforms along with a disciplined approach to ITSM processes, documentation, and patching support.
What we expect
- Minimum of 3–5 years of experience in enterprise infrastructure support roles, ideally within a managed services or MSP environment
- Proven ability to work with monitoring platforms such as LogicMonitor and Azure Monitor, including alert triage, correlation, and escalation
- Hands-on experience with ServiceNow or equivalent ITSM tooling for ticket lifecycle management
- Solid grounding in Windows Server administration (2016/2019/2022) including basic troubleshooting, event log analysis, and service management
- Familiarity with VMware vSphere (VM console access, snapshot management, basic troubleshooting) and Azure VM operations (start/stop, disk management, basic networking)
- Experience with backup validation - specifically Veeam Backup & Replication job status checking and Azure Backup/Recovery Services Vault monitoring
- Willingness to participate in a 1-in-5 weekly on-call rotation providing 24×7 P1 first-response capability
- Strong communication skills for engaging with stakeholders and internal escalation teams
Ideal Candidate Characteristics
Important Attributes
- Technical Breadth: Comfortable operating across both cloud (Azure) and on-premises (VMware/Dell/HPE) environments. You don’t need to be an expert in all areas, but you must be capable of first triage across the full stack.
- Monitoring Discipline: Able to manage high volumes of alerts from LogicMonitor and Azure Monitor, distinguishing between noise and genuine incidents, and escalating appropriately.
- Process Adherence: Strong commitment to ITIL-aligned processes - tickets must be logged, categorised, prioritised, and updated in ServiceNow without exception.
- Patching Awareness: Able to support T2 engineers during patch windows by assisting with pre-patch checks, post-patch validation, and server reboots across both Azure Update Manager and WSUS/SCCM-managed estates.
- On-Call Readiness: Able to respond to P1 alerts within 15 minutes during on-call periods, perform initial triage, and escalate to T2/T3 as needed.
- Documentation Focus: Committed to maintaining and updating runbooks, knowledge base articles, and operational documentation.
Areas of Responsibility
- Monitoring & Alert Management: Manage and triage alerts from LogicMonitor and Azure Monitor across environments. Correlate alerts, close false positives, and escalate genuine incidents. Maintain monitoring dashboards and ensure alert thresholds remain appropriate.
- ServiceNow Ticket Management: Own the ticket lifecycle for incoming incidents and service requests. Triage, categorise, prioritise, and either resolve or escalate. Maintain SLA compliance for response and update targets. Expected to handle 15–25 tickets per day across environments.
- Backup Validation: Perform daily Veeam Backup & Replication job status checks (~1,290 servers) and Azure Backup status validation. Log failures, attempt basic remediation (re-run failed jobs), and escalate persistent failures to Senior Engineers.
- Patching Support: Assist Senior Engineers during monthly patch cycles. This includes pre-patch checks, server reboots, and post-patch validation across WSUS/SCCM (Windows) and Ansible-managed (Linux) estates. Also assist with Azure Update Manager deployment validation and failed-patch reporting.
- VM Troubleshooting: Perform basic troubleshooting on Azure VMs (connectivity, disk, performance) and vSphere VMs (console access, snapshot management, resource contention). Independently resolve P3/P4 VM issues.
- Windows Server Administration: Basic administration including service restarts, event log analysis, disk space management, user access troubleshooting, and DNS/DHCP validation.
- Incident First Response (On-Call): During on-call periods, act as the first responder for all P1 alerts. Perform initial triage, engage vendor support if needed, communicate status updates, and escalate to Senior Engineers/Infrastructure Lead within 30 minutes if resolution is not achievable.
- Documentation & Knowledge Base:Maintain and update operational runbooks, known-error records, and knowledge base articles. Contribute to process improvement initiatives.
Reports to:
Head of Technical Services
On-Call Commitment
his role participates in a shared five-person weekly on-call rotation. Each engineer is primary on-call for one week in every five.
During on-call periods, you are expected to respond to P1 critical alerts within 15 minutes (24×7), perform initial triage and diagnostics, and escalate to Senior Engineers/Infrastructure Lead if the incident cannot be resolved within 30 minutes.
On-call compensation is provided in line with Synapse360’s standard on-call policy.
Experience
Required Technical Skills & Experience Technology Area Required Skill Level LogicMonitor / Monitoring Tools Proficient - alert management, dashboard navigation, threshold tuning Azure Monitor / Log Analytics Awareness to Proficient - alert review, basic KQL queries, VM metrics ServiceNow / Helpdesk Tool Proficient - incident, change, and request management Windows Server (2016/2019/2022) Awareness to Proficient - event logs, services, disk, DNS/DHCP basics VMware vSphere Awareness - VM console, snapshots, basic resource monitoring Azure VMs Awareness - start/stop, disk management, basic networking Veeam Backup & Replication Awareness - job status review, basic re-run of failed jobs Azure Backup / RSV Awareness - backup status validation, policy review WSUS / SCCM Awareness - patch compliance reporting, server group membership PowerShell Awareness - basic script execution for health checks Networking (TCP/IP, VLANs, DNS) Awareness - basic connectivity troubleshooting ITIL Processes Awareness - incident, change, and problem management fundamentals
Desirable Skills
- Experience with HPE ProLiant iLO remote management (console access, health checks)
- Familiarity with Linux fundamentals (file system navigation, service management, log review)
- Exposure to Azure Arc for hybrid server governance
- Experience with Ansible at an awareness level (understanding playbook execution)
- Familiarity with FortiGate firewall basic monitoring
- Knowledge of Azure Sentinel alert review
Education
Qualifications
Desirable:
- Microsoft AZ-900 (Azure Fundamentals) or AZ-104 (Azure Administrator)
- VMware VCA or VCP Foundations
- Veeam VMTSP (Technical Sales Professional)
Close Date 31 Oct 2026
Please click here to read and review our Privacy Policy.