About FacilityOS
FacilityOS is a fast-growing company redefining how facilities operate - bringing safety, security, and daily operations into one unified platform used by organizations around the world.
As we continue to scale globally, we’re building a team of driven, curious people who want to make an impact. You’ll be part of a dynamic, collaborative culture where individuals are trusted to take ownership, solve meaningful problems, and grow in their careers. Our team comes together in-office twice a week to connect, collaborate, and build momentum.
If you’re looking to do your best work alongside a great team in a high-growth environment, FacilityOS is the place to build your career.
About the Role
The Cloud Engineer reports to the Director of Cloud Engineering and joins the CloudOps team responsible for the infrastructure behind FacilityOS’s SaaS platform. This is a generalist role spanning the full stack of cloud operations: infrastructure-as-code, CI/CD, observability, platform reliability, and FinOps across a large serverless Azure environment. You’ll move between provisioning infrastructure, tuning monitoring, respond to alerts to ensure high availability, improving deployment pipelines, , and helping harden the team’s compliance posture.
Key Responsibilities
Infrastructure as Code & Deployment
Provision, configure, and maintain secure, reliable cloud infrastructure using Terraform across multiple environments in segregated data residency/processing regions
Contribute reusable, versioned, and well-documented components to our internal IaC module registry, improving consistency and accelerating delivery across teams
Help move manually configured “ClickOps” resources under IaC management by inventorying existing infrastructure, codifying configurations, validating state, and planning low-risk migrations
Contribute to CI/CD workflows in Azure DevOps to make infrastructure and application deployments secure, repeatable, observable, and efficient
Security & Compliance
Continuously assess and improve our infrastructure security posture, identifying threats, vulnerabilities, misconfigurations, identity and access risks, and policy violations before they reach production
Embed security controls and policy checks into IaC and CI/CD workflows, remediate findings and reduce recurring risk
Ensure infrastructure remains compliant with SOC 2, ISO 27001, HIPAA, and other applicable security and regulatory requirements
Support our identity and access program by using and enforcing managed identities, workload identity federation and minimize secrets.
Site Reliability, Production Support & Incident Response
Provide hands-on production support for cloud infrastructure and services, troubleshooting complex issues across application, platform, network, identity, and data layers
Respond to alerts and participate in the incident response process to quickly mitigate customer impact, restore service, and maintain uptime SLAs
Contribute to issues beyond short-term recovery by investigating contributing factors, identifying root causes, and implementing durable corrective and preventative actions
Build automations and self-healing capabilities that eliminate repetitive operational work, reduce human error, accelerate recovery, and prevent recurring incidents
Maintain and expand observability coverage across all services using Datadog dashboards, monitors, SLOs, logs, traces, and other actionable telemetry
Regularly review alert quality and service coverage to close monitoring gaps, reduce noise, and ensure alerts are tied to meaningful customer and system impact
Contribute to post-incident reviews and follow-up, turning lessons learned into measurable improvements to infrastructure, monitoring, runbooks, architecture, and operational processes
FinOps
Monitor cloud infrastructure costs, investigate anomalies, and provide visibility into spend, usage, and cost drivers
Identify and implement opportunities to reduce waste, right-size resources, improve utilization, and keep infrastructure spending within budget without compromising reliability or security
Partner with engineering teams to encourage cost-aware architecture and establish practical ownership of cloud spend
Developer Enablement & Collaboration
Partner with software development teams to guide cloud onboarding, infrastructure implementation, deployment patterns, and architecture decisions
Provide practical guidance on cloud, security, reliability, observability, and cost-management best practices throughout the development lifecycle
Document infrastructure changes, architectural decisions, runbooks, standards, and operational procedures so teams can work safely and independently
Qualifications
2+ years of experience in a Cloud engineering, DevOps, or SRE role
Deep practical experience and complete fluency with Azure, particularly App Services, Container Apps, Front Door, SQL databases, Cosmos, Service Bus, Key vaults
Experience with a CI/CD platform (Azure DevOps preferred)
Experience with observability/monitoring tools (Datadog preferred)
Experience implementing cloud governance controls from the SOC2 and ISO 27001 frameworks
Proven production support experience, with strong troubleshooting instincts and a methodical approach to diagnosing distributed systems, identifying root causes, and delivering lasting remediation
Strong scripting and automation skills, with experience turning recurring operational problems into reliable automated solutions
Experience in a SaaS or enterprise technology environment preferred
Azure certifications are appreciated
Why work at FacilityOS?
We work hard and play hard and we do both with passion and respect for one another. Our company promotes a fast-paced, fun, friendly, and highly collaborative work environment that provides:
🩺Comprehensive health coverage
🏠A Hybrid work environment
💡Opportunity for advancement and growth
🍕 Catered Events, Snacks, Drinks – You won’t go Hungry!
🥳 Birthday and Life Celebrations
🎉 Two annual parties in a year
FacilityOS Commitment
We believe that a diverse team is the key to innovation and growth. We are an equal opportunity employer that value diversity at our company and encourage all candidates to apply. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.
FacilityOS will accommodate individuals with disabilities through each stage of the recruitment and selection process. Please advise us of any needs when your interview is booked, and we will do our best to meet your needs.
Please note that all candidates must be legally eligible to work in the location in which you applied.
Background and Reference Checks
Any offer of employment may be conditional upon full background checks including a criminal record check, a credit check and employment and educational verifications. A reference check will also be conducted.
We may use artificial intelligence (AI) tools to support parts of the hiring process such as reviewing applications. These tools assist the recruitment team but do not replace human judgment. All advancement and hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us or click here.
FacilityOS thanks all candidates for their interest, however only those selected to continue in the process will be contacted.