AI Engineer at FacilityOS in Toronto
- Company: FacilityOS
- Location: Toronto
- Posted: Sep 23, 2026
- Type: Full-time
- Salary: CA$80K – CA$100K • Offers Bonus
- Experience: 2+ years
Overview
Application About FacilityOS FacilityOS is a fast-growing company redefining how facilities operate—bringing safety, security, and daily operations into one unified platform used by organizations around the world. As we continue to scale globally, we’re building a team of driven, curious people who …
Job description
- Application
- About FacilityOS
- FacilityOS is a fast-growing company redefining how facilities operate—bringing safety, security, and daily operations into one unified platform used by organizations around the world.
- As we continue to scale globally, we’re building a team of driven, curious people who want to make an impact. You’ll be part of a dynamic, collaborative culture where individuals are trusted to take ownership, solve meaningful problems, and grow in their careers. Our team comes together in-office twice a week to connect, collaborate, and build momentum.
- If you’re looking to do your best work alongside a great team in a high-growth environment, FacilityOS is the place to build your career.
- The Cloud Engineer reports to the Director of Cloud Engineering and joins the CloudOps team responsible for the infrastructure behind FacilityOS’s SaaS platform. This is a generalist role spanning the full stack of cloud operations: infrastructure-as-code, CI/CD, observability, platform reliability, and FinOps across a large serverless Azure environment. You’ll move between provisioning infrastructure, tuning monitoring, respond to alerts to ensure high availability, improving deployment pipelines, , and helping harden the team’s compliance posture.
Responsibilities
- Infrastructure as Code & Deployment
- Provision, configure, and maintain secure, reliable cloud infrastructure using Terraform across multiple environments in segregated data residency/processing regions
- Contribute reusable, versioned, and well-documented components to our internal IaC module registry, improving consistency and accelerating delivery across teams
- Help move manually configured “ClickOps” resources under IaC management by inventorying existing infrastructure, codifying configurations, validating state, and planning low-risk migrations
- Contribute to CI/CD workflows in Azure DevOps to make infrastructure and application deployments secure, repeatable, observable, and efficient
- Security & Compliance
- Continuously assess and improve our infrastructure security posture, identifying threats, vulnerabilities, misconfigurations, identity and access risks, and policy violations before they reach production
- Embed security controls and policy checks into IaC and CI/CD workflows, remediate findings and reduce recurring risk
- Ensure infrastructure remains compliant with SOC 2, ISO 27001, HIPAA, and other applicable security and regulatory requirements
- Support our identity and access program by using and enforcing managed identities, workload identity federation and minimize secrets.
- Site Reliability, Production Support & Incident Response
- Provide hands-on production support for cloud infrastructure and services, troubleshooting complex issues across application, platform, network, identity, and data layers
- Respond to alerts and participate in the incident response process to quickly mitigate customer impact, restore service, and maintain uptime SLAs
- Contribute to issues beyond short-term recovery by investigating contributing factors, identifying root causes, and implementing durable corrective and preventative actions
- Build automations and self-healing capabilities that eliminate repetitive operational work, reduce human error, accelerate recovery, and prevent recurring incidents
- Maintain and expand observability coverage across all services using Datadog dashboards, monitors, SLOs, logs, traces, and other actionable telemetry
- Regularly review alert quality and service coverage to close monitoring gaps, reduce noise, and ensure alerts are tied to meaningful customer and system impact
- Contribute to post-incident reviews and follow-up, turning lessons learned into measurable improvements to infrastructure, monitoring, runbooks, architecture, and operational processes
- FinOps
- Monitor cloud infrastructure costs, investigate anomalies, and provide visibility into spend, usage, and cost drivers
- Identify and implement opportunities to reduce waste, right-size resources, improve utilization, and keep infrastructure spending within budget without compromising reliability or security
- Partner with engineering teams to encourage cost-aware architecture and establish practical ownership of cloud spend
- Developer Enablement & Collaboration
- Partner with software development teams to guide cloud onboarding, infrastructure implementation, deployment patterns, and architecture decisions
- Provide practical guidance on cloud, security, reliability, observability, and cost-management best practices throughout the development lifecycle
- Document infrastructure changes, architectural decisions, runbooks, standards, and operational procedures so teams can work safely and independently
Requirements
- 2+ years of experience in a Cloud engineering, DevOps, or SRE role
- Deep practical experience and complete fluency with Azure, particularly App Services, Container Apps, Front Door, SQL databases, Cosmos, Service Bus, Key vaults
- Experience with a CI/CD platform (Azure DevOps preferred)
- Experience with observability/monitoring tools (Datadog preferred)
- Experience implementing cloud governance controls from the SOC2 and ISO 27001 frameworks
- Proven production support experience, with strong troubleshooting instincts and a methodical approach to diagnosing distributed systems, identifying root causes, and delivering lasting remediation
- Strong scripting and automation skills, with experience turning recurring operational problems into reliable automated solutions
- Experience in a SaaS or enterprise technology environment preferred
- Azure certifications are appreciated
Skills
Required
- Terraform
- CI/CD (Azure DevOps preferred)
- SOC2 and ISO 27001 frameworks
- Scripting and automation
- Troubleshooting distributed systems
Preferred
- Azure certifications
Benefits
- We work hard and play hard and we do both with passion and respect for one another. Our company promotes a fast-paced, fun, friendly, and highly collaborative work environment that provides:
- 🩺Comprehensive health coverage
- 🏠A Hybrid work environment
- 💡Opportunity for advancement and growth
- 🍕 Catered Events, Snacks, Drinks – You won’t go Hungry!
- 🥳 Birthday and Life Celebrations
- 🎉 Two annual parties in a year