Position: DevOps Platform Engineer | AI & LLMOps Infrastructure
Location: Onsite- Toronto, ON
Job type: Full-Time/Permanent Hiring
As an Agent & DevOps Platform Engineer, you will design, scale, and secure the foundational infrastructure that powers our next-generation developer platforms and autonomous AI/LLM agent frameworks. This is a highly specialized role bridging the gap between advanced cloud-native Platform Engineering and cutting-edge LLMOps. You will not just build standalone AI applications - you will build the highly automated, multi-account AWS environments, robust CI/CD pipelines, and secure IAM patterns that allow autonomous agents and developers to ship code reliably, securely, and at enterprise scale.
Key Responsibilities
- Infrastructure as Code & Automation: Architect, deploy, and manage highly reliable multi-account AWS environments using Terraform, CloudFormation, or AWS CDK.
- Agent & Tooling Orchestration: Build, deploy, and scale enterprise developer platforms, integrating agentic orchestration frameworks (e.g., LangChain, LlamaIndex) and custom toolchains into production pipelines.
- Bulletproof Identity & Access: Implement advanced AWS IAM architectures under the principle of least privilege, designing secure cross-account access, service-to-service authentication, and Kubernetes service accounts.
- Container Platform Engineering: Package and orchestrate distributed container systems using Docker and enterprise container platforms (Amazon EKS, ECS, or Kubernetes), configuring robust networking, ingress, and workload security.
- LLMOps & Platform Observability: Build the guardrails for production AI workloads, implementing systems for prompt management, tool execution, model evaluation, observability (CloudWatch), and cost-reliability metrics.
- Scale CI/CD & DevSecOps: Treat security as a core software design constraint. Author high-quality reusable pipeline patterns, build automation workflows, and enable decentralized DevOps practices across multiple business units.
- System Ownership & Culture: Act as a technical mentor whose coaching is reflected in clean working code, robust pipelines, and reusable modules. Define and solve core platform bottlenecks in ambiguous environments.
Required Qualifications
- Experience: 3–5+ years of hands-on platform engineering, cloud architecture, and CI/CD operations within complex enterprise environments.
- Programming: Strong proficiency in Python for infrastructure tooling, API integrations, agent services, and advanced production scripting.
- Technical Ecosystems: Working competence in at least one additional development ecosystem (e.g., Java, .NET, Node.js, or Groovy).
- Cloud Native & Containers: Deep production experience with AWS cloud-native deployments and container orchestration via Docker and Kubernetes / Amazon EKS / ECS.
- AWS Security Mastery: Comprehensive knowledge of complex AWS IAM structures, including roles, policies, trust relationships, and secure machine-to-machine authentication.
- AI/LLMOps Engineering: Direct experience building, deploying, or scaling AI/LLM-based tools or agents in or near production. Solid understanding of prompt management, tool usage, and agent testing.
- Methodology: Strong grasp of Git-based workflows, release management at scale, and an uncompromising DevSecOps approach to system design.
Nice-to-Have Skills
- Hands-on usage of advanced AWS ecosystem tools: ECR, Lambda, Bedrock, CloudWatch, S3, Secrets Manager, Systems Manager, and VPC networking.
- Familiarity with Kubernetes internals: Helm, service accounts, ingress controllers, and network policies.
- Prior experience working with framework tools like LangChain or LlamaIndex.
- Experience driving agile delivery models using Scrum, Kanban, or SAFe across multiple teams.