Job Title: AI Platform Engineering Specialist
Location: Montreal (Day 1 onboarding onsite/in office presence 3x/week)
Key responsibilities
- Design, build and operate the AI Gateway's Azure and AWS deployments, taking them from proof of concept to production.
- Develop and extend the Python services (FastAPI / Flask) that provide the Gateway's inference, onboarding and administrative APIs.
- Integrate new model providers and model families, including Azure AI Foundry / Azure OpenAI and AWS Bedrock, covering request signing, streaming responses, failover and quota handling.
- Implement cloud-native authentication and secrets handling - Entra ID with Managed Identity and workload federation, AWS IAM roles and STS - with the goal of eliminating stored credentials.
- Build and evolve the entitlement and authorization data layer across SQL Server and PostgreSQL, including schema changes, migrations and data-correctness controls.
- Own the platform controls that make the Gateway a governance point: rate limiting, token accounting, content guardrails, audit logging and chargeback reporting.
- Deploy and run the service on Kubernetes (on-premises, AKS and EKS) using Helm, GitOps and Terraform, and keep the CI/CD pipelines (Jenkins, GitHub Actions) healthy.
- Build the observability to answer any question about a request after the fact - metrics, logs and dashboards across Prometheus, Grafana, Loki and Snowflake.
- Work with cloud platform, network and security teams on connectivity, egress policy, network controls and architecture review, and produce the evidence those reviews require.
- Support production: participate in on-call, investigate incidents, and drive fixes and hardening back into the code.
- Write tests and documentation as part of delivery, and review peers' changes.
Required qualifications:
- Strong, production-grade Python, including a web framework - FastAPI or Flask - and a real testing discipline.
- Hands-on Kubernetes: deploying, configuring and troubleshooting workloads, not solely reading manifests.
- Practical OIDC / OAuth 2.0: token validation, JWKS, client-credentials flows, claim and audience handling.
- Microsoft Azure, hands-on across at least three of: AKS, Entra ID (app registrations, service principals, Managed Identity / Workload Identity), Azure OpenAI or Azure AI Foundry, Key Vault, Azure Database for PostgreSQL, Azure Cache for Redis, Azure Monitor.
- Amazon Web Services, hands-on across at least three of: IAM and STS / assume-role, SigV4 request signing, Bedrock, EKS, VPC endpoints and private networking, Secrets Manager, CloudWatch.
- Infrastructure as code - Terraform, Bicep or CDK - and CI/CD with Jenkins or GitHub Actions.
- SQL and relational data modelling, including schema migrations.
- Clear written and verbal communication, and the ability to work directly with security, network and platform teams.
Preferred qualifications
- Experience building or operating an API gateway, reverse proxy or multi-tenant platform.
- LLM platform engineering specifics: streaming and server-sent events, token accounting, prompt and response guardrails, model evaluation.
- Kafka and Snowflake for audit and consumption data pipelines.
- Observability depth: Prometheus and PromQL, Grafana, Loki, OpenTelemetry.
- Redis or Valkey beyond basic caching - counters, TTLs, distributed rate-limiter semantics.
- Experience delivering in a regulated enterprise environment with corporate proxies, private networking and strict change control.