Job Summary
We are looking for a DevOps Engineer to build, automate, and operate the cloud infrastructure and software delivery platforms that enable engineering teams to develop, deploy, and run secure, reliable, and scalable applications.
The ideal candidate will have strong experience with AWS, Infrastructure as Code, CI/CD, automation, containers, cloud operations, and DevSecOps. You will play a key role in improving deployment efficiency, system reliability, observability, security, and operational excellence across our technology environment.
You will also have the opportunity to support AI-enabled applications and platforms, including AI governance, model monitoring, and agent/LLM observability.
Key Responsibilities
- Design, build, and maintain AWS cloud infrastructure, CI/CD pipelines, and platform automation.
- Automate application deployments, environment provisioning, infrastructure management, monitoring, and operational processes.
- Develop and maintain Infrastructure as Code (IaC) using tools such as Terraform or similar technologies.
- Build reliable and scalable deployment pipelines that support modern software development practices.
- Support production environments, incident management, troubleshooting, and root cause analysis.
- Implement reliability, availability, resiliency, and disaster recovery best practices.
- Implement DevSecOps practices, security controls, compliance automation, and infrastructure security.
- Work closely with software engineers and architects to ensure applications are production-ready.
- Establish monitoring, logging, alerting, and observability capabilities across applications and infrastructure.
- Develop reusable DevOps and platform capabilities that improve developer productivity and operational consistency.
- Automate repetitive operational tasks and continuously improve engineering workflows.
- Participate in production readiness reviews and help establish operational standards and best practices.
- Support governance, auditability, and compliance requirements through automation and policy-based controls.
- Help support infrastructure and operational requirements for AI-enabled applications, agents, and LLM-based solutions.
Required Qualifications
- 10+ years of experience in DevOps, Cloud Engineering, Platform Engineering, Site Reliability Engineering (SRE), or a related field.
- Strong hands-on experience with AWS cloud services and infrastructure.
- Experience with Infrastructure as Code (IaC), preferably Terraform.
- Strong understanding of CI/CD pipelines, DevOps practices, automation, and cloud infrastructure.
- Experience with containers and container orchestration technologies such as Docker and Kubernetes.
- Experience working with APIs, distributed systems, and modern software delivery practices.
- Strong scripting and automation skills using Python, Bash, PowerShell, or similar languages.
- Understanding of cloud security, networking, IAM, monitoring, and production operations.
- Strong troubleshooting, problem-solving, and incident management skills.
- Ability to collaborate effectively with software engineers, architects, security teams, and other technical stakeholders.
Preferred Qualifications
- Experience with observability and monitoring platforms such as:
- AWS CloudWatch
- OpenTelemetry
- Datadog
- Grafana
- Prometheus
- Splunk
- Experience with governance automation, policy-as-code, auditability, and compliance controls.
- Experience implementing DevSecOps and security telemetry.
- Experience working in regulated or compliance-focused environments.
- Experience with security automation and cloud security best practices.
- Familiarity with AI governance, model monitoring, AI infrastructure, or agent/LLM observability.
- Experience supporting production AI/ML or AI-enabled applications.
Key Skills
AWS | DevOps | Cloud Engineering | Platform Engineering | CI/CD | Terraform | Infrastructure as Code | Docker | Kubernetes | Python | Bash | PowerShell | DevSecOps | Cloud Security | Observability | Monitoring | OpenTelemetry | CloudWatch | Grafana | Prometheus | Splunk | SRE | Automation | Infrastructure Automation | AI/LLM Observability
Pay: $70.00-$100.00 per hour
Application question(s):
- How many years of hands-on AWS experience do you have?
- Do you have experience implementing cloud security or DevSecOps practices?
- Do you have automated security guardrails or checks built into your CI/CD pipelines for AI apps?
- Are you currently monitoring performance, token usage, and costs for your production AI models?
- Do you have hands-on experience with centralized observability or Generative AI integrations?
- Do you have hands-on experience building centralized observability solutions for enterprise environments?
Work Location: Hybrid remote in Markham, ON L3R 9Z7