DevOps Engineer at Knotch focusing on building and scaling infrastructure for AI-driven marketing technology. Collaborate across teams to enhance reliability and performance.
Responsibilities
Design, build, and maintain scalable, secure, and highly available infrastructure across pre-production and production environments
Develop and manage CI/CD pipelines to enable fast, reliable, and repeatable deployments across multiple environments
Own infrastructure as code (IaC) practices using tools like Terraform to ensure consistency and reproducibility
Manage environment lifecycle (development, staging, production), including promotion workflows and configuration management
Partner closely with Engineering, Data, and AI teams to support system performance, reliability, and scalability
Implement and maintain monitoring, logging, and alerting systems to ensure high visibility into system health and performance
Optimize infrastructure for cost, performance, and reliability, especially for compute- and data-intensive AI workloads
Support Kubernetes-based deployments and container orchestration for distributed systems
Contribute to security best practices across infrastructure, including IAM, networking, and application-level protections
Create dashboards and reporting systems to provide visibility into system performance, uptime, and operational metrics
Document architecture, operational processes, and infrastructure decisions to support knowledge sharing and onboarding
Act as a DevOps/SRE partner across teams, helping troubleshoot issues and improve system reliability
Requirements
5+ years of experience in DevOps, Site Reliability Engineering, or Infrastructure Engineering roles within SaaS, PaaS, or cloud-native environments
Prior experience in growth-stage and/or startup environment scaling from $10M to $20M+ ARR with a lean team
Strong experience with Google Cloud Provider (GCP), including IAM, networking, and data services
Hands-on experience with Infrastructure as Code tools such as Terraform
Experience building and maintaining CI/CD pipelines (GitHub Actions, ArgoCD, or similar)
Solid experience with Kubernetes, Docker, and containerized environments
Familiarity with deployment tools such as Helm
Experience with monitoring and observability tools like Prometheus and Grafana
Strong understanding of system reliability, scalability, and performance optimization
Ability to work across multiple systems and priorities in a dynamic environment
Strong documentation and communication skills, with attention to clarity and detail
Supplementary experience supporting AI/ML or data-intensive workloads in production environments (Nice-to-Have)
Familiarity with workflow orchestration or data pipeline tools (Nice-to-Have)
Experience with cost optimization strategies for cloud infrastructure (Nice-to-Have)
Exposure to security frameworks and compliance best practices (Nice-to-Have)
Experience working with distributed or globally deployed systems (Nice-to-Have)
Benefits
Comprehensive medical, dental, and vision insurance eligibility
Maintenance Reliability Engineer optimizing asset reliability, availability, and maintenance costs at Tenaris, a global manufacturer of advanced tubular energy products. Supporting steel mill investments, reliability strategies, and complex maintenance solutions.
AI Security & DevOps Engineer securing AI systems from prototype to production for Casper Studios, an AI services firm. Owning security standards, platform controls, and enterprise client approvals.
Senior DevOps Engineer automating cloud and on - premise infrastructure for Keyfactor’s cryptographic identity security platform. Improving Kubernetes, CI/CD, and Infrastructure as Code delivery.
Student Site Reliability Analyst supporting Sun Life’s financial - services technology reliability and monitoring. Building dashboards, analyzing performance, and helping prevent outages across global 24x7 services.
DevOps co - op interns implementing software fixes, cloud configurations, and automated deployments for Canadian insurer Intact. Roles based in Toronto, Vancouver, and Montréal.
Reliability Engineering Co - Op testing optical switching products and components at Lumentum. Analyzing failures, reliability data, and regulatory qualification results in Ottawa.
Senior DevOps Engineer operating secure AWS infrastructure and CI/CD for SOVRA’s public procurement platform. Improving reliability, observability, compliance, and AI - enabled operations.