Reliability Engineering Architect at Carbon60 managing a team to deliver AWS cloud solutions. Focus on mentoring engineers and integrating AI tools into automated systems.
Responsibilities
Directly manage, coach, and support a pod of 6+ team members
Design and deploy robust AWS cloud architectures
Champion the use of Infrastructure as Code (IaC)
Integrate AI tools and methodologies into day-to-day workflow
Understand and manage team workflows around billable hours and utilization
Participate in on-call rotation to ensure high availability
Lead AWS Well-Architected Reviews with client engineering teams
Requirements
At least one active AWS Professional or Specialty Certification (e.g., Solutions Architect, DevOps Engineer, etc.)
Proven experience managing a team of at least 6 engineers
Deep, hands-on experience with AWS services and Infrastructure as Code (specifically Terraform or CloudFormation)
Strong willingness to learn new technologies and expand into multi-cloud environments
Previous experience in an agency, consultancy, or Managed Services Provider (MSP) environment
Benefits
Competitive compensation package
Retirement Savings Matching Program (RRSP)
Partnership with Perkopolis Discounts
Paid parental leave options
Flexibility & Time Off
Remote first work environment
Flexible work hours & location
Employer-paid health & dental premiums
Mental Health $500 in Health Care Spending Account annually
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.