Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud-native, microservices-based environments.
Responsibilities
Define and implement observability strategies, standards, and governance across applications and platforms
Design and maintain monitoring, alerting, dashboarding, and reporting solutions using Dynatrace or equivalent observability platforms
Establish and drive SRE best practices, including SLIs, SLOs, error budgets, and symptom-based alerting
Partner with engineering and product teams to improve system reliability, performance, and operational maturity
Develop standards for tagging, ownership, dashboard design, access management, and alerting governance
Support non-specialized observability teams through guidance, coaching, and knowledge transfer
Lead technical workstreams, prioritize initiatives, and ensure delivery within defined timelines and budgets
Analyze distributed systems and troubleshoot complex production issues using monitoring and tracing data
Promote documentation, operational rigor, and continuous improvement across engineering teams
Collaborate within a distributed, multilingual environment
Requirements
Significant experience in Site Reliability Engineering within large-scale production environments
Deep understanding of Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and symptom-based alerting
Proven expertise with enterprise observability platforms such as Dynatrace, Datadog, New Relic, or AppDynamics
Strong experience with Application Performance Monitoring (APM), Real User Monitoring (RUM), monitoring agents and instrumentation, alerting strategies, RBAC, SLO management, and tagging and governance models
Strong knowledge of OpenTelemetry (OTEL) and distributed tracing
Experience with composable, microservices-based architectures
Hands-on production experience with AWS and Kubernetes
Experience with infrastructure and operational automation
Practical knowledge of Terraform, Bash scripting, and Python scripting
Experience with CI/CD tools such as GitLab CI or equivalent pipeline/workflow platforms
Demonstrated ability to lead technical initiatives and workstreams
Experience working within complex operational and Agile environments
Strong stakeholder management and collaboration skills
Excellent communication skills in both French and English; French-speaking skills are needed
Strong documentation practices, organizational skills, and attention to detail
High degree of autonomy and ownership
Benefits
Comprehensive insurance plan with Gold, Silver, or Bronze modules; employer contribution up to 80%; short- and long-term disability coverage
Dialogue via Sun Life virtual healthcare services
Employee and Family Assistance Program
Complete mental health support program
$500 Personal Spending Account for healthcare reimbursements, gym memberships, public transit passes, office supplies, or RRSP contributions
RRSP retirement plan with 100% employer matching through DPSP, up to 4%
Flexible vacation policy, including 5 days during the probation period
$30/month Personal Technology Reimbursement, offered from day 1
Winter holiday company closure
Flexible scheduling throughout the year
Growth opportunities, continuous learning, professional growth, and international career opportunities
Inclusion and accessibility support, including reasonable interview accommodations
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.