Senior Site Reliability Engineer at Movable Ink managing multi-cloud content serving infrastructure. Supporting reliability initiatives and collaborative strategies for scaling platform operations.
Responsibilities
Improve the tooling and automation of our infrastructure to minimize manual work, increase performance, and decrease the frequency and severity of incidents
Build, maintain, and support core applications
Monitor our systems for capacity, performance, and troubleshoot issues
Partner with the rest of the SRE team to ensure smooth, continued delivery of our service to clients
Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement
Requirements
Experience in Site Reliability or Software Engineering, building and maintaining scalable, resilient services.
Building the tooling and automation to manage those services, as well as investigating system and application metrics to diagnose and resolve performance issues.
4+ years experience as an SRE or Software Engineer, with a focus on Cloud platforms (AWS/GCP)
Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
Experience and willingness to operate in an on-call environment, evaluating and improving monitoring and alerting systems, and developing run books to investigate and debug issues
Strong experience with infrastructure as code tools. Terraform experience is a major plus
Kubernetes experience, including cluster operations, multi-tenancy strategies, and supporting teams on container orchestration best practices. We use EKS and GKE
Experience with one or more high level programming languages; NodeJS, Go, Ruby, Python, in addition Shell Scripting
Linux experience is a must
Benefits
full range of medical, financial, and/or other benefits
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.