Customer Reliability Engineer at OpsGuru, enhancing customer service through cloud solutions. Leading reliability investigations and employing cutting-edge technology for digital transformation.
Responsibilities
Lead troubleshooting investigations to bring quicker issue resolution to complex problems impacting our customers.
Act as the customer's technical single point of contact, establishing credibility and building impactful relationships
Own core reliability workstreams: incident response, backup/DR validation, security patching, secure access management, and audit log monitoring
Enabling the Operations team to deliver excellent customer service through great documentation and knowledge transfer
Work with a wide variety customers, industries, and bleeding-edge technologies
Develop skills and contributions through educational and personal growth opportunities
Requirements
+7 years of experience in Information technology, preference for cloud engineering, site reliability engineering, infrastructure or platform services operations.
+4 hands-on experience with AWS Cloud infrastructure and ecosystem (EC2, RDS, S3, VPC, IAM, KMS, Backup, CloudTrail, Route53, ECS/EKS, etc.)
Proficiency with Network and Security troubleshooting (Routing, DNS, Security Groups, ACL’s, WAF, VPN/Bastion access).
Proficiency in at least one scripting language such as Ruby, Python, Bash or Powershell
Proficiency in at least one of the following Infrastructure as Code frameworks: Terraform, CDK, CloudFormation
Previous Production On call 24/7 experience, ideally with an incident management/paging tool (e.g., PagerDuty) - MUST
Previous Windows/Linux server administration and patching experience
Familiarity with at least one of the following monitoring platform: CloudWatch, Grafana, New Relic, Datadog, or Prometheus
Conversational level English required
Experience working with Claude Code, Codex or similar AI driven development environment
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.