Site Reliability Engineer improving PayByPhone’s reliable, secure SaaS mobility-payment platform. Driving observability, CI/CD, incident response, disaster recovery, and 99.99% availability.
Responsibilities
Help ensure the PayByPhone platform meets availability and stability requirements
Drive continuous improvements in software quality assurance processes, practices, and culture
Ensure reliability and stability deliverables are implemented on time and on budget and operate 24/7/365 at four nines availability (99.99%) within customer SLAs
Support incident prevention, incident response, and timely remediation
Standardize quality assurance plans and templates for cross-team projects
Assist with QA tooling selection and operational procedures
Help ensure continuity across critical business transactions
Generate and report reliability metrics to stakeholders
Define, refine, and execute the Platform Reliability / DevOps and SRE roadmap
Implement security and compliance tooling, processes, and policies within the CI/CD process
Collaborate on secure and compliant software triage, investigation, resolution, and release processes
Assist with QA within the SDLC, including automation language selection and usage
Build ownership and accountability across teams and contributors
Gather and analyze operating-system and application metrics for performance tuning and fault-finding
Participate in system design consulting for reliability and availability needs
Scope, develop, and test the disaster recovery plan
Own and improve observability, monitoring, and alerting tools; document and train users
Provide production operational support, including documentation and debugging of production issues
Be available to join Sev-1 pages
Assist with post-mortems and on-call bootcamps
Set up and support CI/CD tooling, including GitLab runners
Support platform cost reduction, budget setting, and monitoring
Support logging standards, reliable library usage, and SLA/SLO goals
Participate in on-call responsibilities when needed
Requirements
3+ years of experience in software development, delivery, or site reliability / operations for large, complex software systems spanning legacy and modern stacks
Equivalent experience gained through non-traditional paths is welcome
Bachelor's or higher degree in Computer Science, Computer Engineering, or a related technical field is preferred; equivalent hands-on experience will also be considered
Experience working with high-performing SRE, Ops, or Dev teams
Solid grounding in quality assurance discipline, software quality management, and related frameworks and tools
Working experience with Amazon Web Services (AWS) solution architectures and technologies
Experience building verification and validation practices into end-to-end delivery pipelines
Familiarity with unit, functional requirement, performance, GUI, regression, integration, system load, vulnerability assessment, security, and test automation techniques
Understanding of SaaS multi-tenant and distributed / micro-service architectures
Understanding of DevOps principles, processes, and tools, including IaC, CI/CD, and orchestration
Understanding of cloud computing architecture, services, and platforms
Understanding of web and/or mobile development technologies and programming/scripting languages
Ability to program using one or more high-level languages, such as Python or JavaScript/TypeScript
Experience with distributed storage technologies such as NFS, HDFS, and Amazon S3, as well as dynamic resource management frameworks
Experience with Infrastructure as Code (IaC)
Proactive approach to identifying problems, performance bottlenecks, and areas for improvement
Comfortable working with and supporting cross-functional teams
Strong written communication, including technical documentation and training materials
Ability to participate in on-call responsibilities and maintain a personal data plan to support them
Benefits
Retirement Savings Program (RRSP for Canada / 401(k) for U.S.-based employees)
4 weeks of vacation per year for permanent full-time employees
Up to 15 days of work from anywhere, subject to management and IT Security approval
5 personal days annually
Paid sick days
Comprehensive medical & dental coverage
Employee Assistance Program (EAP)
Professional development, continuous learning, and career progression opportunities
Site Reliability Engineer improving AXON Networks’ AI - driven ISP orchestration platform and high - speed router services. Enhancing cloud - to - device reliability, observability, automation and incident response.
We’re looking for an Azure & Databricks DevOps Engineer (12 - month renewable contract) to support a large - scale Azure data platform initiative. 📍 Hybrid role in
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.