Senior Reliability Engineer improving embedded grid automation reliability for utility-scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Responsibilities
Create reliability initiatives across the lifecycle of critical grid automation embedded and software products
Lead cross-functional teams in applying reliability engineering methodologies to ensure system resilience, regulatory compliance, and long-term field performance
Perform reliability modeling and statistical analysis to predict field performance and support design decisions
Develop and execute hardware and software reliability test plans, including HIL, environmental, and stress testing
Conduct and lead FMEA, FMECA, and Root Cause Analysis for hardware/software failure events
Collaborate with product design, systems engineering, and software teams to embed Design for Reliability from concept through deployment
Analyze field data from deployed devices to identify systemic issues and recommend corrective and preventive actions
Collaborate with utilities and customers on grid asset performance programs and grid modernization initiatives
Develop and maintain reliability KPIs, including MTBF, FIT rate, and availability
Drive predictive and condition-based maintenance strategies using remote diagnostics, SCADA, and sensor data
Align with the Innovation Office and Applications team on reliability priorities, standards leadership, and technology strategy
Monitor regulatory compliance requirements affecting utility asset reliability and cybersecurity
Provide coaching, training, and guidance to internal teams and external stakeholders
Requirements
Bachelor’s degree in electrical engineering, Systems Engineering, or a related discipline
Extensive experience in reliability, systems, or product engineering in the energy or utility sector
Familiarity with protection and control systems, grid-edge devices, substation automation, and utility-grade software solutions
Hands-on experience with reliability engineering tools and standards, such as ReliaSoft, ISO 9001/55000, and IEC 61025
Familiarity with T&D communications protocols and standards, including IEC 61850, DNP3, Modbus, and IEEE C37.118
Experience with data analysis platforms, such as Python, R, Minitab, and Power BI
Understanding of grid reliability challenges in T&D environments
Familiarity with model-based design and HIL testing, including RTDS, PSCAD, or EMTP
Experience working directly with T&D utilities, supporting product deployments, failure investigations, or asset strategies
Familiarity with cybersecurity risk and reliability considerations for critical infrastructure
Master’s degree in Reliability or Systems Engineering is desired
Certified Reliability Engineer (CRE) or similar professional certification is desired
Benefits
Discretionary annual bonus
Culture of ownership, innovation, and impact
Opportunity to shape the future of grid reliability
Opportunity to lead initiatives influencing design, customer outcomes, and industry standards
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.