Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Responsibilities
Create reliability initiatives across the lifecycle of critical grid automation embedded and software products
Lead cross-functional teams applying reliability engineering methodologies to ensure system resilience, regulatory compliance, and long-term field performance
Perform reliability modeling and statistical analysis of energy management and protection/control products
Develop and execute hardware and software reliability test plans, including HIL, environmental, and stress testing
Conduct and lead FMEA, FMECA, and Root Cause Analysis for hardware/software failure events across embedded systems and communication layers
Work with product design, systems engineering, and software teams to embed Design for Reliability from concept through deployment
Analyze field data from deployed devices to identify systemic issues and recommend corrective/preventive actions
Collaborate with utilities and customers on grid asset performance programs and grid modernization initiatives
Develop and maintain reliability KPIs, including MTBF, FIT rate, and availability, and integrate them into product development and quality programs
Drive predictive and condition-based maintenance strategies using remote diagnostics, SCADA, and sensor data
Align with the Innovation Office and Applications team on reliability priorities, working groups, standards leadership, and technology strategy
Monitor regulatory compliance requirements such as NERC, IEEE, and IEC standards affecting utility asset reliability and cybersecurity
Provide coaching, training, and guidance to internal teams and external stakeholders
Requirements
Bachelor’s degree in electrical engineering, Systems Engineering, or a related discipline
Extensive experience in reliability, systems, or product engineering in the energy or utility sector
Familiarity with protection & control systems, grid-edge devices, substation automation, and utility-grade software solutions
Hands-on experience with reliability engineering tools and standards, including ReliaSoft, ISO 9001/55000, and IEC 61025
Familiarity with T&D communications protocols and standards, including IEC 61850, DNP3, Modbus, and IEEE C37.118
Strong analytical skills and experience with Python, R, Minitab, or Power BI
Excellent written and verbal communication skills; able to engage cross-functional teams and utility customers
Deep understanding of grid reliability challenges in T&D environments
Strong systems thinking and problem-solving ability
Collaborative leadership across engineering, field teams, and customers
Data-driven problem-solving mindset with attention to long-term product health
Proactive approach to emerging technologies and continuous improvement
Master’s degree in Reliability or Systems Engineering
Certified Reliability Engineer (CRE) or similar professional certification
Familiarity with model-based design and HIL testing, including RTDS, PSCAD, or EMTP
Experience working directly with T&D utilities, supporting product deployments, failure investigations, or asset strategies
Familiarity with cybersecurity risk and reliability considerations for critical infrastructure
Benefits
Discretionary annual bonus
Culture of ownership, innovation, and impact
Opportunity to shape the future of grid reliability and support smarter, more resilient energy systems
Opportunity to lead initiatives influencing design, customer outcomes, and industry standards
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.