DevOps Engineer/Site Reliability Engineer for the financial sector. Supporting high-impact, mission-critical technology platforms with a focus on reliability and automation.
Responsibilities
Partner with development, operations, and security teams to design and deliver secure, scalable, and resilient infrastructure solutions
Build and maintain automation frameworks for deployment, scaling, and observability
Design and implement CI/CD pipelines , release strategies, and recovery mechanisms
Continuously improve system performance, availability, reliability, and security posture
Provide end-to-end ownership of mission-critical platforms, including production support and root-cause analysis
Proactively monitor systems using observability tools to identify and address performance and reliability improvements
Lead or contribute to incident response , triage, and resolution with a focus on rapid recovery and prevention
Support deployment activities and manage implementation issues through to resolution
Drive adoption of modern engineering practices, tools, and processes to enhance delivery and operational efficiency
Analyze complex technical issues and recommend solutions aligned with business impact
Ensure compliance with enterprise standards and regulatory requirements
Participate in an on call rotation to support production systems (if needed)
Requirements
5+ years of experience in DevOps, SRE, or related roles in hybrid (on-prem and AWS) environments
Strong experience with AWS services (IaaS, PaaS, RDS, observability tools) and Infrastructure as Code ( CDK / TypeScript preferred )
Hands-on experience with configuration management tools (Ansible, YAML)
Strong experience with observability platforms such as Dynatrace, Elasticsearch, and CloudWatch
Proficiency in scripting and development using Python, Bash, and/or JavaScript
Experience implementing automation-first solutions across infrastructure and application layers
Solid understanding of security and compliance practices within regulated industries (financial services preferred)
Experience with Git-based workflows (GitHub preferred)
Working knowledge of ServiceNow and ITSM processes (Incident, Problem, Change, Release, Configuration Management)
Strong experience with RHEL systems administration and clustering technologies (e.g., Veritas Cluster)
Proven ability to support and operate large-scale, mission-critical systems
We’re looking for an Azure & Databricks DevOps Engineer (12 - month renewable contract) to support a large - scale Azure data platform initiative. 📍 Hybrid role in
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.