Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice-management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
Responsibilities
Own complex Azure platform initiatives from architecture and design through implementation, production readiness, and ongoing optimization
Anticipate operational, security, scalability, and reliability risks and implement durable solutions
Establish engineering standards, documentation, reusable automation, and continuous improvement
Act as a senior technical authority for Azure platform architecture, cloud-native DevOps practices, and production operations
Lead the design and delivery of scalable SaaS infrastructure, including technical planning, architecture reviews, implementation, validation, and operational handoff
Define and promote platform standards, reference architectures, guardrails, and engineering best practices
Mentor engineers and lead cross-functional troubleshooting and design discussions
Design, deploy, automate, and manage enterprise Azure infrastructure and networking
Architect, build, and operate cloud-native application platforms using AKS, ACR, App Services, and Function Apps
Design and maintain Infrastructure as Code and configuration automation using Terraform and Ansible
Own Azure DevOps pipelines and automated build, deployment, validation, rollback, and release processes
Develop reusable modules, templates, and golden images using Packer and Azure Compute Gallery
Design, deploy, automate, and support Azure SQL Database, SQL Elastic Pools, PostgreSQL Flexible Server, and Elastic Job Agent
Implement secure infrastructure and delivery practices using managed identities, Key Vault, security groups, secrets management, RBAC, and least-privilege access
Build monitoring, alerting, logging, tracing, and performance management solutions using Datadog and OpenTelemetry connectors
Participate in on-call and incident response; lead incident triage, root cause analysis, corrective actions, and prevention efforts
Enable application integration and event-driven architectures using APIM, Logic Apps, Service Bus, and App Configuration
Support deployment, governance, monitoring, and operation of Azure OpenAI, Azure AI Foundry, AI Foundry Projects, and Document Intelligence
Requirements
Advanced knowledge of systems, cloud architecture, networking, identity, security, and production operations
Demonstrated experience designing, building, deploying, and maintaining scalable SaaS solutions in Microsoft Azure
Deep hands-on experience with Infrastructure as Code, CI/CD, containers, Kubernetes, observability, and automation in enterprise environments
Experience leading complex technical initiatives with multiple stakeholders, dependencies, and production risks
Strong incident management and troubleshooting skills across application, platform, network, security, and data layers
Ability to translate business and engineering needs into pragmatic, secure, and maintainable technical solutions
Clear written and verbal communication, including the ability to explain technical tradeoffs to technical and non-technical partners
Ability to take end-to-end ownership and operate effectively with limited direction
Sound judgment when requirements are ambiguous and ability to escalate material risks with clear options and recommendations
Ability to influence technical direction through expertise, credibility, collaboration, and measurable outcomes
Ability to proactively identify systemic issues and resolve root causes
Ability to build strong partnerships across Engineering, Security, Product, and other business functions
Detail-oriented, security-minded, customer-aware, and committed to continuous improvement
Benefits
Flexible PTO
Medical, Dental, Paid Sick Days, Vision, and Supplemental Coverage
Flexible Spending Account
Health Savings Account
401(k) match
Potential for bonuses based on company performance
Potential for merit increases based on performance
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.