Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
Responsibilities
Build and operate infrastructure behind a large-scale sports betting and media platform
Own critical infrastructure across compute, networking, storage, and cloud services
Drive complex infrastructure migrations and projects across production environments and jurisdictions
Build and maintain platform tooling and automation using ArgoCD, Helm, GitHub Actions, release pipelines, and service onboarding workflows
Support development teams with infrastructure consulting, dependency resolution, architecture reviews, and platform-tool adoption
Design and improve Datadog observability, alerting, dashboards, and runbooks
Provide operational support and incident response through structured debugging and root cause analysis
Participate in on-call rotations
Mentor teammates and contribute to architecture decisions and continuous improvement
Requirements
5+ Years of Experience in a similar role (DevOps, Site Relatability Engineer)
Strong experience operating and troubleshooting Kubernetes in a production Linux environment (cluster lifecycle, networking, storage, scheduling)
Experience working with AWS, GCP, and/or on-premise environments
Proficiency in at least two of: Go, Python, Bash/Shell
Deep understanding of distributed systems, failure modes, networking fundamentals, capacity planning, and performance analysis
Experience with GitOps and CI/CD workflows (ArgoCD, Helm, GitHub Actions, or similar)
Experience with infrastructure-as-code (Terraform, Helm, or equivalent)
Track record of leading complex migrations or infrastructure projects with cross-team dependencies
Strong incident response and troubleshooting skills
Clear technical communication and documentation skills
Experience with service mesh technologies (Istio, Cilium) preferred
Familiarity with distributed storage systems (Ceph, or similar) preferred
Experience with bare-metal Kubernetes or Talos OS preferred
Exposure to regulated environments preferred
Experience with Datadog or comparable observability platforms at scale preferred
Familiarity with PostgreSQL, PgBouncer, or database migration tooling preferred
Benefits
Competitive compensation package
Comprehensive Benefits package
Fun, relaxed work environment
Education and conference reimbursements
Bonus eligibility for most non-sales positions
Best-in-class benefits with personalized physical, financial, and emotional support options
Senior DevOps Engineer building secure, scalable Azure platforms for CARET’s legal and accounting practice - management software. Leading infrastructure, Kubernetes, CI/CD, security, observability, and reliability initiatives.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.