Principal DevOps Engineer architecting NBCUniversal’s Kubernetes-native platform for broadcast production. Leading Go, AWS, Crossplane, GitOps, networking, security, and observability engineering.
Responsibilities
Architect and evolve the Kubernetes-native platform powering NBC broadcast production environments
Design platform infrastructure that automates provisioning, lifecycle management, and cloud infrastructure delivery at enterprise scale
Model broadcast infrastructure as custom resources using Crossplane compositions and custom Go functions
Lead provisioning automation across multi-account AWS environments and on-premises control rooms
Design, build, and maintain Kubernetes operators, controllers, and internal platform APIs in Go
Develop custom Crossplane providers integrating enterprise platforms such as NRCS, Venafi, and Infoblox
Manage resource lifecycles and approval workflows
Lead cloud networking, DNS strategies, cross-account connectivity, VPC topology, and dynamic network routing
Partner with broadcast systems engineers, system integrators, and external vendors
Automate bare-metal compute configurations with Puppet and integrate proprietary vendor solutions
Write RFCs, drive architectural decisions, mentor engineers, and establish CI/CD and testing strategies
Own authorization models, hierarchical RBAC, resource identifiers, and identity integrations
Drive GitOps continuous delivery using Flux, Kustomize, and Helm
Manage configuration-as-code for compute fleets using Puppet
Design observability and alerting stacks
Oversee remote desktop/VDI connectivity, authentication, credential management, and gateway routing
Contribute upstream to open-source projects and improve cloud-native solutions
Requirements
10+ years of experience designing, building, and operating production infrastructure and cloud-native platforms at enterprise scale
Strong proficiency in Go, including systems-level programming and API servers
Deep experience building Kubernetes controllers/operators using controller-runtime and kubebuilder
Expert-level knowledge of Kubernetes, including CRD/XRD generation, operators, informers, admission webhooks, and RBAC
Deep production experience with Crossplane, including composite resources, composition functions, and custom Crossplane providers in Go
Extensive production experience with AWS multi-account architectures, cross-account networking, and identity federation
Experience with EKS, EC2, VPC, IAM, STS, SSM, Secrets Manager, Route 53, and S3
Production experience with GitOps tooling, specifically Flux or ArgoCD
Hands-on Puppet experience, including module development, PuppetDB, Hiera, and r10k
Experience designing REST APIs with middleware patterns and OAuth/JWT authentication
Knowledge of information security, IAM trust chains, least-privilege policies, JWT lifecycles, and secrets abstraction
Experience designing observability platforms using Grafana, Prometheus/Mimir, Loki, OpenTelemetry, Alloy, or Prometheus Node Exporter
Working knowledge of PostgreSQL, SQLite, or similar relational databases, including schema design, migrations, and query optimization
Ability to present architectural decisions, engage with vendors, and write technical documentation
Familiarity with broadcast/media production workflows and live production constraints preferred
Experience with Crossplane function SDK and Kubernetes disaster recovery preferred
Familiarity with VDI solutions, machine identity workflows, and PKI certificate management preferred
Experience with hybrid DNS, software-defined networking, Envoy Gateway, or Gateway API preferred
Familiarity with k6, KUTTL, SOPS, Air, kind, or colima preferred
Ability to script in Bash and PowerShell
Active open-source contributions, particularly in the CNCF/Kubernetes ecosystem, preferred
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.
Senior Site Reliability Engineer securing AuthZed’s cloud infrastructure and authorization platform, including SpiceDB. Building Kubernetes guardrails, supply - chain security, vulnerability management, and incident response.
Principal SRE leading Cerence’s Site Reliability Engineering function for cloud - native automotive AI systems. Owning reliability strategy, incident escalation, observability, automation, and SLO governance.
Senior Data Scientist developing anomaly detection, dashboards, and alerts for General Motors vehicle reliability. Supporting engineering, quality, warranty, and software teams with production analytics.
DevOps Specialist scaling AWS infrastructure and automating deployments for Portage CyberTech’s digital trust, identity, privacy, and security solutions. Maintaining reliable cloud environments and CI/CD pipelines.
SRE Manager leading platform engineering teams across reliability and developer experience. Improving Lightspeed’s global cloud commerce SaaS platform through scalable infrastructure and automation.
SRE Manager leading platform engineering, reliability, and Developer Experience for Lightspeed’s global cloud POS SaaS platform. Guiding teams, infrastructure automation, and strategic delivery.
Gestionnaire SRE dirigeant deux équipes de plateforme chez Lightspeed. Assurant la fiabilité et l’automatisation de sa plateforme SaaS mondiale de commerce infonuagique.
DevOps Engineer managing AWS, Kubernetes, CI/CD, and observability for StellarTech’s scalable global EdTech products. Improving reliability, security, and production operations across multiple products.