Senior DevOps Engineer modernizing AWS infrastructure and agentic automation for Big Viking Games’ live-service games. Improving reliability, security, observability, and production operations.
Responsibilities
Design, build, and operate agentic tooling that performs diagnosis, remediation, provisioning, routine maintenance, and operational follow-up
Build and maintain API integrations, MCP-style servers, webhook-driven orchestration, serverless functions, and permission-controlled automation
Convert manual runbooks, SOPs, recurring maintenance tasks, and repetitive operational chores into automation
Establish production automation guardrails, including least-privilege scopes, dry-run modes, approval paths, logging, audit trails, rollback plans, and escalation rules
Use agentic coding tools for IaC authoring, migration scripts, incident analysis, log analysis, documentation, and operational troubleshooting
Improve engineering team use of automation and AI-assisted tooling
Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms
Drive infrastructure modernization while maintaining uptime for live games
Implement and maintain Infrastructure as Code
Improve CI/CD pipelines, release workflows, deployment reliability, and environment management
Operate containerized workloads and GitOps-based deployment patterns
Improve local development, build, test, deploy, and production support workflows
Rebuild and improve logging, metrics, alerting, dashboards, and operational visibility
Detect silent failures, pipeline failures, data freshness issues, and degraded production behavior
Maintain and monitor data pipelines between game source databases, Snowflake, and downstream analytics and reporting systems
Manage secrets and credential lifecycles, API key rotation, access controls, environment variable governance, and least-privilege practices
Support incident response, root cause analysis, remediation planning, and post-incident improvements
Automate repeatable incident response and remediation
Create executable documentation, runbooks, and SOPs
Improve backup, restore, disaster recovery, access review, and production-readiness practices
Requirements
7+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, platform engineering, or a similar role
Strong hands-on experience designing, maintaining, and improving production infrastructure
Experience supporting live production systems where uptime, reliability, data integrity, and performance matter
Experience working with mature or legacy systems that predate modern cloud-native patterns
1+ year building with agentic coding tools, tool-calling systems, workflow automation, or infrastructure automation that takes real action
Experience with systems such as agents wired into pipelines, MCP-style integrations, automated diagnosis, automated remediation, provisioning workflows, or production-safe infrastructure tooling
Experience identifying repetitive operational work and turning it into reliable automation
Strong hands-on experience with AWS or similar cloud platforms
Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar, including shared state and team-based workflows
Experience with containerized applications, especially Docker
Experience with container orchestration such as Kubernetes, ECS, EKS, or similar, including debugging real production issues
Experience with GitOps and declarative deployment tools such as ArgoCD, Flux, or equivalent
Experience with CI/CD tooling, version control, deployment automation, and modern release workflows
Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting
Experience establishing observability, including choosing tooling, defining alerts, tuning alert noise, and creating useful dashboards
Experience with relational databases such as MariaDB, MySQL, or Postgres, including replication, backup, verified restore, and schema changes against systems that stay online
Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices
Self-starting and proactive about identifying and eliminating operational toil
Inventive but pragmatic approach to automation
Strong problem-solving skills for complex infrastructure or production issues
Comfortable working directly with engineers to improve build, deploy, and operational workflows
Practical ownership mindset with ability to prioritize, execute, and close loops
Strong communication with technical and non-technical stakeholders
Calm under pressure during incidents, escalations, and ambiguous production issues
Benefits
Group Retirement Savings Plan matching and participation
Comprehensive benefits package, including health, dental, and vision coverage
Health and Wellness spending account
Generous time off policies
Opportunity to support long-running live-service games with established player communities
Deep ownership of cloud modernization, DevOps automation, security improvement, and agentic infrastructure workflows
A high-impact role with meaningful ownership over reliability, performance, and engineering operations
DevOps Engineer building AWS platforms, CI/CD pipelines, and infrastructure automation for SWTCH’s North American EV charging solutions. Improving deployment reliability, observability, and cloud architecture.
Reliability Engineering Manager leading SAP PM and asset - reliability strategy. Improving equipment performance across Apotex’s global pharmaceutical manufacturing sites.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.