DevOps Engineer, Cloud Infrastructure, Live Games

Posted last month

Apply Now

Resume Score

Check how well your resume matches this job before you apply.

Sign in to check score

About the role

  • Senior DevOps Engineer at Big Viking Games focusing on cloud infrastructure and automation for live-service games. Responsible for maintaining uptime and improving deployment processes in a hybrid environment.

Responsibilities

  • Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows.
  • Drive infrastructure modernization while maintaining uptime for live games with active player communities — every improvement ships while the plane is flying.
  • Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows — reducing manual toil and human error.
  • Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies.
  • Maintain and monitor data pipelines between game source databases (MariaDB), the Snowflake data warehouse, and downstream analytics and reporting systems — ensuring pipeline health, freshness, and alerting when data stops flowing.
  • Improve CI/CD pipelines, release workflows, and deployment reliability so development teams can ship safely and frequently.
  • Own secrets and credential lifecycle management across platforms — including API key rotation, access controls, environment variable governance, and least-privilege practices.
  • Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms.
  • Improve observability across the stack: logging, metrics, alerting, dashboards, and operational visibility — with particular attention to early detection of silent failures in data pipelines and production systems.
  • Support incident response, root cause analysis, remediation planning, and post-incident improvements.
  • Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency.
  • Create clear documentation, runbooks, SOPs, and repeatable processes for infrastructure and DevOps workflows.

Requirements

  • 5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
  • Strong hands-on experience with AWS or similar cloud platforms.
  • Experience designing, maintaining, and improving production infrastructure — including comfort with legacy systems that predate modern cloud-native patterns.
  • Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar.
  • Experience with containerized applications, especially Docker.
  • Experience with CI/CD tools, version control, deployment automation, and modern release workflows.
  • Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
  • Experience supporting production systems where uptime, reliability, and performance matter — especially systems that cannot tolerate extended downtime.
  • Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes.
  • Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.
  • Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages.
  • Ability to work closely with software engineers to improve build, deploy, and operational workflows.
  • Comfort creating documentation, runbooks, and repeatable operating processes.
  • Strong communication skills with both technical and non-technical stakeholders.
  • Practical ownership mindset with the ability to prioritize, execute, and close loops.
  • Nice to Have
  • Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms.
  • Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products.
  • Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools.
  • Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools.
  • Experience with Redis, Memcached, queues, workers, or event-driven systems.
  • Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
  • Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments.
  • Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
  • Experience improving cloud cost management, tagging, resource optimization, or infrastructure governance.
  • Experience with container orchestration platforms such as Kubernetes, ECS, EKS, or Nomad.
  • Experience operating infrastructure that supports AI/ML workflows, API integrations, or automation platforms (Make.com, webhook-driven orchestration, MCP servers).
  • Experience using AI tools such as Claude, ChatGPT, Gemini, or similar platforms to improve DevOps workflows, documentation, troubleshooting, and automation.
  • Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on.

Benefits

  • Group Retirement Savings Plan matching and participation.
  • Comprehensive benefits package, including health, dental, and vision coverage.
  • Health and Wellness spending account.
  • Generous time off policies.
  • Opportunity to support long-running live-service games with established player communities.
  • Exposure to cloud modernization, DevOps automation, security improvement, and AI-enabled infrastructure workflows.
  • A high-impact role with meaningful ownership over reliability, performance, and engineering operations.

Job type

Full Time

Experience level

Mid levelSenior

Salary

CA$95,000 - CA$115,000 per year

Degree requirement

Bachelor's Degree

Tech skills

AWSCloudDockerETLGrafanaJenkinsKubernetesLinuxMariaDBMySQLPostgresPrometheusRedisTerraform

Location requirements

HybridTorontoCanada

Report this job

Found something wrong with the page? Please let us know by submitting a report below.