Senior DevOps Engineer at Big Viking Games focusing on cloud infrastructure and automation for live-service games. Responsible for maintaining uptime and improving deployment processes in a hybrid environment.
Responsibilities
Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms that support our live games, data systems, internal tools, and AI-powered operational workflows.
Drive infrastructure modernization while maintaining uptime for live games with active player communities — every improvement ships while the plane is flying.
Build, maintain, and improve automation for deployments, environment management, provisioning, secrets rotation, and operational workflows — reducing manual toil and human error.
Implement and maintain Infrastructure as Code using tools such as Terraform, CloudFormation, CDK, or similar technologies.
Maintain and monitor data pipelines between game source databases (MariaDB), the Snowflake data warehouse, and downstream analytics and reporting systems — ensuring pipeline health, freshness, and alerting when data stops flowing.
Improve CI/CD pipelines, release workflows, and deployment reliability so development teams can ship safely and frequently.
Own secrets and credential lifecycle management across platforms — including API key rotation, access controls, environment variable governance, and least-privilege practices.
Support and improve the infrastructure that powers AI and automation tooling, including API integrations, MCP servers, serverless functions, webhook reliability, and orchestration platforms.
Improve observability across the stack: logging, metrics, alerting, dashboards, and operational visibility — with particular attention to early detection of silent failures in data pipelines and production systems.
Support incident response, root cause analysis, remediation planning, and post-incident improvements.
Help manage cloud spend, infrastructure usage, resource tagging, and environment efficiency.
Create clear documentation, runbooks, SOPs, and repeatable processes for infrastructure and DevOps workflows.
Requirements
5+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, or a similar role.
Strong hands-on experience with AWS or similar cloud platforms.
Experience designing, maintaining, and improving production infrastructure — including comfort with legacy systems that predate modern cloud-native patterns.
Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar.
Experience with containerized applications, especially Docker.
Experience with CI/CD tools, version control, deployment automation, and modern release workflows.
Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting.
Experience supporting production systems where uptime, reliability, and performance matter — especially systems that cannot tolerate extended downtime.
Experience with relational databases (MariaDB, MySQL, Postgres) and comfort working adjacent to data pipelines and ETL processes.
Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices.
Strong problem-solving skills and the ability to investigate complex infrastructure or production issues, including silent failures and data pipeline outages.
Ability to work closely with software engineers to improve build, deploy, and operational workflows.
Comfort creating documentation, runbooks, and repeatable operating processes.
Strong communication skills with both technical and non-technical stakeholders.
Practical ownership mindset with the ability to prioritize, execute, and close loops.
Nice to Have
Experience in gaming, live-service products, SaaS, digital products, or other high-availability consumer platforms.
Experience supporting live games, virtual worlds, multiplayer systems, or real-time online products.
Experience with GitHub Actions, GitLab CI, Jenkins, CircleCI, Buildkite, or similar CI/CD tools.
Experience with Datadog, Grafana, Prometheus, CloudWatch, ELK, OpenTelemetry, or similar observability tools.
Experience with Redis, Memcached, queues, workers, or event-driven systems.
Experience with Snowflake, data warehouse connectivity, ETL monitoring, or data pipeline reliability.
Experience with serverless platforms (Netlify Functions, Vercel, AWS Lambda) and multi-platform hosting environments.
Experience with disaster recovery, backup strategies, incident management, load testing, and performance tuning.
Experience with container orchestration platforms such as Kubernetes, ECS, EKS, or Nomad.
Experience operating infrastructure that supports AI/ML workflows, API integrations, or automation platforms (Make.com, webhook-driven orchestration, MCP servers).
Experience using AI tools such as Claude, ChatGPT, Gemini, or similar platforms to improve DevOps workflows, documentation, troubleshooting, and automation.
Experience working in small, high-leverage engineering teams where infrastructure ownership is broad and hands-on.
Benefits
Group Retirement Savings Plan matching and participation.
Comprehensive benefits package, including health, dental, and vision coverage.
Health and Wellness spending account.
Generous time off policies.
Opportunity to support long-running live-service games with established player communities.
Exposure to cloud modernization, DevOps automation, security improvement, and AI-enabled infrastructure workflows.
A high-impact role with meaningful ownership over reliability, performance, and engineering operations.
Senior SRE operating Kubernetes and cloud infrastructure for Penn Entertainment’s sports betting and media platforms. Leading migrations, automation, observability, and incident response across regulated production services.
DevOps Engineer operating multi - cloud Kubernetes infrastructure for InfluxData’s time - series platform. Automating operations and supporting highly available distributed services.
Senior Reliability Engineer improving embedded protection, control, and software products for utility grids. Leading reliability testing, failure analysis, and modernization initiatives for resilient energy systems.
Senior Reliability Engineer improving embedded grid automation reliability for utility - scale energy systems. Leading testing, failure analysis, KPIs, and modernization initiatives with utilities.
Site Reliability Expert leading observability and SRE for Valtech, an experience innovation company. Improving reliability across cloud - native, microservices - based environments.
Staff DevOps Engineer owning reliable, scalable infrastructure for Nexxa’s AI systems. Supporting machine learning workloads across heavy - industry operations.
Staff SRE securing and scaling IAM systems at RBC, a Canadian bank. Designing resilient infrastructure, automating operations, and leading incident response.
DevOps Engineer designing production - style CI/CD, cloud, and infrastructure tasks for Your Software Supplier. Reviewing AI - generated solutions and ensuring correctness and reproducibility.