Senior DevOps Engineer modernizing AWS infrastructure and agentic automation for Big Viking Games’ live-service games. Improving reliability, security, observability, and production operations.
Responsibilities
Design, build, and operate agentic tooling that performs diagnosis, remediation, provisioning, routine maintenance, and operational follow-up
Build and maintain API integrations, MCP-style servers, webhook-driven orchestration, serverless functions, and permission-controlled automation
Convert manual runbooks, SOPs, recurring maintenance tasks, and repetitive operational chores into automation
Establish production automation guardrails, including least-privilege scopes, dry-run modes, approval paths, logging, audit trails, rollback plans, and escalation rules
Use agentic coding tools for IaC authoring, migration scripts, incident analysis, log analysis, documentation, and operational troubleshooting
Improve engineering team use of automation and AI-assisted tooling
Monitor, maintain, and improve cloud infrastructure across AWS, Netlify, Vercel, and related platforms
Drive infrastructure modernization while maintaining uptime for live games
Implement and maintain Infrastructure as Code
Improve CI/CD pipelines, release workflows, deployment reliability, and environment management
Operate containerized workloads and GitOps-based deployment patterns
Improve local development, build, test, deploy, and production support workflows
Rebuild and improve logging, metrics, alerting, dashboards, and operational visibility
Detect silent failures, pipeline failures, data freshness issues, and degraded production behavior
Maintain and monitor data pipelines between game source databases, Snowflake, and downstream analytics and reporting systems
Manage secrets and credential lifecycles, API key rotation, access controls, environment variable governance, and least-privilege practices
Support incident response, root cause analysis, remediation planning, and post-incident improvements
Automate repeatable incident response and remediation
Create executable documentation, runbooks, and SOPs
Improve backup, restore, disaster recovery, access review, and production-readiness practices
Requirements
7+ years of experience in DevOps, infrastructure engineering, cloud engineering, site reliability engineering, platform engineering, or a similar role
Strong hands-on experience designing, maintaining, and improving production infrastructure
Experience supporting live production systems where uptime, reliability, data integrity, and performance matter
Experience working with mature or legacy systems that predate modern cloud-native patterns
1+ year building with agentic coding tools, tool-calling systems, workflow automation, or infrastructure automation that takes real action
Experience with systems such as agents wired into pipelines, MCP-style integrations, automated diagnosis, automated remediation, provisioning workflows, or production-safe infrastructure tooling
Experience identifying repetitive operational work and turning it into reliable automation
Strong hands-on experience with AWS or similar cloud platforms
Proficiency with Infrastructure as Code tools such as Terraform, CloudFormation, CDK, Pulumi, or similar, including shared state and team-based workflows
Experience with containerized applications, especially Docker
Experience with container orchestration such as Kubernetes, ECS, EKS, or similar, including debugging real production issues
Experience with GitOps and declarative deployment tools such as ArgoCD, Flux, or equivalent
Experience with CI/CD tooling, version control, deployment automation, and modern release workflows
Strong understanding of Linux systems, networking, cloud security, monitoring, logging, and operational troubleshooting
Experience establishing observability, including choosing tooling, defining alerts, tuning alert noise, and creating useful dashboards
Experience with relational databases such as MariaDB, MySQL, or Postgres, including replication, backup, verified restore, and schema changes against systems that stay online
Security-aware mindset with practical experience in secrets management, credential rotation, access control, vulnerability reduction, and least-privilege practices
Self-starting and proactive about identifying and eliminating operational toil
Inventive but pragmatic approach to automation
Strong problem-solving skills for complex infrastructure or production issues
Comfortable working directly with engineers to improve build, deploy, and operational workflows
Practical ownership mindset with ability to prioritize, execute, and close loops
Strong communication with technical and non-technical stakeholders
Calm under pressure during incidents, escalations, and ambiguous production issues
Benefits
Group Retirement Savings Plan matching and participation
Comprehensive benefits package, including health, dental, and vision coverage
Health and Wellness spending account
Generous time off policies
Opportunity to support long-running live-service games with established player communities
Deep ownership of cloud modernization, DevOps automation, security improvement, and agentic infrastructure workflows
A high-impact role with meaningful ownership over reliability, performance, and engineering operations
Support and Deployment Engineer deploying and optimizing ATEME video delivery solutions for broadcast and streaming customers. Troubleshooting Linux, networking, compression, and streaming systems across North America.
Senior Site Reliability Engineer deploying Kubernetes - based AI infrastructure on NVIDIA - certified hardware for Mirantis. Ensuring reliable, secure, scalable cloud operations and customer delivery.
DevOps Engineer building and maintaining cloud infrastructure, automation, and CI/CD pipelines for Calliere's software platform. Operating containers, observability tooling, and production workloads across public clouds.
Senior DevOps Engineer improving Sherweb’s cloud - based IT solution delivery through CI/CD, IaC, GitOps, and AI automation. Supporting secure platforms, operational transitions, and developer self - service.
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.