Senior DevOps Engineer responsible for AWS infrastructure, ensuring system reliability at RAZR Marketing. Collaborating on operational excellence and disaster recovery for enterprise-scale applications.
Responsibilities
Design, implement, and maintain infrastructure using Pulumi.
Automate infrastructure provisioning and configuration management.
Manage environment configurations across dev, staging, and production.
Implement version control and code review processes for infrastructure changes.
Develop reusable infrastructure components and modules.
Manage and optimize AWS services including ECS, Lambda, RDS, S3, CloudFront, and VPC configurations.
Implement AWS best practices for security, performance, and cost optimization.
Monitor and maintain system health across multiple environments (dev, staging, production).
Conduct regular infrastructure audits and implement improvements.
Design, implement, and maintain disaster recovery plans for critical systems and databases.
Develop and execute backup strategies for RDS databases, S3 data, and application configurations.
Conduct regular DR drills and validate recovery procedures.
Document and maintain runbooks for disaster recovery scenarios.
Implement multi-region failover strategies for high-availability services.
Perform routine system maintenance including patching, updates, and security hardening.
Manage database maintenance tasks including backups, performance tuning, and capacity planning.
Monitor system performance and proactively address potential issues.
Support application deployments for Angular containers, NestJS servers, and Java Spring Boot services.
Implement and maintain AWS security best practices including IAM policies, security groups, and encryption.
Maintain comprehensive monitoring solutions using CloudWatch, application logs, and custom metrics.
Monitor and optimize AWS costs through resource right sizing and reserved capacity planning.
Create and maintain operational documentation, runbooks, and architecture diagrams.
Requirements
Bachelor's degree in Computer Science, Information Technology, or related field.
7+ years of experience in DevOps, systems administration, or cloud operations roles.
Solid experience with infrastructure as code tools, particularly Pulumi, Terraform or Cloud Formation.
Strong hands-on experience with AWS services and cloud architecture.
Proven experience designing and implementing disaster recovery solutions.
Solid understanding of database administration, particularly PostgreSQL/RDS.
Experience with containerization using Docker and orchestration with ECS or Kubernetes.
Proficiency in scripting languages: Python, Bash, or Ruby for automation tasks.
Strong understanding of networking concepts, VPCs, security groups, and load balancers.
Experience with monitoring and logging tools (CloudWatch, ELK stack, or similar).
Knowledge of security best practices and compliance requirements.
Excellent problem-solving skills and ability to troubleshoot complex system issues.
Strong communication skills and ability to document technical procedures clearly.
Experience with banking or financial services systems is a plus.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.
Team Leader guiding Remote’s global SRE platform for compliant international employment. Leading engineers and reliability across Kubernetes, AWS, observability, and infrastructure.
AWS and DevOps Engineer establishing secure, automated environments for a bilingual nonprofit digital platform. Managing deployment, monitoring, recovery, and operational handover.