Director of SRE leading global infrastructure and reliability teams for Blackpoint Cyber's cyber defense services. Focusing on cost efficiency, automation, and team leadership.
Responsibilities
Lead the design, implementation, and management of scalable, reliable, and highly available cloud-based infrastructure (AWS/Azure)
Establish SRE best practices, including monitoring, incident response, capacity planning, and performance tuning
Improve observability, monitoring, and alerting, ensuring quick detection and resolution of reliability issues
Drive automation-first approaches, reducing manual intervention through Infrastructure-as-Code (IaC) and CI/CD pipelines
Lead a team of SREs, applying Blackpoint Cyber's management values of Coach, Model, Care, in defining business-critical outcomes, creating action plans, and supporting the team in achieving them
Continue hands-on contributions in an SRE role
Design, implement, and support key infrastructure, including automated attack infrastructure deployment, isolated identity and productivity environments, and secure data storage
Establish and apply security hygiene and monitoring policies to meet Blackpoint Cyber security requirements
Monitor and optimize cloud spending, ensuring cost-effective resource utilization without compromising reliability
Manage and mentor a global team of SREs, DevOps engineers, and cloud infrastructure specialists
Collaborate with security teams to ensure compliance, security hardening, and disaster recovery readiness
Requirements
10+ years of experience in SRE, DevOps, or Cloud Infrastructure roles
5+ years of experience in people management, leading SRE team
Strong experience with AWS, Azure, or GCP, with expertise in cost management and scaling strategies
Proficiency in Infrastructure-as-Code (IaC) (e.g., Terraform, CloudFormation, Pulumi)
Hands-on experience with CI/CD pipelines, Kubernetes, and container orchestration
Expertise in monitoring, logging, and observability tools (e.g., Prometheus, Grafana, Datadog, Splunk)
Proven ability to optimize cloud costs (COGS) while maintaining reliability and performance
Strong leadership, collaboration, and problem-solving skills
Experience with SLA/SLO/SLIs will be valuable
A general understanding of the modern AI tooling landscape and how SRE can use that to increase velocity and improve stability
Support and Deployment Engineer deploying and optimizing ATEME video delivery solutions for broadcast and streaming customers. Troubleshooting Linux, networking, compression, and streaming systems across North America.
Senior Site Reliability Engineer deploying Kubernetes - based AI infrastructure on NVIDIA - certified hardware for Mirantis. Ensuring reliable, secure, scalable cloud operations and customer delivery.
DevOps Engineer building and maintaining cloud infrastructure, automation, and CI/CD pipelines for Calliere's software platform. Operating containers, observability tooling, and production workloads across public clouds.
Senior DevOps Engineer improving Sherweb’s cloud - based IT solution delivery through CI/CD, IaC, GitOps, and AI automation. Supporting secure platforms, operational transitions, and developer self - service.
Senior DevOps / Cloud Infrastructure Engineer needed for hybrid role in North York, ON. Requires 10+ years experience with GCP, AWS, Kubernetes, Terraform, and CI/CD.
Staff Site Reliability Engineer strengthening AWS and Kubernetes resilience for Caseware, a fintech company building audit and accounting software. Driving secure delivery, observability, and incident management.
Application Reliability Engineer supporting Innodata’s Google Cloud enterprise applications. Restoring production services, managing deployments, and enhancing microservices for a global AI data engineering company.