Platform Engineering Manager overseeing hybrid and cloud infrastructures for a tech company. Leading team improvements in production and developer experience across multiple regions.
Responsibilities
Design, document, implement and maintain hybrid and multiple cloud infrastructures whilst working closely with other cloud and engineering teams to implement projects and infrastructure improvements.
Communicate and work alongside various application and development teams to increase our uptime and maintain SLO and SLA.
Attend internal and third-party meetings covering incidents or service improvements.
Become a subject matter expert in our Product offerings, how they are served by infrastructure be it cloud or physical and ensure right sizing / scalability.
Ensure all systems are observable, promote a culture of fix before fail.
Bring positive and eager attitude to the team.
Draw on best practices and your knowledge of internal and external business issues to improve products or services.
Manage a team of 2 other platform engineers, having proactive coaching and development discussions with them.
Work with the CTO to develop and maintain the roadmap for the platform engineering team.
Manage the time of the team ensuring projects are delivered on time and any delays are proactively communicated to stakeholders.
Requirements
Excellent customer facing/customer service skills.
Able to troubleshoot/work under pressure, meet deadlines.
Experience with Observability Systems, e.g., New Relic.
A thorough understanding and commercial experience of AWS cloud technologies, e.g., EC2, ECS, Route53, S3.
A specific focus on AWS serverless components and deployment pipelines e.g., Lamda, API Gateway, DyanmoDB, Auroura, Terraform, Cloudformation etc.
Various Database architecture experience, e.g., MariaDB, Cassandra.
Proficiency in two or more scripting languages, e.g., Python, Bash, PowerShell.
Understanding of Site Reliability Engineering and key concepts.
Senior Platform Engineer building scalable backend and cloud infrastructure for ExaCare’s AI - powered post - acute care platform. Improving reliability, developer velocity, and healthcare admission workflows.
Principal Platform Engineer leading SkyWatch’s satellite - data platform and AI agent infrastructure. Owning architecture, customer - driven roadmap delivery, and platform engineering leadership.
Platform Engineer securing Just Eat Takeaway.com’s global food - delivery edge infrastructure. Building gateways, automation, and resilient traffic routing across production environments.
Director leading Blackpoint Cyber’s cloud - based Unified Security Posture data platform for cybersecurity solutions. Driving platform roadmap, reliability, APIs, data engineering, and team growth.
Ingénieur logiciel principal intégrant des plateformes, API et solutions IA chez EDC. Gouvernance technique, sécurité, résilience et mentorat dans une société canadienne de financement du commerce.
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.