Platform Engineer supporting production MySQL environments at Bold Commerce, enhancing reliability and operational maturity while collaborating with Engineering teams.
Responsibilities
Support the reliability, scalability, and operational maturity of production MySQL environments and core production infrastructure
Manage and improve database replication, failover, backups, recovery, upgrades, and performance optimization
Contribute to cloud infrastructure, automation, observability, and operational tooling initiatives
Support infrastructure-as-code and automation efforts using Terraform and related technologies
Improve monitoring, alerting, and operational visibility across systems and applications
Troubleshoot database, infrastructure, and application-related issues in production environments
Partner with Engineering and Platform teams to improve reliability, scalability, operational efficiency, and production readiness
Support incident response, root cause analysis, and remediation efforts
Contribute to operational best practices and continuous platform improvements
Participate in an on-call rotation supporting production systems
Requirements
5+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, Infrastructure Engineering, Database Reliability, or similar roles
Strong experience managing and supporting production MySQL database infrastructure, including replication, failover, backup/recovery, upgrades, and performance optimization
Experience supporting highly available cloud infrastructure in environments such as GCP
Exposure to or experience working with technologies such as Kubernetes, Terraform, Docker, CI/CD pipelines, and infrastructure automation
Strong Linux systems administration, troubleshooting, and incident response experience
Experience with monitoring, observability, and alerting platforms
Scripting or automation experience using Bash, Go, Python, JavaScript, or similar languages
Familiarity with security, access management, and operational best practices
Experience with configuration management tools such as Ansible is considered an asset
Strong troubleshooting, communication, and problem-solving skills
Comfortable operating independently in a lean, fast-moving environment with broad ownership and a focus on reliability, scalability, and continuous improvement
Benefits
Employer Paid Health & Dental Benefits - starting day 1!
Annual Health Benefit ($1,000 per year) to help you thrive!
Virtual mental health and EAP platform for support anytime
Backend Platform Engineer scaling Hapiko’s real - time APIs, queues, infrastructure and observability. Building reliable, secure cloud systems for Spin Master’s voice - activated kids’ sticker printer.
Azure Platform Engineer operating secure, reliable AKS infrastructure for Smile Digital Health’s FHIR - based healthcare data platform. Owning Kubernetes networking, CI/CD, observability, security, and production operations.
Senior Platform Engineer building scalable backend and cloud infrastructure for ExaCare’s AI - powered post - acute care platform. Improving reliability, developer velocity, and healthcare admission workflows.
Principal Platform Engineer leading SkyWatch’s satellite - data platform and AI agent infrastructure. Owning architecture, customer - driven roadmap delivery, and platform engineering leadership.
Platform Engineer securing Just Eat Takeaway.com’s global food - delivery edge infrastructure. Building gateways, automation, and resilient traffic routing across production environments.
Director leading Blackpoint Cyber’s cloud - based Unified Security Posture data platform for cybersecurity solutions. Driving platform roadmap, reliability, APIs, data engineering, and team growth.
Ingénieur logiciel principal intégrant des plateformes, API et solutions IA chez EDC. Gouvernance technique, sécurité, résilience et mentorat dans une société canadienne de financement du commerce.
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.