Senior Site Reliability Engineer for Hopper's Cloud FinOps team managing cloud infrastructure. Focused on optimizing cost efficiency and system reliability for fintech solutions.
Responsibilities
Work on projects that will drive a higher cost efficiency, such as:
Reduce our network egress costs by removing unnecessary headers.
Ensure that our warehouse data is in use and select the most efficient storage for it. E.g., cold storage for buckets with infrequent retrieval.
Ensure that autoscaling for both databases and compute is well optimized.
Work on improving the current cost attribution to ensure all teams have clear visibility into their costs.
You will also participate in providing support to incidents and be part of on-call rotation for platform incidents, as each engineering team has their own on-call rotation (Team is scattered across America and Europe, so you can sleep at night!). You will also contribute to solving doubts and problems engineers might face with our infrastructure and approving PRs that require Platform supervision.
You will be part of a small and highly efficient team of SREs.
Requirements
Strong background in SRE, DevOps, Software Engineering or Systems engineering
Troubleshooting skills
System design with good analytical capabilities
Good communication skills
Knowledge of major cloud providers, preferably Google Cloud
SQL knowledge
Containers, Kubernetes, and related tooling like Kustomize and Helm
Service Mesh, preferably with Istio
Networking knowledge. DNS, TLS, certificates, ingresses, etc.
Observability with log collection, metrics, APM, etc. preferably Datadog
Security knowledge, IAM, RBAC, network security, etc.
Knowledge on authentication and authorization technologies
CI/CD
Database technologies
Competent in scripting with Bash and Python or other scripting languages
Benefits
Well-funded and proven startup with large ambitions, competitive salary and upsides of pre-IPO equity packages.
Hopper covers 100% of the premiums for group insurance plan.
Hopper offers life, short term and long term disability coverage.
HSA that covers eligible medical and dental expenses.
All employees and dependents have access to Dialogue’s telemedicine services, anytime, anywhere.
All employees have access to an RRSP plan with automatic pre-tax withdrawals per pay.
Please ask us about our very generous parental leave, much above industry standards!.
Unlimited PTO.
Carrot Cash travel stipend.
Access to co-working space on demand through FlexDesk AND Work-from-home stipend.
Entrepreneurial culture where pushing limits and taking risks is everyday business.
Open communication with management and company leadership.
DevOps Intern supporting CI/CD, cloud infrastructure, and automation for Ludia’s mobile game studio. Improving reliability and developer tools in production game environments.
Manager leading global SRE teams for Akamai's distributed Cloud IAM services. Improving reliability, scalability, security, and usability through cloud - native tooling and software.
Cloud Engineer supporting Kinaxis’s AI - powered supply chain orchestration platform reliability. Automating cloud infrastructure, deployments, and production operations across Canadian locations.
Senior DevOps Engineer owning CI/CD, Kubernetes, and cloud infrastructure for Bounteous, a global AI services firm. Automating secure, reliable platforms across the DevOps lifecycle.
DevOps Engineer owning CI/CD and app releases for a gamified sports training platform. Maintaining React Native, Expo/EAS, Supabase, Next.js, and React delivery workflows.
Site Reliability Engineer managing AWS and Kubernetes reliability for Rentsync’s rental - property software products. Leading incident response, observability, automation, and infrastructure hardening.
Senior DevOps Engineer securing Boeing Canada’s Azure, Kubernetes, and on - premise platforms. Leading CI/CD, infrastructure automation, reliability, compliance, and technical mentorship.
Site Reliability Engineer automating enterprise release orchestration and delivery operations for Sun Life. Supporting platform reliability, Kubernetes automation, and transition to a future release management solution.