Infrastructure Engineer focusing on high availability and systems reliability in high throughput data services. Collaborate with infrastructure and product teams while operating Kubernetes systems.
Responsibilities
Collaborate deeply with our infrastructure and product teams to enforce org-wide practices for emitting and collecting telemetry across a wide range of services, both internal and external facing.
Own and operate the Kubernetes infrastructure of the observability team.
Work within the Observability team to ensure industry-standard deployment and reliability practices are used.
Orchestrate and scale systems such as VictoriaMetrics, OpenTelemetry Collector, and Vector.
Requirements
5+ years of experience in a Site Reliability Engineering role
Experience operating and supporting clustered applications in production environments
Hands-on experience deploying and managing applications in Kubernetes (k8s) environments
Working knowledge of PostgreSQL, including administration, performance tuning, and troubleshooting
Proficiency with at least one Infrastructure as Code (IaC) tool (e.g., Terraform, Pulumi, OpenTofu, or equivalent)
Experience with telemetry tooling such as OpenTelemetry, VictoriaMetrics, Grafana, Prometheus.
Experience with AWS services is a plus
Strong documentation and communication skills is a plus
Senior Infrastructure Engineer building secure, programmable infrastructure for Turnkey’s autonomous - economy platform. Automating Kubernetes, Terraform, and Go systems across security - critical environments.
Senior DevOps Engineer operating AWS, Kubernetes, and blockchain infrastructure for Startale’s onchain finance products. Owning reliability, security, deployment, and production operations for StartaleApp and Strium.
Senior Cloud Infrastructure Engineer scaling AWS networking and storage for Mecka AI’s robotics and embodied AI data infrastructure. Leading reliability, security, transfer, and cost optimization.
Senior infrastructure developer owning AWS environments, internal applications, and enterprise integrations. Supporting Benevity’s technology platform that enables companies and employees to take social action.
Senior developer owning AWS infrastructure, security, and integrations for Benevity’s technology platform. Supporting internal applications that help companies and employees take social action.
Infrastructure Engineer building secure, automated Azure environments for PLATO, Canada’s largest Indigenous - owned software testing and technology services company. Applying Terraform, IaC, and DevOps practices across cloud infrastructure.
Infrastructure support engineer handling incidents, troubleshooting, monitoring, and service requests for Genpact’s enterprise technology services. Supporting hybrid operations on rotational night shifts in Montreal.
Senior Infrastructure Engineer rebuilding Stream’s real - time platform as it migrates from AWS to GCP. Owning Kubernetes, PostgreSQL scaling, cloud efficiency, and reliability.