Senior Platform Engineer building resilient Go and Kubernetes platforms for Virtasant’s cloud-native products. Owning reliability, infrastructure as code, observability, and developer delivery tools.
Responsibilities
Design, build, and operate production Kubernetes clusters, including networking, workload isolation, and multi-region topologies
Work with Kubernetes internals, resource quotas, scheduling, NetworkPolicy enforcement, and custom controllers or operators
Implement and operate service mesh capabilities, including mTLS, workload authentication and authorization, and traffic management
Optimize containerized workloads for performance, cost, and resource efficiency
Write, refactor, and maintain production Go services, controllers, and middleware
Build HTTP, REST, and gRPC interfaces for internal engineering teams
Write unit and integration tests
Lead incident response, investigate root causes, and author postmortems
Troubleshoot production systems using logs, metrics, traces, and profiling tools
Diagnose and resolve distributed-system performance and reliability problems
Define and drive SLOs and actionable alerting
Own infrastructure as code and maintain reusable modules
Build and improve CI/CD and GitOps delivery workflows
Balance developer velocity with reliability, security, and compliance
Build and maintain metrics, dashboards, alerting policies, and distributed tracing
Partner with product, security, and infrastructure teams on requirements and architecture
Contribute to design reviews, set technical direction, and mentor engineers
Requirements
6+ years of professional experience in software, platform, infrastructure, or site reliability engineering
Significant experience operating production distributed systems
Demonstrated experience building and operating production Kubernetes platforms
Production experience writing Go and ability to use Go as the primary day-to-day language
Experience taking ambiguous system designs through production
Degree in Computer Science, Engineering, or a related field, or equivalent practical experience
Strong understanding of Kubernetes internals, including CNI networking, NetworkPolicy, resource management, and cluster behavior under load
Hands-on experience with a service mesh such as Istio, Envoy, or Linkerd, plus mTLS and workload identity
Solid Linux fundamentals, including cgroups and resource management
Infrastructure as code at scale using Terraform or equivalent
Production experience with GCP, AWS, or Azure
Experience with Prometheus, Grafana, OpenTelemetry, and query languages such as PromQL
Production experience with relational databases, including PostgreSQL or managed Postgres-compatible services, and replication and failover
Docker and container tooling experience
Strong debugging and performance profiling skills
Strong analytical and problem-solving ability
Clear written and verbal communication
Ability to work independently within a distributed team
Comfort working in a fast-moving, highly technical environment
Senior Software Engineer building scalable Python and AWS infrastructure for d1g1t’s institutional wealth management platform. Driving implementation, deployment, troubleshooting, and large - dataset performance.
Sr. Engineer position focusing on vulnerability remediation and cyber security for Genpact in Canada. Collaborating with teams and leading the cyber security team to meet business objectives.
Director of Infrastructure Engineering at Outschool leading infrastructure and platform team towards innovation and stability. Fostering a culture of technical excellence and ownership in a remote setup across U.S. and Canada.
IT Engineer responsible for IT infrastructure design and support in automotive software. Collaborating with teams to ensure efficient technology operations.
Senior Infrastructure Engineer leading infrastructure operations and stability for Buffer's social media tools. Focusing on modernization and developer tooling within a fully remote team.
Senior Data Engineer designing and implementing scalable data lakehouse infrastructure at TRM Labs. Collaborating across teams to optimize data workflows and ETL processes.