Senior Platform Engineer building and operating GitLab Orbit, a Rust-based knowledge graph service for GitLab’s DevSecOps platform. Improving distributed-system reliability, observability, cloud infrastructure, and data workflows.
Responsibilities
Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment
Improve deployment, monitoring, and operations across GitLab.com, Dedicated, and Self-Managed deployments using Kubernetes, Helm, Terraform, and AWS or GCP
Automate operational work and build tools for safer deployments, upgrades, recovery, capacity management, and service maintenance
Strengthen observability through metrics, logs, traces, dashboards, alerts, and service-level indicators
Collaborate with SRE teams on incident response, on-call readiness, runbooks, and troubleshooting workflows
Investigate production issues and address underlying causes involving concurrency, partial failures, retries, consistency, idempotency, performance, and multi-tenant isolation
Build and improve the graph query engine, SDLC and code indexing pipelines, cloud storage integrations, API, and MCP surfaces
Design reliable, scalable, and cost-aware data workflows using S3, ClickHouse, NATS, and Siphon
Own changes from technical design through rollout and iteration, documenting constraints and trade-offs
Collaborate asynchronously with product, data, infrastructure, security, delivery, AI, and SRE teams
Requirements
Experience designing, building, and operating production backend services with strong Rust skills or clear evidence of ability to ramp up in a Rust-first, performance-sensitive codebase
Experience with distributed-system design, including concurrency, failure handling, consistency, messaging, data partitioning, scalability, and multi-tenant isolation
Hands-on knowledge of AWS, GCP, or both, including cloud networking, identity and access management, compute, storage, and object storage such as Amazon S3
Experience deploying and troubleshooting applications on Kubernetes with Helm
Experience contributing to repeatable, reviewable infrastructure changes using Terraform or similar infrastructure-as-code tools
Experience improving reliability, observability, maintainability, and on-call readiness of backend services
Experience diagnosing issues across application, data, orchestration, and infrastructure layers
Strong system design skills, including architectural decisions, documenting constraints, and aligning trade-offs
Ability to work autonomously in ambiguous environments by identifying problems, driving solutions, and taking ownership
Ability to learn and apply new languages, tools, and frameworks such as Ruby, Go, or TypeScript and Vue
Excellent written communication and asynchronous collaboration skills
GitLab expects all team members to incorporate AI into their daily workflows
Benefits
Benefits to support your health, finances, and well-being
Flexible Paid Time Off
Team Member Resource Groups
Equity Compensation & Employee Stock Purchase Plan
Senior AI Platform Developer building scalable AI services and LLM workflows for MaintainX’s industrial work execution platform. Improving reliability, observability, performance, and cost efficiency.
Infrastructure team lead building and operating Spare’s GCP and Kubernetes platform for on - demand transit. Leading developers while improving reliability, security, AI SRE, and cloud cost efficiency.
Senior AI Platform Developer building reusable, secure AI - agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Driving Azure - based platform capabilities across orchestration, evaluation, observability, and governance.
AI Platform Developer building Azure - based agent infrastructure for Petal, a Canadian healthcare orchestration and billing company. Creating secure, observable, governed AI services for product teams.
Kubernetes/DevOps Engineer operating Kubernetes infrastructure for a petabyte - scale social media platform. Building distributed systems and machine - learning workloads for Capgemini Engineering’s client.
Lead AI Platform Engineer securing and operating EQ Bank’s enterprise AI platforms. Driving Azure engineering, observability, automation, governance, and production readiness.
Platform Software Engineer supporting Invoice Simple, EverCommerce’s invoicing SaaS for small businesses. Building cloud infrastructure, CI/CD pipelines, and reliable production systems.
Software Engineer building cloud infrastructure and deployment systems for Invoice Simple, EverCommerce’s invoicing platform for small businesses. Improving reliability, observability, and production operations.
Senior cybersecurity engineer securing L3Harris’s mission - critical platform - management systems for military and cruise ships. Hardening Linux, networks, software, and infrastructure across the product lifecycle.