About the Role
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.
The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.
An overview of this role
As a Senior Platform Engineer on the Orbit team, you'll help build and scale GitLab Orbit, a high-impact knowledge graph data service that supports agents, analytics, and architecture-level features across GitLab.com, Dedicated, and Self-Managed deployments. You'll join a small, senior, Rust-first team and focus on how backend code behaves within a distributed, cloud-native system, making graph capabilities reliable, observable, secure, and easy for other teams and agents to use.
In this role, you'll own meaningful parts of the service end to end. You'll design and implement backend services, data workflows, and interfaces while improving multi-tenant behavior, performance, resilience, and operational readiness. You'll also contribute to the cloud infrastructure and deployment patterns needed to run the service effectively.
In your first year, you'll take ownership of improving how GitLab Orbit runs in production, with a focus on operational automation, observability, incident readiness, and distributed-system reliability. You'll also contribute to key areas such as the graph query engine, indexing pipelines, cloud storage integrations, and multi-tenant behavior. Through thoughtful system design, better tooling, clear runbooks, and shared context, you'll help reduce single points of failure and improve how we build, deploy, and operate the service.
What you’ll do
- Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
- Improve the deployment, monitoring, and operations of GitLab Orbit across GitLab.com, Dedicated, and Self-Managed deployments using Kubernetes, Helm, Terraform, and cloud services from Amazon Web Services, Google Cloud Platform, or both.
- Automate recurring operational work and build tools that make deployments, upgrades, recovery, capacity management, and service maintenance safer and more efficient.
- Strengthen observability across application, data, orchestration, and infrastructure layers by improving metrics, logs, traces, dashboards, alerts, and service-level indicators to track availability, error rates, and incident response time while collaborating with site reliability engineering teams to improve incident response, on-call readiness, runbooks, and troubleshooting workflows.
- Investigate production issues, address underlying causes, and write backend code that handles concurrency, partial failures, retries, consistency, idempotency, performance, and multi-tenant isolation.
- Build and improve the graph query engine, software development lifecycle (SDLC) and code indexing pipelines.
Requirements
Rust programming
Proficiency in Rust is essential for building backend services.
Cloud-native systems
Experience with distributed cloud-native environments is required.
Kubernetes expertise
Knowledge of Kubernetes for deployment and operations is necessary.
Operational automation
Ability to automate operational tasks to enhance efficiency.
Nice to Have
Familiarity with Amazon Web Services or Google Cloud Platform is a plus.
Experience in site reliability engineering practices is beneficial.
Benefits
Remote work
Flexible remote work options are available.
Learning opportunities
Access to continuous learning and development resources.