About the Role
Overview
As a Senior Platform Engineer on the Orbit team, you'll help build and scale GitLab Orbit, a high-impact knowledge graph data service that supports agents, analytics, and architecture-level features across GitLab.com, Dedicated, and Self-Managed deployments. You'll join a small, senior, Rust-first team and focus on how backend code behaves within a distributed, cloud-native system, making graph capabilities reliable, observable, secure, and easy for other teams and agents to use.
In this role, you'll own meaningful parts of the service end to end. You'll design and implement backend services, data workflows, and interfaces while improving multi-tenant behavior, performance, resilience, and operational readiness.
In your first year, you'll take ownership of improving how GitLab Orbit runs in production, with a focus on operational automation, observability, incident readiness, and distributed-system reliability.
What You’ll Do
- Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
- Improve the deployment, monitoring, and operations of GitLab Orbit across GitLab.com, Dedicated, and Self-Managed deployments using Kubernetes, Helm, Terraform, and cloud services from Amazon Web Services, Google Cloud Platform, or both.
- Automate recurring operational work and build tools that make deployments, upgrades, recovery, capacity management, and service maintenance safer and more efficient.
- Strengthen observability across application, data, orchestration, and infrastructure layers by improving metrics, logs, traces, dashboards, alerts, and service-level indicators.
- Investigate production issues, address underlying causes, and write backend code that handles concurrency, partial failures, retries, consistency, idempotency, performance, and multi-tenant isolation.
Requirements
Rust programming
Experience in Rust is essential for building backend services.
Cloud-native systems
Familiarity with distributed cloud-native environments is required.
Kubernetes expertise
Proficiency in Kubernetes for deployment and operations is necessary.
Operational automation
Ability to automate operational tasks to enhance efficiency.
Nice to Have
Familiarity with Amazon Web Services or Google Cloud Platform is a plus.
Knowledge of site reliability practices can be beneficial.
Benefits
Remote work
Flexible remote work options are available.
Career development
Opportunities for career growth and development are provided.