About the Role
About Level AI
Level AI is on a mission to turn every customer interaction into a strategic advantage. Our AI-native platform helps enterprises transform contact centers from cost centers into engines of customer intelligence, operational efficiency, and business growth. By combining advanced AI with deep domain understanding of customer experience, Level AI empowers teams to unlock actionable insights, automate workflows, and deliver more consistent, higher-quality support across the customer journey.
Headquartered in Mountain View, California, Level AI is a Series C company backed by leading investors including Battery Ventures and ENIAC. Our platform leverages Large Language Models and Custom Small Language Models (SLMs) to power AI Agents across the entire CX journey—customer-facing agents, agent-assist, and backend automation—along with deep conversation analytics for QA, coaching, and insights.
About The Role
The Senior SRE will be positioned at the intersection of backend engineering, infrastructure operations, and FinOps. The role is explicitly broader than a traditional DevOps engineer and explicitly more hands-on than a pure architect.
What You'll Do
- Infrastructure cost efficiency and FinOps: Own the continued reduction of Kubernetes overprovisioning, drive right-sizing programs, and maintain the cost telemetry that backend teams use to make decisions.
- GPU throughput optimization: Run a structured experimentation program on on-premise GPU clusters, partnering with AI service owners.
- Backend enablement: Build the tooling, dashboards, and processes that let backend teams from other groups own their own cost and reliability budgets.
- Reliability instrumentation: Ensure that surface area is captured properly for both cost-at-scale and reliability.
- Selective security workstreams: Take on a defined slice of the active security work.
Who You Are
This role explicitly requires 4-5 years of hands-on systems experience. We are not looking for someone who will lean entirely on AI tooling to discover what to do; we are looking for someone who already knows what to ask, and can use AI tooling as a force multiplier on top of that judgement.
- Backend engineering depth: Production experience in Python, Go/Rust, comfortable owning services end to end.
- Kubernetes at scale: Scheduler behavior, resource requests/limits, HPA/VPA, node pool design.
- Cloud and on-premise infrastructure: GCP fluency, IaC (Terraform), CI/CD.
- GPU workload understanding: Familiarity with throughput profiling, batching, KV-cache behavior.
- Observability and reliability: Metrics, traces, logs, SLOs.
- FinOps mindset: Demonstrated history of converting infrastructure choices into measurable cost outcomes.
- Security baseline: Able to take on platform-security workstreams.
Compensation
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are made by our team.
Requirements
Hands-on systems experience
This role requires 4-5 years of hands-on systems experience.
Backend engineering depth
Production experience in Python, Go/Rust, and comfortable owning services end to end.
Kubernetes expertise
Experience with Kubernetes at scale, including scheduler behavior and cost-aware autoscaling.
Cloud infrastructure knowledge
Fluency in GCP, IaC (Terraform), and CI/CD processes.
Observability skills
Ability to implement metrics, traces, logs, and SLOs effectively.
Nice to Have
Familiarity with throughput profiling and GPU utilization metrics.
A demonstrated history of converting infrastructure choices into measurable cost outcomes.
Ability to handle platform-security workstreams.
Benefits
Remote work options
Flexible work arrangements to support a healthy work-life balance.
Learning budget
Opportunities for professional development and continuous learning.
Health insurance
Comprehensive health insurance plans for employees.