About the Role
Job Description
As a Staff Software Engineer (IC4) within the Systems team, you will act as a principal technical architect and champion for our foundational infrastructure. You will be responsible for defining the technical direction of our core distributed systems, solving complex scaling paradigms, and engineering high-throughput, low-latency platforms that power our global SaaS ecosystem. You will operate as a strategic partner to engineering leadership, balancing bleeding-edge system design with long-term operational excellence.
Impact You Will Create
- Define Foundational Architecture: Architect, pioneer, and own next-generation, multi-tenant cloud-native distributed infrastructure from zero to scale, establishing design patterns that serve as the blueprint across the entire engineering organization.
- Guarantee Tier-0 Production Excellence: Set the technical standard for mission-critical systems, engineering multi-region fault tolerance, advanced disaster recovery, and deep observability to guarantee 99.999% availability at scale.
- Act as a Multi-Org Force Multiplier: Guide and grow senior technical talent across multiple teams. Raise the technical bar through foundational design reviews, steering committee contributions, and cross-functional technical alignment.
- Pioneer AI-Driven Platform Evolution: Strategic integration and optimization of autonomous AI systems, LLM infrastructure, and next-gen AI tools to radically accelerate developer velocity and enhance platform capabilities.
Roles & Responsibilities
- Strategic Infrastructure Lifecycle Ownership: Drive the long-term roadmap for our core systems. Own the full lifecycle of platform and infrastructure initiatives, transforming highly ambiguous business visions into concrete, scalable technical realities.
- High-Performance Distributed Systems Architecture: Design, build, and optimize next-generation backend services, high-throughput streaming pipelines, and robust microservices capable of handling millions of concurrent connections.
- System Bottleneck & Performance Engineering: Lead deep-dive architectural refactoring and performance tuning across the entire stack, optimizing kernel/OS interactions, network protocols, memory management, and database query execution paths.
- Global Production Governance: Champion systemic resilience by architecting advanced service meshes, rate-limiting frameworks, circuit breakers, multi-region failover protocols, and automated chaos engineering practices.
- Cross-Functional & Executive Alignment: Partner closely with Principal Engineers, SRE Directors, and Product Executives to translate product roadmaps into long-term infrastructure capacity and architectural readiness.
- Engineering Standards & Governance: Define, document, and enforce company-wide engineering standards for distributed systems, advanced design patterns, and systemic security controls.
Qualifications
Skills & Competencies
- Advanced Systems Engineering: Master-level fluency in backend systems languages (e.g., Go, Java, Rust, C++), concurrent programming, asynchronous execution frameworks, and low-latency microservices design.
- Advanced System Design & Distributed Systems: Deep understanding of distributed systems theory (e.g., CAP theorem, consensus protocols like Raft/Paxos, replication, partitioning, and sharding).
- Data Infrastructure Rigor: Expert-level knowledge of data modeling, transaction isolation levels, and tuning distributed storage systems (Relational, NoSQL, NewSQL, distributed caches like Redis, and message buses like Kafka/Pulsar).
- AI & Next-Gen Infrastructure Fluency: Hands-on experience architecting vector databases, optimizing hardware accelerators (GPUs/TPUs), or building backend infrastructure optimized for AI model hosting and routing.
- Systems Telemetry & Observability: Expert capability in designing structured logging, distributed tracing, and metric aggregation models across vast distributed fleets.
Requirements
Advanced Systems Engineering
Master-level fluency in backend systems languages and low-latency microservices design.
Distributed Systems Knowledge
Deep understanding of distributed systems theory and consensus protocols.
Data Infrastructure Expertise
Expert-level knowledge of data modeling and tuning distributed storage systems.
AI Infrastructure Experience
Hands-on experience with architecting vector databases and optimizing hardware accelerators.
Nice to Have
Ability to guide and grow senior technical talent across multiple teams.
Experience in architectural refactoring and performance tuning.
Benefits
Remote Work
Flexible work arrangements to support work-life balance.
Health Insurance
Comprehensive health coverage for employees.