About the Role
About Gray Swan
Gray Swan is on a mission to empower the world to use AI safely and securely. We evaluate AI models for the leading frontier labs along with building real-time threat detection and adaptive adversarial red teaming agents for teams deploying AI.
We're a team of approximately 50 people, well-funded, growing quickly. Our work directly influences how the world deploys AI agents and systems at scale.
The Role
Gray Swan is looking for an Infrastructure Engineer to build and scale the systems that power our AI security platform. You'll design the backend services, distributed infrastructure, and cloud architecture that enable Gray Swan to build AI systems and for customers to safely deploy frontier AI models at scale.
This role is ideal for an engineer who enjoys solving infrastructure challenges across reliability, scalability, observability, and performance. You'll work closely with machine learning engineers, product engineers, and security researchers to ensure our platform remains fast, resilient, and secure as we grow.
You'll have significant ownership over foundational systems and the opportunity to influence technical direction in a rapidly evolving AI startup.
What You’ll Do:
- Design, build, and maintain highly available backend services and distributed systems that power Gray Swan's AI security platform.
- Own cloud infrastructure across Kubernetes, AWS, networking, storage, and compute to ensure reliable production environments.
- Build scalable APIs, internal platform services, and infrastructure tooling that improve developer productivity and system reliability.
- Improve system observability through logging, metrics, tracing, dashboards, and automated alerting.
- Optimize performance, latency, and infrastructure costs while maintaining reliability and security.
- Partner closely with machine learning, security, and product engineering teams to deliver production-ready infrastructure for AI workloads.
Who You Are:
- 5+ years of experience building backend infrastructure or distributed systems in production environments.
- Strong programming skills in C/C++, Go, Python, Rust, or Java.
- Experience operating services on Kubernetes and modern cloud platforms such as AWS, GCP, or Azure.
- Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures.
- Experience designing APIs, microservices, asynchronous systems, and event-driven architectures.
- Comfortable debugging complex production issues and improving reliability through automation and operational excellence.
- Passionate about writing clean, maintainable code and building infrastructure that other engineers love using.
- Excited to work in a fast-moving startup with significant ownership and ambiguity.
Bonus Points If You Have:
- Experience supporting machine learning or LLM infrastructure.
- Familiarity with infrastructure-as-code tools.
- Experience with Kafka, Redis, PostgreSQL, ClickHouse, or similar distributed data systems.
- Experience building internal developer platforms or platform engineering tooling.
- Knowledge of cloud security, infrastructure hardening, or zero-trust architectures.
- Previous experience at a high-growth startup or building products from zero to one.
- Interest in AI safety, cybersecurity, or adversarial machine learning.
If you don’t have 100% of these, you should still seriously consider applying. We care more about what you can do than your credentials.
You’ll Thrive Here If You:
- You thrive on ownership and solving hard problems. You're energized by ambiguity, enjoy building systems from the ground up, and take pride in delivering reliable solutions from design through production.
- You think at scale. You enjoy designing resilient infrastructure, optimizing performance, and building systems that are secure, observable, and built to grow.
- You collaborate across disciplines. You work effectively with machine learning, security, and product engineering teams.
Requirements
Backend infrastructure experience
5+ years of experience building backend infrastructure or distributed systems in production environments.
Programming skills
Strong programming skills in C/C++, Go, Python, Rust, or Java.
Cloud platform experience
Experience operating services on Kubernetes and modern cloud platforms such as AWS, GCP, or Azure.
Networking knowledge
Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures.
Nice to Have
Experience supporting machine learning or LLM infrastructure.
Familiarity with infrastructure-as-code tools.
Experience with Kafka, Redis, PostgreSQL, ClickHouse, or similar distributed data systems.
Benefits
Ownership
Significant ownership over foundational systems.
Fast-paced environment
Opportunity to work in a rapidly evolving AI startup.