About the Role
About
You'll own the data backbone of Interfere. Every signal the product reasons about (events, traces, logs, runtime behavior, code, session data) flows through systems you'll build, store, and query. The product's intelligence is only as good as the data underneath it, and that data is only useful if it's accurate, fast, and affordable at scale. That's your job. You'll work close to the AI/ML and product teams, designing the pipelines, schemas, and storage everything else sits on, and making the early architectural decisions the next ten engineers will inherit.
What You'll Do
Concretely, this looks like building:
- High-throughput ingestion that handles billions of product events without dropping data, slowing down, or melting the infra bill.
- Real-time stream processing for detection and diagnosis, where the work has to be correct and low-latency at the same time.
- Storage and query systems (ClickHouse; columnar, time-series, vector, whatever the problem calls for) that stay fast as customer data grows by orders of magnitude.
- The indexing and retrieval infrastructure that lets agents and models find the right context at the moment they need it.
- Schema, taxonomy, and data-quality systems that hold up as event shapes evolve and new product surfaces appear.
- The cost, observability, and reliability layer for our own data systems, because observability for our customers starts with observability of ourselves.
Who You Are
You've built and operated production data infrastructure at meaningful scale, with real throughput, real cost pressure, and real consequences when it breaks. You're fluent in distributed-systems tradeoffs: streaming vs batch, consistency vs latency, full fidelity vs sampling, storage vs compute. You pick the right answer for the situation rather than the one you read most recently. You take an ambiguous data or infrastructure problem, define the next useful step, and ship without waiting for a fully specified plan. You treat cost as a feature. A system that works at 1x and burns the company at 100x isn't finished. You can explain pipeline behavior, failure modes, and tradeoffs clearly enough that engineers, AI researchers, and PMs can make the right call quickly.
Nice to Have
- Deep experience with high-throughput streaming or stream-processing systems (Kafka, Flink, Kinesis, Materialize).
- Background in observability, telemetry, or session-replay data systems.
- Built vector or hybrid retrieval infrastructure for AI/ML use cases.
- Comfort across the full data lifecycle: ingestion, transformation, storage, query, retention, deletion.
- Fluent in Go, Rust, Python, TypeScript, or whichever tool the throughput actually demands.
Requirements
Data infrastructure experience
You've built and operated production data infrastructure at meaningful scale.
Distributed systems knowledge
You're fluent in distributed-systems tradeoffs and can choose the right approach for the situation.
Problem-solving skills
You can take ambiguous data problems and define the next useful step.
Cost management
You treat cost as a feature and ensure systems are efficient.
Communication skills
You can explain complex pipeline behaviors and tradeoffs clearly.
Nice to Have
Deep experience with high-throughput streaming or stream-processing systems.
Background in observability, telemetry, or session-replay data systems.
Built vector or hybrid retrieval infrastructure for AI/ML use cases.
Comfort across the full data lifecycle from ingestion to deletion.
Fluent in Go, Rust, Python, TypeScript, or whichever tool the throughput demands.