5+ years of backend, platform, infrastructure, or distributed systems engineering experience in a production environment.
Deep hands-on TypeScript experience. We run TypeScript services on Bun, and we care about people who can design clean, reliable backend systems rather than only write application code.
Strong understanding of queues, workers, retries, rate limits, idempotency, distributed locks, webhooks, and failure handling.
Experience operating systems under real production pressure: incidents, provider outages, customer escalations, noisy alerts, slow dependencies, and incomplete data.
Comfort with cloud-native infrastructure. Our stack includes GCP, Kubernetes, Cloud Tasks, Pub/Sub, RabbitMQ, Redis, Postgres, Firestore, ClickHouse, Prometheus, Grafana, and OpenTelemetry.
Ability to debug across layers: application code, logs, metrics, queues, provider APIs, databases, network behaviour, and deployment state.
Product-minded judgement. You can separate internal detail from customer-safe explanation, and you understand when a “small infra change” has customer, security, or cost implications.
High ownership and low ego. We want people who take responsibility for ambiguous systems, ask sharp questions, and make the team better through clear thinking and kindness.
A bias toward durable fixes. You do not stop at “restart the pod” if the real issue is retry amplification, missing backpressure, an unsafe parser assumption, or a provider limit we do not model.