Sr Site Reliability Engineer
signoz · Remote
About SigNoz SigNoz is an open-source observability platform that helps modern engineering teams monitor, debug, and optimize their applications with deep visibility into metrics, traces, and logs — all in one place. We're built natively on OpenTelemetry and offer both self-hosted and cloud options, so teams can run observability the way they want, without vendor lock-in. We are growing fast and building core developer infra products. And we are not fooling around: 27,000+ GitHub stars 800+ customers 7,000+ members in our Slack community Role: Sr Site Reliability Engineer (SRE) We're looking for an SRE to own the reliability, scalability, and operability of the SigNoz cloud platform. You'll keep a petabyte-scale observability system fast and dependable — making sure the people who trust us to watch their systems can always trust ours. The platform team handles infra, scalability of SaaS, ingest pipelines, staging environments, automation, and the operational backbone of the product. This is a deeply hands-on role for someone who understands what actually breaks in production at scale — and enjoys fixing it for good. What we're looking for Kubernetes at scale — not just "I've deployed to k8s," but real fluency with the nuances and gotchas: resource tuning, autoscaling behavior, networking, stateful workloads, upgrades, and the failure modes that only show up under load Working knowledge of ClickHouse — operating it, tuning queries, and understanding its behavior at scale — is a strong plus Knowledge of Golang is a plus (most of our stack and tooling is in Go) Familiarity with OpenTelemetry and running large-scale data ingest pipelines is a plus What you'll work on You'll work with a high-caliber team across areas like: Reliability of the SigNoz cloud platform: SLOs/SLIs, error budgets, incident response, and on-call practices that don't burn people out Scaling the ingest path — making it robust to bursts while maintaining data freshness SaaS auto-scalability and capa