Cloud Software Engineer - Observability Platform
Boost your chances before you apply.
ClickHouse
US
Summary
Design, build, and operate distributed systems that power observability across ClickHouse Cloud. You will own the reliability, performance, and cost-efficiency of telemetry pipelines while participating in on-call rotations.
Job Description
ClickHouse is looking for an experienced engineer to join our Observability team. We build and operate the telemetry platform that powers both internal monitoring and the observability features our customers rely on. Our systems ingest trillions of events per day with sustained throughput in the tens of millions per second. Engineers on the team are hybrid software, systems, and infrastructure engineers who ensure this platform is reliable, scalable, and efficient. We work closely with product and infrastructure teams and play a key role in major engineering initiatives across the company.
We're looking for someone who thrives in fast-paced environments, isn't afraid to get hands-on during incidents, and knows when to automate the pain away. While experience in roles like Software Engineer, SRE, Systems Engineer, or DevOps is valuable, we care most about your problem-solving skills and mindset. If you enjoy tackling complex challenges across system design, infrastructure, automation, and incident response—while helping us scale with confidence—you’ll fit right in.
What you’ll do
-
Design, build, and operate distributed systems that power observability across ClickHouse Cloud
-
Own reliability, performance, and cost-efficiency of our telemetry pipeline and storage systems
-
Take part in the on-call rotation and help drive root-cause resolution and long-term fixes
-
Build tooling and automation to eliminate repetitive operational work
-
Help shape the roadmap for observability by identifying bottlenecks and scaling challenges
-
Collaborate with other engineering teams to improve their observability posture
-
Contribute to design discussions, architecture reviews, and mentor teammates
What we’re looking for
-
Strong bias for action and ownership — you ship, fix, and improve systems proactively
-
Great production debugging skills and a problem-solving mindset
-
Strong communication skills; comfortable working in a remote, async-friendly team
-
Experience balancing system performance, reliability, and cost
-
Ability to iterate quickly: build MVPs, collect feedback, and improve continuously
Requirements
-
5+ years building and running production systems at scale
-
Proficiency in Golang
-
Experience with Kubernetes, Helm, ArgoCD, and Terraform or similar IaC tools
-
Comfortable working with at least one major cloud provider (AWS, GCP, Azure)
-
Experience with OpenTelemetry, Prometheus, Grafana, or similar tools
-
Experience with ClickHouse preferred
#LI-Remote
ClickHouse is a fast, open-source columnar database built for real-time data processing and analytics at scale. ClickHouse Cloud delivers the query speed and concurrency that applications demanding instant insight from large volumes of data require. As AI agents become more embedded in software, generating higher query volumes at tighter latency, ClickHouse provides a high-throughput, low-latency engine purpose-built for that workload.
Founded
2021
Company size
501-1,000 employees
Industry
Software Development
Org type
Privately Held
Headquarters
Palo Alto, California
ClickHouse is a fast, open-source columnar database built for real-time data processing and analytics at scale. ClickHouse Cloud delivers the query speed and concurrency that applications demanding instant insight from large volumes of data require. As AI agents become more embedded in software, generating higher query volumes at tighter latency, ClickHouse provides a high-throughput, low-latency engine purpose-built for that workload.
Founded
2021
Company size
501-1,000 employees
Industry
Software Development
Org type
Privately Held
Headquarters
Palo Alto, California