Site Reliability Engineer
Harrison Clarke · Palo Alto, CA
Apply directly on Harrison Clarke’s careers site — no account needed.
About the role
We're partnering with a well-funded Series A AI infrastructure startup building a cloud-native platform designed to support highly available, distributed systems at scale.
As an early engineering hire, you'll play a key role in improving the reliability, scalability, and operational maturity of the platform. You'll work closely with software engineers to automate infrastructure, strengthen observability, and ensure production systems remain resilient as the company grows.
Key Responsibilities
- Operate and scale Kubernetes environments across AWS, Azure, or GCP.
- Build and maintain infrastructure using Terraform, Helm, and GitOps practices.
- Enhance platform reliability through monitoring, alerting, logging, and observability improvements.
- Automate deployments, scaling, recovery processes, and day-to-day operational tasks.
- Troubleshoot complex production issues across infrastructure, networking, and distributed systems.
- Improve production readiness, resilience, security, and overall platform performance.
- Partner with engineering teams to embed operational best practices into the development lifecycle.
- Support incident response, capacity planning, and disaster recovery initiatives.
Requirements
- 5-10 years' experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure.
- Strong experience running Kubernetes workloads in production.
- Hands-on experience with AWS, Azure, or GCP.
- Proven experience with Infrastructure as Code using Terraform.
- Familiarity with Helm, GitOps, Argo CD, or similar deployment tooling.
- Solid understanding of Linux, networking, DNS, load balancing, and cloud security.
- Experience building and maintaining CI/CD pipelines.
- Strong knowledge of observability tooling such as Prometheus, Grafana, and OpenTelemetry.
- Experience supporting distributed systems, including technologies such as Kafka, Redis, PostgreSQL, or similar.
- Proficiency in Go, Python, Bash, or another scripting/programming language.
- Excellent troubleshooting skills across application, infrastructure, and network layers.
Description sourced from the public LinkedIn listing — this role isn't indexed from the company's career page yet.
Skills
- Go
- Python
- Bash
- Terraform
- Kubernetes
- AWS
- Azure
- GCP
- Helm
- Prometheus
- Grafana
- Kafka
- Redis
- PostgreSQL
- Linux
Never be applicant #200 again
Every job here is indexed straight from company career pages — often hours after it opens, before it reaches the big boards. Create a free account and get your best matches in a twice-daily digest.
- Your best matches, twice a day
- No duplicates, no ghost jobs, no recruiter spam
- Every job free to browse — pay only when you apply
Free account — no card required
93 151 live jobs · 17 788 companies tracked · 5 869 added today