Site Reliability Manager

Karsun Solutions LLC · Herndon, VA, US

Full-timePublished Jul 23, 2026

Apply directly on Karsun Solutions LLC’s careers site — no account needed.

About the role

Overview:
Summary
We are seeking a highly skilled and experienced Site Reliability Manager to join our team to ensure the reliability, scalability, and performance of our systems and services. You will lead a team of engineers focusing on three core pillars: Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate must reside in DMV area and be available to work on site in office or customer locations in this area. Must have demonstrated experience in having performed this role for at least 3 years.
Responsibilities:
What You'll Be Doing:
  • Team Leadership: Lead an 8-20 person service delivery team (Service Support specialist, DevSecOps, and Site Reliability engineers), mentoring them to foster a culture of learning and innovation.
  • Pillar 1: Application Reliability (Core Focus): Take joint ownership of production reliability, standardize observability and error handling using Datadog, develop SLOs and KPIs to measure system performance, and conduct incident post-mortems and root cause analyses.
  • Pillar 2: DevSecOps: Define and implement best practices for infrastructure as code, deployment automation, and drive vulnerability management and resolution.
  • Pillar 3: Platform Lifecycle Management: Oversee the end-to-end platform lifecycle, collaborate with cross-functional teams to design scalable and fault-tolerant architectures, and drive continuous improvement initiatives to enhance system efficiency.
Qualifications and Education:
Required Qualifications:
  • Bachelor’s degree in Computer Science, Engineering, or a related field; Master's degree preferred.
  • 10+ years of experience in a similar role managing a team of site reliability engineers and delivering in the AWS cloud platform.
  • 5+ years of experience supporting operations and maintenance for cloud-native applications in production that are fault-tolerant, self-healing, scalable, and highly available.
  • Deep understanding of the AWS cloud computing platform and containerization technologies (e.g., Docker, Kubernetes).
  • Strong knowledge of infrastructure as code tools (e.g., Terraform, Ansible, ArgoCD) and CI/CD pipelines.
  • Experience with Datadog as the primary logging, monitoring, and observability platform (alongside AWS Cloudwatch).
  • Excellent communication and interpersonal skills, with the ability to collaborate effectively with cross-functional teams.
  • Strong problem-solving and analytical skills, with a keen attention to detail.
  • Ability to obtain and maintain a Public Trust clearance.
  • Certifications such as AWS Certified DevOps Engineer are a plus.
Compensation:
The proposed salary range for this role is $****** to $******* USD. The salary range provided is a good faith estimate representative of all experience levels. Karsun considers several factors when extending an offer, including but not limited to, the role, function and associated responsibilities, a candidate’s work experience, location, education/training, and key skills.

Description sourced from the public Indeed listing — this role isn't indexed from the company's career page yet.

Skills

  • AWS
  • Docker
  • Kubernetes
  • Terraform
  • Ansible
  • ArgoCD
  • Datadog
  • GitHub Actions

Never be applicant #200 again

Every job here is indexed straight from company career pages — often hours after it opens, before it reaches the big boards. Create a free account and get your best matches in a twice-daily digest.

  • Your best matches, twice a day
  • No duplicates, no ghost jobs, no recruiter spam
  • Every job free to browse — pay only when you apply
Get my matched jobs

Free account — no card required

93 117 live jobs · 17 641 companies tracked · 146 added today

Similar jobs

Site Reliability Manager — Karsun Solutions LLC · Real Job Offers