Senior DevOps Engineer
Apply directly on Agile Resources, Inc.’s careers site — no account needed.
About the role
Salary: Up to $130K (depending on experience) + bonus
Work Authorization: Due to client requirements, only candidates who do not require sponsorship now or in the future will be considered.
Benefits:
- Unlimited PTO
- Full healthcare coverage for employees + family (medical, dental, vision, life, and supplemental insurances)
- Short- and Long-Term Disability (STD/LTD)
- HSA & FSA options
This Senior DevOps Engineer role owns the reliability, security, and day-to-day operations of a production Kubernetes platform and the delivery pipelines that promote containerized applications into production. Success in this position looks like stable clusters, predictable releases through CI/CD and GitOps, and rapid, root-cause-driven incident resolution. You will set practical platform standards and continuously improve automation, observability, and operational readiness.
Responsibilities:
- Own production Kubernetes clusters end-to-end, including architecture decisions, upgrades, capacity planning, and ongoing administration.
- Drive incident response and troubleshooting across Kubernetes networking, storage, scheduling, and workload behavior; lead root-cause analysis and implement durable fixes.
- Build and maintain secure Kubernetes platform foundations, including RBAC, secrets handling, network policies, pod security controls, and image governance practices.
- Develop and operate CI/CD workflows for containerized applications using GitHub Actions or equivalent, ensuring repeatable builds, testing gates, and controlled promotions.
- Implement and run GitOps-based deployments with tooling such as Argo CD or equivalent to enable auditable, consistent configuration and release management.
- Create and maintain Helm charts and Kubernetes manifests (YAML) that support standardized, reusable application delivery patterns.
- Establish and enforce platform standards for release practices, Docker image creation, versioning, and operational readiness; publish clear runbooks and documentation.
- Improve platform reliability through automation, proactive monitoring/alerting, and performance tuning of cluster resources and workloads.
- Partner with engineering teams to harden deployment patterns, reduce operational risk, and streamline the path from pull request to production rollout.
Required Skills:
- 5+ years of experience in DevOps, SRE, Platform Engineering, or a closely related role supporting production systems.
- Deep, hands-on ownership of a production Kubernetes platform, including cluster administration, troubleshooting, architecture, networking, storage, workloads, resources, and security.
- Experience owning CI/CD for containerized applications using GitHub Actions/Workflows or equivalent, with strong understanding of release controls and promotion strategies.
- Hands-on GitOps experience using Argo CD or equivalent, including managing deployment state through version-controlled configuration.
- Strong working knowledge of Docker, Helm, and Kubernetes YAML, including building images and packaging/deploying applications reliably.
- Proven ability to diagnose and resolve complex production issues and implement corrective actions that prevent recurrence.
Preferred Skills:
- Direct OpenShift administration experience, including operating managed OpenShift environments (ideally on IBM Cloud).
- AWS and EKS experience, including architecture, migration planning/execution, or steady-state production operations.
- Experience supporting Java application delivery through automated pipelines, including Maven-based build and deploy workflows.
- Python experience for automation, operational tooling, or reliability improvements.
- Experience operating in regulated or compliance-sensitive environments (e.g., healthcare or financial services).
- Demonstrated success establishing or materially improving platform standards, including infrastructure-as-code, release practices, Docker image practices, documentation/runbooks, and engineering guardrails.
Description sourced from the public LinkedIn listing — this role isn't indexed from the company's career page yet.
Skills
- Kubernetes
- Docker
- Helm
- GitHub Actions
- Terraform
Never be applicant #200 again
Every job here is indexed straight from company career pages — often hours after it opens, before it reaches the big boards. Create a free account and get your best matches in a twice-daily digest.
- Your best matches, twice a day
- No duplicates, no ghost jobs, no recruiter spam
- Every job free to browse — pay only when you apply
Free account — no card required
93 119 live jobs · 17 642 companies tracked · 136 added today