SRE Technical Lead- Bristol
FDM Group · Bristol, ENG, GB
Apply directly on FDM Group’s careers site — no account needed.
About the role
About The Role
As a senior individual contributor, you will provide technical leadership rather than people management. You will work closely with platform, infrastructure, engineering, and support teams to implement SRE principles, define reliability standards, reduce operational toil through automation, and establish meaningful service health measurements. You will help accelerate the organisation's transition from reactive production support towards a proactive, engineering-led reliability model, ensuring reliability is designed into services rather than addressed after incidents occur
- Partner with the Head of SRE to implement and embed the organisation's SRE strategy, operating model, and reliability standards.
- Act as a technical authority for reliability engineering, providing guidance and expertise across application, platform, and infrastructure teams.
- Define, implement, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Critical User Journeys (CUJs) to establish meaningful service reliability metrics.
- Drive adoption of SLO-based decision making, supporting teams in balancing reliability, delivery velocity, and operational risk.
- Identify opportunities to reduce operational toil through automation, including runbook automation, self-healing capabilities, deployment improvements, and recovery processes.
- Design and implement observability best practices across logging, metrics, tracing, alerting, and dashboarding.
- Support and improve incident management processes, participating in major incident response and post-incident reviews while driving root cause analysis and preventative actions.
- Work with engineering teams to improve service resilience, availability, scalability, and recoverability through proactive engineering improvements.
- Analyse reliability trends and operational data to identify systemic issues and recommend long-term solutions.
- Contribute to the development of reliability standards, frameworks, and technical roadmaps across both legacy and modern technology environments.
- Champion engineering excellence and reliability best practices through mentoring, knowledge sharing, and collaboration with technical teams.
- Support technology transformation initiatives by ensuring operational resilience and reliability requirements are embedded throughout delivery programmes.
About You
- Significant hands-on experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or a similar reliability-focused role.
- Strong background in supporting and improving business-critical production environments.
- Proven experience implementing SRE practices, including SLIs, SLOs, error budgets, and observability frameworks.
- Experience working within large, complex enterprise environments.
- Strong engineering mindset with practical experience delivering automation and operational improvements.
- Deep understanding of incident management, problem management, and operational resilience principles.
- Experience designing and implementing monitoring, logging, alerting, and distributed tracing solutions.
- Strong troubleshooting and root cause analysis skills across complex technology stacks.
- Ability to influence technical teams and stakeholders without formal line management responsibility.
- Strong communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences.
- Comfortable working in evolving environments where processes and capabilities are still being established.
- Pragmatic and outcome-focused approach to solving reliability and operational challenges.
- Experience helping establish or mature SRE capabilities within an organisation.
- Exposure to large-scale technology transformation programmes.
- Experience working with legacy platforms alongside cloud-native technologies.
- Experience standardising observability practices across multiple teams and toolsets.
- Familiarity with financial services or other highly regulated environments.
- Experience building automation solutions using scripting and Infrastructure as Code practices.
- Experience mentoring engineers or contributing to reliability communities of practice.
- Knowledge of cloud platforms, containerisation, orchestration technologies, and modern platform engineering practices.
About Us
Diversity and Inclusion
FDM Group is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, colour, religion, sex, sexual orientation, national origin, age, disability, veteran status or any other status protected by federal, provincial or local laws.
Why join us
- Career coaching, mentoring and access to upskilling throughout your entire FDM career
- Assignments with global companies and opportunities to work abroad
- Opportunity to re-skill and up-skill into new areas, develop non-linear career paths and build a skillset within your field
- Annual leave and work-place pension
Description sourced from the public Indeed listing — this role isn't indexed from the company's career page yet.
Skills
- Kubernetes
Never be applicant #200 again
Every job here is indexed straight from company career pages — often hours after it opens, before it reaches the big boards. Create a free account and get your best matches in a twice-daily digest.
- Your best matches, twice a day
- No duplicates, no ghost jobs, no recruiter spam
- Every job free to browse — pay only when you apply
Free account — no card required
93 114 live jobs · 17 642 companies tracked · 188 added today