Role Overview
We are seeking a DevOps and Reliability Lead with 6 to 9 years of experience to own ETPL's infrastructure, deployment, and reliability engineering practice across both ETPL Digital and ETPL AI. This is a senior technical leadership role with full ownership of the systems, processes, and standards that keep ETPL's products running reliably, securely, and at scale.
The DevOps and Reliability Lead will define and operate ETPL's CI/CD pipelines, cloud infrastructure, containerisation strategy, monitoring and alerting framework, and incident response practice, ensuring that engineering teams can ship with confidence and that institutional clients experience consistent, high-availability service.
ETPL's products serve operationally sensitive environments, including educational institutions running academic cycles, lending platforms processing financial transactions, and cooperative governance systems managing member data, where downtime and performance degradation carry real institutional consequences. The Lead will bring both the technical depth to architect robust infrastructure and the operational discipline to build a reliability culture across the engineering organisation.
Key Responsibilities
- Own and evolve ETPL's cloud infrastructure across AWS / Azure, including compute, networking, storage, database services, and security configurations, ensuring environments are scalable, cost-optimised, and production-ready.
- Design, build, and maintain CI/CD pipelines across ETPL's product portfolio, enabling fast, reliable, and consistent delivery of software from development through staging to production.
- Define and implement infrastructure-as-code (IaC) practices using Terraform, Pulumi, or equivalent tooling, ensuring all infrastructure is version-controlled, reproducible, and auditable.
- Own ETPL's containerisation and orchestration strategy, managing Docker-based build standards and Kubernetes cluster operations across product environments.
- Build and maintain a comprehensive observability stack, including centralised logging, metrics collection, distributed tracing, and alerting, using tools such as Prometheus, Grafana, ELK Stack, Datadog, or equivalent.
- Define and own service level objectives (SLOs) and service level indicators (SLIs) for ETPL's production systems, working with product and engineering teams to align reliability targets with institutional client expectations.
- Lead incident response across ETPL's production environment, owning the on-call framework, incident classification, escalation paths, post-incident reviews, and structured follow-through on remediation actions.
- Conduct capacity planning and performance modelling for ETPL's product infrastructure, anticipating growth requirements and ensuring systems can scale ahead of demand.
- Establish and enforce security and compliance standards across the infrastructure layer, including network security, secrets management, access control, vulnerability scanning, and data protection practices appropriate to ETPL's institutional client obligations.
- Manage database infrastructure operations, including backups, replication, failover configuration, and performance tuning, in coordination with product engineering teams.
- Drive the adoption of DevOps culture and practices across ETPL's engineering teams, including developer self-service, deployment ownership, and shared accountability for production reliability.
- Evaluate, select, and govern the use of infrastructure tooling, managed services, and third-party platform integrations across the engineering organisation.
- Lead and develop a small team of DevOps and infrastructure engineers, setting clear expectations, reviewing work quality, and building capability within the function.
Experience and Profile
- 6 to 9 years of progressive experience in DevOps, site reliability engineering, or infrastructure engineering roles.
- Demonstrated experience owning cloud infrastructure and CI/CD pipelines for production SaaS or enterprise technology products at meaningful scale.
- Proven hands-on experience with containerisation and Kubernetes in a production environment.
- Experience building and operating observability stacks and leading structured incident response processes.
- Experience implementing infrastructure-as-code and bringing discipline to infrastructure management within a growing engineering organisation.
- Prior experience in a multi-product or multi-tenant SaaS environment is strongly preferred.
- Experience managing or mentoring junior DevOps or infrastructure engineers is expected at this level.