Senior Production Reliability Engineer – Cloud Infrastructure & Automation
Westbury Partners Sydney, AustraliaSenior Production Reliability Engineer – Cloud Infrastructure & Automation
Drive reliability across critical production systems by engineering scalable cloud infrastructure, automating operations, strengthening observability, leading incident response, and enabling engineering teams to deliver safer, resilient services.
What You'll Do:
- Design, operate, and continuously improve highly available, scalable, secure production infrastructure.
- Build automation for provisioning, deployments, configuration, and operational workflows.
- Establish meaningful SLOs, SLIs, reliability metrics, and error budgets.
- Strengthen monitoring, logging, alerting, tracing, and overall observability.
- Troubleshoot complex production issues and participate in on-call incident response.
- Lead post-incident reviews and turn recurring problems into lasting improvements.
- Improve deployment safety, rollback strategies, change management, and release processes.
- Identify infrastructure bottlenecks, reliability risks, technical debt, and opportunities for optimisation.
- Support capacity planning, disaster recovery, performance testing, and resilience initiatives.
- Create infrastructure-as-code, runbooks, documentation, and operational best practices.
Your responsibilities will include:
- You’ll take ownership of critical production environments while partnering with software, security, and infrastructure teams. You’ll automate repetitive operational work, improve system resilience, reduce toil, and help engineering teams adopt reliability-focused practices.
- You’ll work across cloud platforms, Linux, containers, Kubernetes, infrastructure-as-code, CI/CD, observability, networking, databases, and distributed systems. Your work will directly contribute to improved availability, faster incident recovery, safer deployments, scalability, and infrastructure efficiency.
Why Join Us:
- This is an opportunity to solve challenging production infrastructure problems while having a direct influence on how systems are designed, deployed, monitored, and operated.
- You’ll be empowered to build meaningful automation, improve engineering practices, strengthen reliability, and create tools that make life easier for development teams while delivering a better experience for customers.
About You:
- You’re an experienced reliability, infrastructure, DevOps, or production engineer with strong Linux and cloud expertise. You’re comfortable writing code or scripts in Python, Go, Bash, or similar languages and have hands-on experience with technologies such as Kubernetes, Docker, Terraform, and modern CI/CD platforms.
- You understand networking, DNS, HTTP/TLS, load balancing, databases, distributed systems, monitoring, and observability. You’re analytical, collaborative, calm during incidents, and naturally driven to automate problems rather than repeatedly work around them.
- You take ownership, communicate clearly, enjoy solving complex technical challenges, and continuously look for ways to make production systems safer, simpler, and more reliable.
#SiteReliabilityEngineering
#SRE
#DevOps
#CloudInfrastructure
#Kubernetes
#Terraform
#InfrastructureAsCode
#CloudEngineering
#Observability
#ProductionEngineering
#Automation
#IncidentResponse
#PlatformEngineering
#ReliabilityEngineering
#CloudNative
Please Stay Alert to Potential Scams
We would like to remind you that eFinancialCareers is a job board and does not conduct hiring or ask for payment or any financial details as part of the job application process.
If you receive any suspicious messages claiming to be from us or a hiring company, we urge you not to click on any links and not to reply to the message itself.
Instead, please report the message to our support team at support@efinancialcareers.com.
It is advisable to always verify job offers directly with the hiring company.