Senior Site Reliability Engineer – AUS Region
Boost your chances before you apply.
Nucleus Security
Summary
You will maintain highly available, secure, and performant production systems while leveraging AI-assisted tools to accelerate troubleshooting and incident response. Additionally, you will manage Kubernetes clusters, infrastructure as code, and CI/CD pipelines to reduce operational toil and improve system reliability.
Job Description
Senior Site Reliability Engineer – AUS Region
Are you looking for more in life than just building another web app? Does upending cyber security resonate with you? We're a growth stage cyber security startup that is paving the way forward for how vulnerability management is run in large enterprise organizations. For our customers, vulnerability management has always been a game of catch up, with limited asset coverage and manual processes. Nucleus’ core goal is to build a fast and scalable platform that solves these problems and many more so that vulnerability management isn't just possible, it's easy. We're looking for a passionate Senior Site Reliability Engineer to join our growing team of engineers AUS region.
What You Will Do
-
Maintain Reliable, Secure, and AI-Assisted Production Operations
Keep production systems highly available, secure, patched, and performant. Use AI-assisted tooling to accelerate troubleshooting, identify risks, analyze incidents, and improve operational response.
-
Build and Maintain Kubernetes, Cloud, and DevOps Infrastructure
Own and improve Kubernetes clusters, containerized workloads, Infrastructure as Code, CI/CD pipelines, and cloud infrastructure. Leverage AI-assisted development and automation tools to improve delivery speed, configuration quality, and operational consistency.
-
Build Observability and Automation That Reduces Toil
Improve monitoring, alerting, logging, dashboards, and automated remediation to identify issues earlier and reduce repetitive operational work. Apply AI and intelligent automation to correlate signals, surface anomalies, assist with root-cause analysis, and automate common SRE workflows.
Expectations of Your Experience
Required Qualifications:
-
8+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, or related field.
-
Strong hands-on experience with cloud platforms, including AWS, GCP, Azure, and/or OpenShift (OCP).
-
Deep experience with Kubernetes, containers, and production container orchestration.
-
Experience building and maintaining highly available, scalable, and secure production infrastructure.
-
Strong experience with Infrastructure as Code, preferably Terraform, and configuration/automation tools such as Ansible.
-
Strong scripting and automation skills using Python, Bash, or similar languages.
-
Experience building and maintaining CI/CD pipelines using GitHub, GitLab, Bitbucket, or similar platforms.
-
Strong experience with observability and monitoring platforms such as Prometheus, Grafana, Loki, CloudWatch, or equivalent tools.
-
Experience with incident response, root-cause analysis, production troubleshooting, and reliability engineering practices.
-
Experience using AI-assisted engineering tools to improve infrastructure automation, troubleshooting, documentation, code generation, or operational workflows.
-
Ability to identify opportunities where AI and automation can reduce operational toil, improve signal detection, and accelerate incident investigation.
-
Strong understanding of Linux, networking, security, cloud architecture, and distributed systems.
-
Ability to provide technical leadership, mentor engineers, and help drive a culture of automation, reliability, and continuous improvement.
-
Must be located in the AUS region and have citizenship in the country you currently reside.
Minimal Requirements
-
Minimum 8 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.
-
Strong hands-on experience with cloud providers like AWS, GCP, Azure, and OpenShift (OCP).
-
Strong experience with Kubernetes, Infrastructure as Code (terraform, tofu, CloudFormation), and automation (Python, bash, PowerShell).
-
Proven experience supporting highly available production systems, including observability, incident response, troubleshooting, and reliability improvements.
Preferred Qualifications:
-
Experience integrating LLMs or AI-enabled tools into engineering or operational workflows.
-
Familiarity with AI-assisted log analysis, anomaly detection, incident summarization, or root-cause investigation.
-
Experience building internal automation or tooling that combines APIs, scripting, infrastructure data, and AI models.
-
Understanding of how to use AI safely in production engineering environments, including data security, access controls, validation, and human review.
Additional Information
At Nucleus we are committed to achieving excellence in our field by combining diversity, collaboration, teamwork, and pride in our work. All qualified applicants will receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or disability.
Nucleus is the exposure management platform enterprise and government security leaders rely on to reduce cyber risk at scale. By unifying asset, vulnerability, and threat data, Nucleus automatically prioritizes and mitigates critical exposures, cutting high-priority risk by 60% within three months
Founded
2019
Company size
51-200 employees
Industry
Computer and Network Security
Org type
Privately Held
Headquarters
Sarasota, Florida
Nucleus is the exposure management platform enterprise and government security leaders rely on to reduce cyber risk at scale. By unifying asset, vulnerability, and threat data, Nucleus automatically prioritizes and mitigates critical exposures, cutting high-priority risk by 60% within three months
Founded
2019
Company size
51-200 employees
Industry
Computer and Network Security
Org type
Privately Held
Headquarters
Sarasota, Florida