Role Style & Setup
This is a hands-on consulting engagement, rather than a pure review or advisory assignment.
You'll work alongside the existing Infrastructure team, balancing operational support with targeted improvement work.
You'll be expected to get under the bonnet of the environment, understand how it currently operates, identify areas holding it back, and turn those findings into practical improvements.
Key Responsibilities
- Assess and optimise the existing HPC environment, including compute, scheduling, storage and network performance
- Provide senior-level Linux administration and troubleshooting across the infrastructure
- Investigate performance, availability and capacity issues using monitoring and observability tooling
- Deliver infrastructure improvements through defined work packages and sprint-based activity
- Improve automation, reliability and operational processes across the platform
- Document findings, recommendations and technical changes for the wider team
Skills & Experience Required
- Strong commercial experience administering RHEL/Linux in demanding technical environments
- Proven HPC experience, ideally working with Slurm and distributed compute environments
- Experience with containers and/or virtualisation
- Strong automation capability using Ansible, Python and Bash
- Comfortable diagnosing complex infrastructure and performance problems
- Able to work independently and communicate technical findings clearly
Please note, we’re looking for someone who is available to start in the next 2 weeks.