Senior HPC and Cloud Engineer - AI/ML Compute Infrastructure
Senior HPC andamp; Cloud Engineer - AI/ML Compute Infrastructure Oxford | Hybrid (3 days in office) | Competitive, DOE Sentinel is recruiting for several senior/staff-level engineers to design, build and operate a hybrid GPU compute environment, combining on-prem HPC clusters with public cloud infrastructure for large-scale AI research workloads. Responsibilities: Build and operate high-performance GPU training/inference clusters, including scheduling, isolation and automated life cycle management Design high-throughput data paths across compute and storage, including parallel filesystems (eg Lustre) Benchmark and resolve performance bottlenecks across compute, network and orchestration layers Implement observability, resilience and security controls for a compliance-conscious research environment Work with research and applied ML teams to forecast GPU/storage capacity and streamline experimentation pipelines Requirements: Experience with HPC/GPU clusters, including a strong understanding of GPU architecture, high-speed networking and distributed training performance Cloud platform experience (Azure, GCP, AWS or other) Experience with containerisation/Kubernetes Working knowledge of IaC and CI/CD (eg Terraform, Argo CD) ..... full job details .....
Perform a fresh search...
-
Create your ideal job search criteria by
completing our quick and simple form and
receive daily job alerts tailored to you!