Senior Kubernetes Platform Engineer
Firmus Technologies · Sydney, New South Wales, Australia
New
senior
platform engineer
Apply on Firmus Technologies →
AI FactoryOS Operations
AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.
AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.
The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability. It also builds the shared services the estate's own operation depends on, and runs them. Operating the estate every day is what shows how the platform behaves under real load and under failure, and the function works with the engineering teams that build it to turn what it finds into permanent fixes and design improvements.
Senior Kubernetes Platform Engineer
Role Summary
Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Kubernetes Platform Engineer runs the Kubernetes estate that every product and every tenant runs on: the platform controller layer, the virtual cluster platform tenants are provisioned onto, the Kubernetes environment baselines used across the estate, the GPU integration layer, and the automated tenant onboarding and release pipeline.
This is a hands-on senior role with deep technical expertise . This role owns the platform lifecycle execution across the fleet : keeping clusters healthy and current across the fleet, keeping tenant workloads running through upgrade and failure, and restoring control planes to service when they degrade . Automation is a first-class part of the role , in the operational tooling, guarded remediation and fleet orchestration that make estate-scale operation possible, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour.
Key Responsibilities
Responsible for the reliable operation, automation and continuous improvement of the multi-tenant Kubernetes platform that Firmus' products and tenants run on, spanning every site in the estate.
Build the operational tooling, guarded remediation and orchestration that automate operations at fleet scale, and contribute operator and controller requirements, and code where agreed, to AI Infrastructure's platform backlog with the production evidence behind them.
Execute the Kubernetes cluster lifecycle across the fleet, including provisioning, patching, upgrade and decommissioning, running the deployment and upgrade mechanisms built by AI Infrastructure through the agreed staged or canary path, and holding estate version compliance and retirement coordination.
Operate and recover Kubernetes control planes carrying live tenant workload, including etcd state, certificate rotation, failed upgrades and corrupted resources.
Operate the virtual cluster platform and multi-tenant isolation patterns that tenants are provisioned onto, and the automated tenant onboarding and release pipeline that lands new tenants safely and repeatably.
Operate the GPU integration layer for Kubernetes (for example the NVIDIA GPU Operator), including device plugins, GPU scheduling and driver coordination.
Diagnose and resolve scheduling failures, CNI and CSI faults, admission rejections and resource contention from first principles, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
Run the Kubernetes baselines in production carrying the admission policy, workload identity and network policy content set by the Senior Platform Security Engineer, and hold the operational acceptance requirements those baselines have to meet before they enter production.
Provide the deepest technical expertise for Kubernetes faults across the estate, diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure , and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results.
Lead technical recovery during major Kubernetes incidents under the incident commander, drive the changes that remove repeat causes through the problem record, and share the after-hours escalation roster for the Kubernetes estate.
Skills & Experience
Required Skills
Strong skills in platform and infrastructure engineering, with 8+ years of experience overall and substantial ownership of production Kubernetes platforms in a 24/7 environment.
Deep experience operating Kubernetes at fleet scale, including cluster lifecycle, upgrades and multi-cluster management.
Strong experience with Kubernetes control plane internals, including etcd, the API server, controllers, schedulers, and certificate and credential rotation.
Strong experience writing Kubernetes controllers, operators or admission logic in a production setting.
Experience with multi-tenant or virtual cluster patterns (for example vCluster or equivalent), including tenant isolation at the Kubernetes layer.
Experience operating GPU-enabled Kubernetes, including device plugins, GPU scheduling and driver coordination (for example the NVIDIA GPU Operator).
Strong skills in infrastructure automation, infrastructure-as-code and GitOps practices (for example OpenTofu or Terraform, Ansible, Argo CD), with change delivered through peer review, automated testing and progressive rollout.
Strong experience with scripting or programming for operational automation and tooling, such as Go, Python or Bash.
Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, post-incident review, and the production of runbooks that others can execute successfully.
Solid understanding of Kubernetes security fundamentals, including admission control, workload identity and network policy.
Clear technical judgement and communication skills, with the ability to produce documentation, design notes and escalations that other engineers can act on.
Preferred Experience
Experience operating Kubernetes for GPU or HPC workloads at scale.
Experience with automated tenant or customer onboarding pipelines in a multi-tenant platform.
Experience with vendor Kubernetes distributions or reference architectures for accelerated computing.
Familiarity with DPU or SmartNIC-based networking as it relates to Kubernetes CNI design.
Experience contributing to open-source Kubernetes ecosystem projects.
A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training.
Expected Outcomes
Fleet-wide cluster upgrades executed through the staged path with no unplanned tenant-visible outage, and version compliance held across the estate.
Tenant onboarding running as a routine operation rather than a project.
Control plane recovery tested against live-equivalent conditions and executable by an engineer who did not write the procedure.
Kubernetes escalations falling as operational automation and guarded remediation take on the common faults.
Platform defects and operability gaps evidenced into the platform engineering backlog and closed permanently, with repeat causes falling.
Location & Reporting
Location : Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
On-call: The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function.
Posted 2026-09-21