NVIDIA — DevOps / SRE jobs
50 open roles
NVIDIA
✕ Clear filter
DevOps and Site Reliability Engineers build and run the infrastructure that keeps software alive at scale — CI/CD, Kubernetes, cloud platforms, and observability. As AI systems move to production, SRE and platform engineering are more critical than ever. This board aggregates open DevOps and SRE roles daily from public ATS feeds, linked to the original posting.
-
Senior Site Reliability EngineerNVIDIA · Israel, Yokneam
-
Senior Deep Learning Software Infrastructure EngineerNVIDIA · US, CA, RemoteRemote · US
-
Senior Site Reliability Engineering - StorageNVIDIA · Israel, Yokneam
-
Senior Site Reliability Engineer in Test, SDETNVIDIA · China, Shanghai
-
Harness and Platform Engineer, AI Safety and Security EngineeringNVIDIA · 6 Locations
-
Senior Network Site Reliability EngineerNVIDIA · 3 Locations
-
Senior ASIC Front End Infrastructure EngineerNVIDIA · 4 Locations
-
Principal Site Reliability EngineerNVIDIA · US, CA, Santa ClaraNew
-
Compute Platform Engineer, License Scheduling - EDA InfrastructureNVIDIA · 5 Locations
-
Senior Silicon Reliability EngineerNVIDIA · US, CA, Santa Clara
-
Senior Infrastructure Engineer - Infrastructure Security and Core ServicesNVIDIA · US, CA, Santa ClaraNew
-
Senior Site Reliability Engineer, DGX CloudNVIDIA · 2 Locations
-
Senior FC Verification & Simulation Infrastructure EngineerNVIDIA · Israel, Yokneam
-
Staff Site Reliability Engineer - AI Platform RuntimeNVIDIA · US, CA, Santa ClaraNew
-
Senior Network Reliability Engineer - DGX CloudNVIDIA · 2 Locations
-
DevOps Engineer, DOCANVIDIA · 2 Locations
-
Senior Compute Platform Engineer, LSF - EDA InfrastructureNVIDIA · 3 Locations
-
Principal Engineer, Cloud Site Reliability EngineeringNVIDIA · US, CA, Santa Clara
-
Director, DOCA DevOps and IntegrationNVIDIA · Israel, Yokneam
-
Senior Site Reliability Engineering, StorageNVIDIA · India, Bengaluru
-
Senior Infrastructure Engineer - AI, Automation, Observability and MonitoringNVIDIA · US, CA, Santa ClaraNew
-
Senior Site Reliability Engineer - HPCNVIDIA · 3 Locations
-
Senior AI Infrastructure Engineer - EDA InfrastructureNVIDIA · 5 Locations
-
Senior Site Reliability Engineer - StorageNVIDIA · US, CA, Santa Clara
-
Senior Platform Engineer, Network InfrastructureNVIDIA · 4 Locations
-
DevOps EngineerNVIDIA · 2 Locations
-
Senior DevOps Engineer - E-commerceNVIDIA · India, Pune
-
Senior Storage Platform EngineerNVIDIA · US, CA, Santa ClaraNew
-
Senior Staff Site Reliability Operations Technical LeadNVIDIA · US, NC, DurhamNew
-
Deep Learning Compiler CI/Infrastructure EngineerNVIDIA · 2 Locations
-
Senior Staff Site Reliability EngineerNVIDIA · India, Bengaluru
-
Senior HPC DevOps Engineer, NCSNVIDIA · 2 Locations
-
Senior DevOps EngineerNVIDIA · 3 Locations
-
Manager, Infrastructure Engineering and DevOpsNVIDIA · Israel, Yokneam
-
Emulation Front-End Infrastructure EngineerNVIDIA · US, CA, Santa Clara
-
Senior Staff Site Reliability Engineer - Compute Core EngineeringNVIDIA · 2 Locations
-
Solutions Architect, AI Factory Infrastructure DevOpsNVIDIA · 4 Locations
-
Senior Staff Platform EngineerNVIDIA · US, CA, Santa Clara
-
Senior Reliability EngineerNVIDIA · US, CA, Santa Clara
-
Senior Staff Site Reliability OperationsNVIDIA · US, WA, SeattleNew
-
Senior Network Infrastructure EngineerNVIDIA · 5 Locations
-
Senior Director, Reliability EngineeringNVIDIA · US, CA, Santa Clara
-
Solutions Architect - DevOpsNVIDIA · Australia, RemoteRemote
-
Software Manager, Robotics Platform EngineeringNVIDIA · China, ShanghaiNew
-
Systems Quality and Reliability Engineer - LPUNVIDIA · US, CA, Santa Clara
-
Senior Staff Client Platform EngineerNVIDIA · 2 Locations
-
Senior DevOps Engineer, Cloud Simulation InfrastructureNVIDIA · US, CA, Santa Clara
-
Director, Marketing Platform Engineering and PersonalizationNVIDIA · US, CA, Santa Clara
-
Senior Data Infrastructure Engineer, AI PerformanceNVIDIA · 2 Locations
-
DevOps and Automation Engineer – ICPENVIDIA · Israel, YokneamNew
DevOps / SRE — FAQ
- What does a DevOps Engineer do?
- Automates software delivery (CI/CD), manages cloud infrastructure, and bridges development and operations — typically using tools like Terraform, Kubernetes, Docker, GitHub Actions, and major cloud platforms.
- DevOps vs. SRE?
- Site Reliability Engineering (Google's model) applies software engineering to operations — focusing on reliability, error budgets, and eliminating toil. DevOps is broader, emphasizing automation and culture. In practice the roles overlap heavily.
- What skills do DevOps / SRE roles require?
- Linux, cloud (AWS/GCP/Azure), containerization (Docker/Kubernetes), IaC (Terraform/Pulumi), CI/CD pipelines, monitoring (Prometheus/Datadog), and scripting (Python/Bash).