NVIDIA — DevOps / SRE jobs
50 open roles
NVIDIA
✕ Clear filter
DevOps and Site Reliability Engineers build and run the infrastructure that keeps software alive at scale — CI/CD, Kubernetes, cloud platforms, and observability. As AI systems move to production, SRE and platform engineering are more critical than ever. This board aggregates open DevOps and SRE roles daily from public ATS feeds, linked to the original posting.
-
Senior System Software Engineer - Windows DevOps and Test LabsNVIDIA · India, Pune
-
Build and DevOps Engineer for CompilersNVIDIA · 2 Locations
-
Senior DevOps Software Engineer - Chip DesignNVIDIA · 2 Locations
-
Software DevOps Engineer, NetworkingNVIDIA · US, CA, Santa Clara
-
Senior Backend Platform Engineer - Profiling ServicesNVIDIA · 2 LocationsNew
-
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server SystemsNVIDIA · US, CA, Santa Clara
-
Senior DFX Infrastructure EngineerNVIDIA · US, CA, Santa Clara
-
Senior DevOps EngineerNVIDIA · India, Pune
-
Senior Platform Engineer, Network Infrastructure - DGX CloudNVIDIA · 6 Locations
-
Senior AI Infrastructure EngineerNVIDIA · 4 Locations
-
Senior DevOps Engineer, GeForce NOWNVIDIA · 2 Locations
-
Senior Technical Program Manager, Cloud Infrastructure NPINVIDIA · 2 Locations
-
Senior Product Quality and Reliability EngineerNVIDIA · Israel, Yokneam
-
Harness and Platform Engineer, AI Safety and Security EngineeringNVIDIA · 6 LocationsNew
-
Senior System Reliability EngineerNVIDIA · US, CA, Santa Clara
-
Senior Network Site Reliability EngineerNVIDIA · 3 Locations
-
Senior FPGA Prototyping Platform EngineerNVIDIA · US, CA, Santa Clara
-
Senior ASIC Front End Infrastructure EngineerNVIDIA · 4 Locations
-
Senior Backend Platform EngineerNVIDIA · US, CA, Santa Clara
-
Senior Silicon Reliability EngineerNVIDIA · US, CA, Santa Clara
-
Senior AI Platform EngineerNVIDIA · Israel, YokneamNew
-
Senior Site Reliability Engineering - StorageNVIDIA · India, Bengaluru
-
Senior Site Reliability EngineerNVIDIA · 2 Locations
-
Senior Solution Architect, Cloud Infrastructure - DevOpsNVIDIA · India, Mumbai
-
Electro Optics Reliability EngineerNVIDIA · Israel, Yokneam
-
Senior Network Reliability Engineer - DGX CloudNVIDIA · 2 Locations
-
Senior Solution Architect, Cloud Infrastructure-DevOpsNVIDIA · India, Mumbai
-
Director, DOCA DevOps and IntegrationNVIDIA · Israel, Yokneam
-
Senior Research-Ops & DevOps EngineerNVIDIA · 2 Locations
-
DevOps Engineer, DFX SoftwareNVIDIA · India, Bengaluru
-
Senior Site Reliability Engineer - HPCNVIDIA · 3 Locations
-
Senior Deep Learning Infrastructure EngineerNVIDIA · 3 LocationsNew
-
DevOps EngineerNVIDIA · 2 Locations
-
Senior DevOps Engineer - E-commerceNVIDIA · 2 Locations
-
Senior Solutions Architect, Cloud Infrastructure and DevOps - NVISNVIDIA · Japan, RemoteRemote
-
Deep Learning Compiler CI/Infrastructure EngineerNVIDIA · 2 Locations
-
Senior Site Reliability Engineer, AIOPsNVIDIA · US, CA, Santa Clara
-
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFXNVIDIA · US, CA, Santa Clara
-
Senior Infrastructure Engineer – Bazel Remote ExecutionNVIDIA · US, CA, Santa ClaraRemote · US
-
Senior Engineer, Build and DevOps - ADINVIDIA · 6 Locations
-
Senior Reliability EngineerNVIDIA · US, CA, Santa Clara
-
Senior Systems Software Engineer, Cloud Infrastructure and DevelopmentNVIDIA · US, CA, Santa ClaraNew
-
Senior Software and DevOps Engineer, DOCA VerificationNVIDIA · 2 Locations
-
Automation Infrastructure EngineerNVIDIA · Israel, Yokneam
-
Site Reliability Engineer - Hardware InfrastructureNVIDIA · US, CA, Santa Clara
-
Senior Director, Reliability EngineeringNVIDIA · US, CA, Santa Clara
-
Senior Data Center Infrastructure EngineerNVIDIA · 4 LocationsNew
-
Systems Quality and Reliability Engineer - LPUNVIDIA · US, CA, Santa Clara
-
Senior Staff Client Platform EngineerNVIDIA · 2 Locations
-
Senior DevOps Engineer, Cloud Simulation InfrastructureNVIDIA · US, CA, Santa Clara
DevOps / SRE — FAQ
- What does a DevOps Engineer do?
- Automates software delivery (CI/CD), manages cloud infrastructure, and bridges development and operations — typically using tools like Terraform, Kubernetes, Docker, GitHub Actions, and major cloud platforms.
- DevOps vs. SRE?
- Site Reliability Engineering (Google's model) applies software engineering to operations — focusing on reliability, error budgets, and eliminating toil. DevOps is broader, emphasizing automation and culture. In practice the roles overlap heavily.
- What skills do DevOps / SRE roles require?
- Linux, cloud (AWS/GCP/Azure), containerization (Docker/Kubernetes), IaC (Terraform/Pulumi), CI/CD pipelines, monitoring (Prometheus/Datadog), and scripting (Python/Bash).