About this role
NVIDIA is seeking a Senior System Software Engineer to design, build, and operate scalable Software Defined Networking (SDN) solutions for its AI Clouds, which support GPU-accelerated workloads such as hyperscale multi-node training, inference, and cloud gaming. This role covers the full lifecycle of the SDN stack, from development to production reliability engineering.
Responsibilities
- Design and develop next-generation multi-tenant cloud SDN control and data plane software using OVS, OVN, and OpenFlow.
- Build Infrastructure-as-a-Service virtual network orchestration using gRPC and REST to support BMaaS, VMaaS, and Kubernetes.
- Drive upstream contributions to OVN-Kubernetes and related open-source projects.
- Develop software for network observability, including monitoring, telemetry, and performance analysis.
- Operate and support OVS-OVN based SDN solutions in large-scale AI Cloud environments.
- Maintain CI/CD pipelines (GitLab) and implement GitOps approaches for secure cloud infrastructure integration.
- Collaborate with SRE and DevOps teams on production readiness and incident management.
Required Qualifications
- BS/MS in Computer Science or a related technical field, or equivalent experience.
- 5+ years of experience in software development for large-scale distributed environments.
- Expert-level knowledge of OVN, OVS, OpenFlow, and modern network protocols.
- Strong programming skills in C and Go, with advanced scripting in Bash and Python.
- Deep knowledge of Kubernetes and practical experience with CNIs (OVN-Kubernetes).
- Hands-on experience with Infrastructure-as-Code (Ansible, Terraform, ArgoCD, Flux) and CI/CD pipelines.
- Experience developing secure, high-performance services using gRPC and REST with TLS.
- Strong knowledge of datacenter routing, switching, and Linux host/VM networking.
Preferred Qualifications
- Contributions to open-source projects (OVS, OVN, Kubernetes networking).
- Experience with hardware acceleration (GPU, DPU) for networking.
- Practical experience with major cloud providers (AWS, Azure, GCP) and hybrid/multi-cloud deployments.
- SRE/DevOps expertise in incident management and service reliability.
- Experience with observability tools like Prometheus, Grafana, Jaeger, OpenTelemetry, and ELK.
Skills for this role
SDNOVSOVNOpenFlowCGoBashPythonKubernetesgRPCRESTAnsibleTerraformArgoCDFluxGitLabLinux NetworkingDatacenter RoutingSwitching