You and the process

With a person
Anna, in-house recruiter
9 years hiring engineers
30 minutes with a real recruiter
They read your CV with you, on a call, and say where the offers are being lost.
Didn’t find what you were looking for? Tell us what to build
Be the first to open itNo views yet
NVIDIA

Senior System Software Engineer, Performance - CUDA Driver

  • Hybrid
  • 6+ years

Salary

$184,000 - 287,500/ year

AI summary

For members

The whole posting in a few lines. Sign up to read it here and on every role you open.

Sign Up to Read

Description

We are looking for senior systems software engineers to make the CUDA Driver faster, more efficient, and ready for the next generation of accelerated computing. Our team develops performance-critical CUDA Driver features and systems software improvements that help AI, deep learning, HPC, and other CUDA-powered applications realize more of the performance available from NVIDIA GPUs!

Some of the hardest performance problems emerge not within one component , but at the boundaries among systems software, CPUs, interconnects, and GPUs. In this role, you will trace those problems from real workloads through the software and hardware stack, design and ship production features , and deliver validated optimizations for current and emerging platforms.

You will combine hands-on engineering with broad technical influence, collaborating across CUDA software, hardware architecture, frameworks, applications, product, and customer-facing teams. You will lead cross-layer investigations, mentor engineers, and help set performance direction. Your work will help developers get more useful computing from NVIDIA GPUs today while helping shape the CUDA software and GPU architectures NVIDIA builds next .

What you'll be doing:

Develop and ship performance-centric CUDA Driver features and programming-model capabilities from design through validation.

Diagnose complex performance problems through workload analysis, focused measurement and modeling, cross-layer root-cause isolation, and application-level validation .

Optimize critical CUDA primitives, memory management and movement, CPU–GPU coordination, and interconnect paths for latency, throughput, bandwidth, efficiency, and scalability.

Establish performance goals for current and future platforms, characterize as new platforms come online , close software and hardware gaps, and drive performance readiness through release s .

Translate workload and platform evidence into CUDA API, programming model, system software, and future GPU architecture recommendations.

Set subsystem performance direction and mentor engineers tackling complex systems and performance challenges.

Influence technical decisions with clear performance evidence and tradeoffs and strengthen implementations through rigorous design and code reviews.

What we need to see:

A BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field—or equivalent practical experience — with at least 7 + years of relevant systems-software development experience.

Strong production C/C++ systems-programming experience, including delivery of substantial features, optimizations, or production fixes in a complex codebase.

Strong operating systems and concurrency foundations, including threads, synchronization, processes, virtual memory, and user/kernel interactions.

Strong computer architecture foundations, including processors, memory hierarchy, caching and coherence, data movement, and system interconnects.

Demonstrated success improving real software performance: measuring behavior, identifying bottlenecks , implementing optimizations , and profil ing t o prove their effectiveness .

Sound technical judgment, ownership of ambiguous problems, and clear communication across organizational and disciplinary boundaries.

Direct CUDA or GPU experience is valuable but is not required when accompanied by deep systems software , operating systems , computer architecture , and performance-engineering foundations.

Ways to stand out from the crowd:

Experience developing GPU or accelerator drivers, runtimes, kernel software, firmware, compilers, or other performance-critical low-level systems.

Experience with pre-silicon analysis, platform bring-up, performance modeling, or hardware/software co-design.

Systems-level performance experience with AI/DL, HPC, graphics, automotive, robotics, or similarly demanding workloads.

Evidence of technical inventions such as software-performance patents, novel production designs, or measurement-backed recommendations that influenced hardware revision o r future architecture.

Python or another scripting language used for focused experimentation, data analysis, or visualization.

LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

Responsibilities

  • Applications for this job will be accepted at least until September 6, 2026.
  • This posting is for an existing vacancy.
  • NVIDIA uses AI tools in its recruiting processes.
  • NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Where you’d work

Part of the week in the office

You can work from

  • United States

About the company

NVIDIA

Office in United States

Also hiring in Munich, Germany, France, Bristol, United Kingdom and 13 more places

273 of their 660 open roles are remote

Your chances

Worth a look before you spend an evening tailoring a CV for it.

  • 21 checks run
  • 2 red flags

Still hiring?

14 checks

1 red flag

How crowded?

7 checks

1 red flag

Fits Me

How well does this role fit you?

Answer a few questions or drop your CV, and every role gets a fit score with the reasons, this one first.

  • Your field
  • Level
  • Stack
  • Work model
  • Salary floor
  • Must-haves
Details12 facts · Role, Location, Compensation, Employment
Tech stack
  • Python
  • C++
  • CUDA
Seniority
Senior
Type
Full-time
Equity
Equity offered
Region
United States
Pay period
Annual
Show 6 more factsShow less

Role

Category
Development
Seniority
Senior
Experience
7+ years
Tech stack
  • Python
  • C++
  • CUDA

Location

Work model
Hybrid
Region
United States
Office
  • United States
Remote from
  • United States

Compensation

Salary
$184,000 - 287,500 / year
Pay period
Annual
Equity
Equity offered

Employment

Type
Full-time

Something wrong with this vacancy?

Similar vacancies

  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Staff Software Engineer

    Salary by agreement

    • Remote · Europe, Portugal
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer Intern

    Salary by agreement

    • Hybrid · Romania, Berlin
    • Junior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Senior Software Engineer

    Salary by agreement

    • Remote
    • Senior
  • Early Window: Be the first to open itNo views yetCloses in
    Company hidden

    Software Engineer

    Salary by agreement

    • Hybrid · Toronto
    • Junior

Share this vacancy

What's wrong with it?

The employer never sees who reported.

Reason