About the role
HOW TO APPLY: Apply directly through the official NVIDIA Careers portal using the application link NVIDIA is hiring a System Software Engineer for its Local AI team in Pune. The role focuses on build… Subscribe to Premium to view the full description and apply.
Responsibilities
- Partner with NVIDIA software, research, architecture, and product teams to define technical strategies for local AI on RTX and DGX systems.
- Build and optimize local AI inference stacks for RTX, RTX Pro, and DGX GPUs.
- Develop high-performance inference software focused on performance, stability, scalability, and low latency.
- Architect and develop modern inference runtimes and execution stacks.
- Work with inference frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, and TensorRT-RTX.
- Support AI workloads including LLMs, vision-language models, text-to-speech, automatic speech recognition, and diffusion models.
- Perform end-to-end optimization of AI models, data pipelines, and inference runtimes.
- Optimize software for current and next-generation GPU architectures.
- Apply model optimization techniques including quantization, pruning, sparsity, and knowledge distillation.
- Enable efficient deployment of large AI models on local and edge devices.
- Perform system-level debugging and performance optimization.
- Analyze performance-versus-accuracy trade-offs.
- Develop infrastructure for performance and accuracy sweeps.
- Analyze benchmarking results to identify performance gaps and drive fixes.
- Establish engineering guidelines for model and inference-backend bring-up.
- Ensure new models and inference backends are production ready.
Requirements
- 2+ years of relevant professional experience.
- Bachelor's, Master's, or PhD in Computer Science, Software Engineering, Mathematics, or a related field, or equivalent experience.
- Excellent C++ programming skills.
- Strong C++ debugging abilities.
- Strong understanding of data structures and algorithms.
- Strong understanding of machine learning concepts.
- Proven experience working with AI inference pipelines and applications.
- Experience with ML/DL frameworks such as Llama.cpp, vLLM, PyTorch, WinML, DXCGC, or TensorRT.
- Strong interest in inference backends and runtime internals.
- Understanding of scheduling and memory management.
- Understanding of KV-cache behavior and graph execution.
- Knowledge of quantization and hardware-aware optimization.
- Strong analytical and problem-solving skills.
- Ability to multitask effectively in a fast-moving technical environment.
- Strong written and verbal communication skills.
Benefits
- Competitive salary and comprehensive benefits package.
- Opportunity to work on cutting-edge AI and GPU technologies.
- Exposure to NVIDIA RTX and DGX platforms.
- Opportunity to work on large language models and generative AI workloads.
- Collaboration with software, research, architecture, and product teams.
- Opportunity to contribute to high-performance AI infrastructure.
- Exposure to next-generation GPU architectures.
- Opportunity to work on production-grade local and edge AI systems.
Required Skills
Frequently Asked Questions
What is the salary for System Software Engineer - Local AI?
The listed salary for System Software Engineer - Local AI at NVIDIA is ₹18-30 LPA (Estimated; company has not disclosed the official salary).
Where is this System Software Engineer - Local AI role located?
This position is based in Pune, Maharashtra, India (On-site).
What experience is required for System Software Engineer - Local AI?
Candidates are expected to have 2+ years of experience for this role.
How do I apply for System Software Engineer - Local AI at NVIDIA?
Use the Apply button on this HireDoor job page to submit your application for System Software Engineer - Local AI at NVIDIA.
What are the key benefits for System Software Engineer - Local AI?
Key benefits include: Competitive salary and comprehensive benefits package.; Opportunity to work on cutting-edge AI and GPU technologies.; Exposure to NVIDIA RTX and DGX platforms.; Opportunity to work on large language models and generative AI workloads.; Collaboration with software, research, architecture, and product teams.; Opportunity to contribute to high-performance AI infrastructure.; Exposure to next-generation GPU architectures.; Opportunity to work on production-grade local and edge AI systems..
Redirecting to sign up…