Develop and maintain hardware abstraction layers and runtime interfaces for NVIDIA’s computing platforms. Collaborate with cross-functional teams to enhance reliability and performance.
Responsibilities
Extend and maintain hardware abstraction layers and core system libraries used across the platform.
Design and implement drivers, runtimes, and data movement/aggregation pipelines supporting workload execution.
Build and maintain runtime interfaces for launching, monitoring, and managing workloads.
Improve platform reliability through automation, error reporting, diagnostics, and operational tooling.
Debug and resolve complex sequencing, initialization, and runtime issues across multi-component systems.
Partner cross-functionally with hardware engineering, compiler teams, and data center operations to bring features from prototype to production.
Support new platform bring-up and NPI (New Product Introduction) efforts for new boards and silicon.
Contribute to engineering excellence through documentation, tooling improvements, code reviews, and knowledge sharing.
Requirements
A Masters Degree in Computer Science, Computer Engineering, Electrical Engineering, related STEM field or equivalent experience.
5+ years of relevant work experience
Strong proficiency in modern C++ (design, implementation, debugging, and performance considerations).
Experience designing, maintaining, and refactoring software libraries and APIs with long-term support in mind.
Comfort working in large, multi-repository or multi-component codebases with layered dependencies.
Demonstrated ability to lead or drive triage of difficult reliability issues and produce clear root-cause analysis.
Ability to clearly communicate software architecture and design tradeoffs, including using diagrams and written design docs.
Low-level platform software experience (e.g., firmware/boot flows, RTOS, BMCs/MCUs, RISC-V, or closely related system software).
Linux systems experience that includes driver or kernel-adjacent interfaces (e.g., VFIO or similar subsystems).
Hardware bring-up and/or system triage experience (fault analysis, system diagnostics, or validation support in lab environments).
Principal Agentic Engineer architecting AI - native customer - experience systems for APPLY, an agentic CX partner for global consumer and entertainment brands. Leading full - stack architecture, coding - agent delivery, cloud platforms, and AI - powered applications.
Java Full Stack Developer building mission - critical risk technology solutions for Morgan Stanley’s global financial services business. Designing web applications, APIs, databases, and event - driven systems in Montreal’s hybrid environment.
Senior Software Engineer building and maintaining OpenShift applications for RBC, Canada’s largest bank. Supporting data transformation, scalable services, production reliability, and DevOps practices.
Software Developer building AI - enabled cybersecurity solutions for Arctic Wolf’s global security services. Developing scalable cloud - native web applications with React, TypeScript, Go/Python, and AWS.
Fullstack Software Engineer building Snowflake’s cloud data engineering applications. Developing scalable ingestion, pipeline authoring, and observability experiences with Python, TypeScript, React, NodeJS, and Java.
Frontend Engineer building intuitive web applications for Cambio’s real estate decarbonization software platform. Collaborating across engineering, product, and design to deliver accessible interfaces and automated testing.
Software Engineer building partner integrations for Liftoff’s AI - powered mobile adtech platform. Designing scalable solutions, troubleshooting issues, and optimizing tools for mobile app marketers and publishers.
Senior Full - Stack Software Engineer building AI - driven cloud insurance platforms for Guidewire. Developing Java microservices, front - end configuration tools, and scalable SaaS services.