Senior Engineer optimizing deep learning inference on edge hardware for autonomous vehicles and robotics at NVIDIA. Collaborating with automotive OEMs and addressing complex optimization challenges.
Responsibilities
Address customer and partner optimization challenges by engaging directly with automotive OEMs and robotics associates to analyze, debug, and improve deep learning models on NVIDIA platforms
Own performance benchmarking by driving efforts to achieve leading results on MLPerf Edge and industry benchmarks, defining methodology and ensuring reproducibility
Evaluate emerging model architectures by analyzing DL architectures, including vision encoders, multi-modal VLMs, for compilation feasibility, memory footprint, and latency on target SOCs
Collaborate across teams by partnering with compiler, runtime, and hardware teams to connect model-level insight with platform capabilities
Deliver TensorRT and compiler-stack solutions for edge by creating and deploying inference solutions on Jetson, DRIVE, and GPU + ARM platforms for AV and robotics workloads.
Develop Proofs of Readiness (PORs) and work closely with compiler team on Torch-TRT, MLIR-TRT, and related frameworks to bridge performance gaps.
Requirements
Master’s degree or equivalent experience in Computer Science, Electrical Engineering, or a related field
12 + years of industry experience with over 8 years in deep learning model optimization, inference engineering, or neural network compilation
Adept at interpreting and reasoning about model architectures at the operator/kernel level
Over 5 years of validated expertise in embedded/edge software, experience delivering production inference solutions within power-limited, latency-sensitive deployment environments
Deep knowledge of current DL architectures: transformers, attention variants, vision encoders (ViT), multi-modal/vision-language model frameworks, and experience with diffusion models and/or state space models
Expert knowledge of GPU architecture fundamentals, CUDA, and low-level performance optimization using heterogeneous computing
Experience with TensorRT, compiler IRs, or equivalent inference optimization toolchains
Solid understanding of embedded operating system internals (QNX/Linux), memory management, C/C++, and embedded/system software concepts
Background in parallel programming (e.g., CUDA, OpenMP) and experience reasoning about memory hierarchies, data movement, and compute utilization
Demonstrated capability to collaborate directly with external partners and customers in a deep technical role, solving their workload issues, identifying performance problems, and providing solutions within production limitations.
Software Engineer building Felix’s WordPress healthcare website and design system. Developing growth tooling, APIs, analytics, experimentation, and automated QA for Canadian healthcare customers.
Software Developer building production AI agents and MCP integrations for Euna Solutions’ cloud - based public - sector budgeting software. Deploying secure LLM features with Python or TypeScript in a hybrid Oakville team.
Full Stack Engineer building JavaScript and Python applications for Ledgebrook, an InsurTech MGA modernizing specialty insurance. Developing scalable APIs, LLM tools, and underwriting technology.
Staff Engineer designing AWS integrations, microservices, and AI automation for EverCommerce’s service - commerce SaaS platform. Providing technical leadership and hands - on development across distributed teams.
Principal Agentic Engineer architecting AI - native customer - experience systems for APPLY, an agentic CX partner for global consumer and entertainment brands. Leading full - stack architecture, coding - agent delivery, cloud platforms, and AI - powered applications.
Java Full Stack Developer building mission - critical risk technology solutions for Morgan Stanley’s global financial services business. Designing web applications, APIs, databases, and event - driven systems in Montreal’s hybrid environment.
Senior Software Engineer building and maintaining OpenShift applications for RBC, Canada’s largest bank. Supporting data transformation, scalable services, production reliability, and DevOps practices.
Software Developer building AI - enabled cybersecurity solutions for Arctic Wolf’s global security services. Developing scalable cloud - native web applications with React, TypeScript, Go/Python, and AWS.
Fullstack Software Engineer building Snowflake’s cloud data engineering applications. Developing scalable ingestion, pipeline authoring, and observability experiences with Python, TypeScript, React, NodeJS, and Java.