Senior AI Performance Engineer (CUDA / GPU / NVIDIA Stack)

Contract
Remote within US
$65-70/hr
No visa sponsorship

Job description

Job Title: Senior AI Performance Engineer (CUDA / GPU / NVIDIA Stack)
Duration: Min 12+ Months
Location: 100% Remote
Job Description:
This is a hands-on engineering role, requiring deep expertise in CUDA, GPU architecture, and performance profiling.
Key Responsibilities
• Profile and optimize AI/ML workloads across multi-GPU and multi-node systems
• Identify bottlenecks across compute, memory, networking, and orchestration layers
• Optimize CUDA kernels (memory coalescing, shared memory usage, occupancy tuning)
• Improve inference performance using TensorRT, Triton, DeepStream, NeMo
• Analyze and improve latency, throughput, GPU utilization, and memory efficiency
• Work on distributed AI systems using Apache Ray, NCCL, Kubernetes GPU scheduling
• Build benchmarking frameworks and performance monitoring systems
• Collaborate with AI, DevOps, and Infrastructure teams for system-wide optimization
Required Skills
• Strong hands-on CUDA programming and GPU performance optimization
• Deep understanding of GPU architecture and memory hierarchy
• Experience with Nsight, CUDA profiling tools, performance benchmarking
• Hands-on experience with NVIDIA ecosystem (Triton, TensorRT, NeMo, DeepStream)
• Experience with distributed AI systems (multi-GPU, multi-node, NCCL, Ray)
• Experience working with AI models such as YOLO, GPT, LLaMA, Transformers
• Strong understanding of AI system performance metrics (latency, throughput, utilization)
Preferred
• Experience working at NVIDIA or similar GPU/AI infrastructure companies
• Experience with real-time video / Vision AI systems
• Experience with large-scale production AI deployments
Interview Process (Mandatory)
• Candidates will receive a technical handout 1 day before interview
• 90-minute deep-dive demo discussion (NOT theoretical)
• Candidate must explain:
o Bottleneck identification approach
o GPU optimization strategies
o System-level performance improvements


More information

Experience level

Expert and leadership (8+years)

Job skills

CUDA programming

GPU architecture knowledge

Performance benchmarking

Distributed AI systems

AI model experience

Certifications

NVIDIA Certified CUDA Developer

Performance Optimization Specialist

AI/ML Professional Certificate

Languages

English

Company overview

company-logo
Brillfy Technology Inc.

IT Services and IT Consulting·201-1,000 employees

Welcome to Brillfy, where we are passionate about creating value through technology. As a leading technology company, we believe in the transformative power of innovation and strive to unlock its full potential for businesses worldwide. Our team of skilled professionals combines expertise with a deep understanding of the values of technology. We are committed to delivering cutting-edge solutions that drive our clients' growth, efficiency, and success. At Brillfy, excellence, integrity, and customer satisfaction are at the core of everything. Our unwavering dedication to providing exceptional service and solutions sets us apart. We take the time to listen to our clients, understand their unique business needs, and collaborate closely with them to deliver tailored and effective results.