waypointjobs

Trading Interview

GPU Performance Engineer | Experienced Hire

northern, KY

Check who can apply and the requirements below before continuing.

Availability awaiting confirmation

We are waiting for a fresh update from the source. This page preserves the last received job details; current availability is not confirmed.

Job description

Job Type Full-time

Posted 5 months ago

The role

Job description

Overview

We are looking for aGPU Performance Engineerto build highly optimized CUDA kernels for low-latency inference. This role is focused on workloads where off-the-shelf runtimes and vendor libraries do not fully exploit the structure of the model, and where custom kernels, memory layouts, and execution strategies can deliver meaningful gains.

You will work closely with quantitative researchers and engineers to understand model structure,identifycomputational bottlenecks, and turn mathematical ideas into production-grade GPU implementations. You will use your understanding of GPU hardware to help shape models that are both mathematically effective and efficient to run. The problems span compact neural networks, tree-based models, and other structured inference workloads where latency, throughput, and efficiency all matter.

This role is a strong fit for someone who enjoys low-level optimization, performance analysis, and translating abstract models into hardware-efficient code.

What you'll do

Design, implement, and optimize custom CUDA kernels for latency-critical inference workloads

Develop fine-grained GPU implementations tailored to specific model structures

Analyze quantitative research models and computational bottlenecks to identify opportunities for parallelization and hardware-efficient execution

Collaborate directly with quantitative researchers to translate mathematical models into high-performance compute pipelines

Optimize end-to-end inference performance through kernel tuning, memory-layout design, execution strategy, I/O optimization, and precision tradeoffs

Profile and benchmark GPU performance

Improve latency and throughput in production inference systems

Contribute to GPU architecture decisions and performance best practices

What we're looking for Strong proficiency in writing and optimizing CUDA kernels

Solid programming experience in C/C++ (preferred)

Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs

Ability to reason about numerical stability, precision, performance tradeoffs, and how model design choices affect hardware efficiency

Strong problem-solving skills and comfort working with low-level systems

Preferred qualifications

PhD in mathematics, physics, computer science, engineering, or related quantitative field

Strong background in linear algebra, probability, numerical methods, or scientific computing

Experience working with quantitative research teams or financial models

Demonstrated ability to improve real-world inference performance beyond baseline framework or library implementations

Familiarity with PTX-level behavior, tensor core utilization, or architecture-specific tuning

Exposure to ONNX Runtime, TensorRT, Triton, TVM, or similar systems

Exposure to neural networks, tree-based models (e.g., LightGBM), state space models (e.g., Mamba architectures), and experience with kernel fusion, custom operators, model compilation, or graph-level optimization

About Susquehanna

Susquehanna is a global quantitative trading firm powered by scientific rigor, curiosity, and innovation. Our culture is intellectually driven and highly collaborative, bringing together researchers, engineers, and traders to design and deploy impactful strategies in our systematic trading environment. To meet the unique challenges of global markets, Susquehanna applies machine learning and advanced quantitative research to vast datasets in order to uncover actionable insights and build effective strategies. By uniting deep market expertise with cutting-edge technology, we excel in solving complex problems and pushing boundaries together.

Strong proficiency in writing and optimizing CUDA kernels

Solid programming experience in C/C++ (preferred)

Deep understanding of GPU architecture, including memory hierarchy, SIMT execution, occupancy, and latency/throughput tradeoffs

Ability to reason about numerical stability, precision, performance tradeoffs, and how model design choices affect hardware efficiency

Strong problem-solving skills and comfort working with low-level systems

Preferred qualifications

PhD in mathematics, physics, computer science, engineering, or related quantitative field

Strong background in linear algebra, probability, numerical methods, or scientific computing

Experience working with quantitative research teams or financial models

Demonstrated ability to improve real-world inference performance beyond baseline framework or library implementations

Familiarity with PTX-level behavior, tensor core utilization, or architecture-specific tuning

Exposure to ONNX Runtime, TensorRT, Triton, TVM, or similar systems

Exposure to neural networks, tree-based models (e.g., LightGBM), state space models (e.g., Mamba architectures), and experience with kernel fusion, custom operators, model compilation, or graph-level optimization

About Susquehanna

Susquehanna is a global quantitative trading firm powered by scientific rigor, curiosity, and innovation. Our culture is intellectually driven and highly collaborative, bringing together researchers, engineers, and traders to design and deploy impactful strategies in our systematic trading environment. To meet the unique challenges of global markets, Susquehanna applies machine learning and advanced quantitative research to vast datasets in order to uncover actionable insights and build effective strategies. By uniting deep market expertise with cutting-edge technology, we excel in solving complex problems and pushing boundaries together.

#J-18808-Ljbffr

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Explore related searches

Current related jobs

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive business …

Listing review due 2026-10-08View job

Kohler

WhatJobs

Engineer, New Product Integration

kohler, WI

Salary not specified

Engineer, New Product Integration Work Mode: Onsite Location: Onsite, four days per week - Kohler, WI Opportunity This is mo…

Listing review due 2026-10-08View job

GE Vernova

WhatJobs

Principal Engineer - AI Engineering

niskayuna, NY

See pay details in description

Job Description Summary GE Vernova is embracing cutting-edge technologies to streamline operations, improve customer experiences, and drive gr…

Listing review due 2026-10-08View job

GE Vernova

WhatJobs

Lead Product Safety and Compliance Engineer

rochester, NY

See pay details in description

Job Description Summary The Product Safety & Compliance Engineer works directly within the engineering development team to ensure our products…

Listing review due 2026-10-08View job

Hobbs Brook Real Estate

WhatJobs

Commercial Facilities Engineer

waltham, MA

$30.88 to $38.61 per hour

Job Description: Hobbs Brook Real Estate LLC is an innovative commercial real estate leader with a portfolio of forward-thinking, sustainable pr…

Listing review due 2026-10-08View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing …

Listing review due 2026-10-08View job