waypointjobs

Long Ridge Partners

Machine Learning Performance Engineer

new york, NY

Check who can apply and the requirements below before continuing.

About this opportunity

Long Ridge Partners lists this Machine Learning Performance Engineer opportunity in new york, New York. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Machine Learning Performance Engineer (Inference)

High-Frequency Trading Firm

Compensation: $600,000-1.5 million total

About the Opportunity

A leading high frequency trading firm is hiring a Machine Learning Performance Engineer to sit at the intersection of quantitative research and high-performance production systems. In his role, you'll architect inference pipelines that operate at the physical limits of hardware, driving the speed, efficiency, and reliability of ML inference so predictive models consistently achieve microsecond-level latency.

GPU usage across the firm's trading teams has grown roughly 100x in the past year as deep learning has moved from a supporting signal to the core of how strategies are built. That growth has outpaced the decision-making around it. Strategies get pushed onto GPUs by default, without anyone systematically asking whether GPU is the right target at all. This role owns that question end to end: benchmark the workload across CPU, GPU, and FPGA, decide the architecture on evidence, then optimize and deploy against it.

You will also have the chance to revisit existing models that never reached production, some of which stalled for hardware or deployment reasons, and run them through different environments to determine where they belong.

What You'll Do

Benchmarking & Strategy

Lead the technical evaluation of inference platforms across CPUs, GPUs, and FPGAs to guide infrastructure deployment decisions

Benchmark trading workloads across architectures before compute is committed, and identify where performance gains actually come from - code-level or hardware-level

Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing

Assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle

Infrastructure & Deployment Feasibility

Work with Infrastructure teams to understand the thermal, power, and operational constraints of hardware platforms, and design inference strategies for latency-critical trading strategies that fit within those envelopes

Consider the interaction between trading workloads, compute requirements, hardware selection, and fleet utilization

Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon

Implement advanced model reduction techniques - quantization, pruning, distillation - to ensure compact memory footprints and numerical stability

Prioritize optimization for low-latency, event-level inference workloads that meet real-time trading requirements

Cross-Functional Collaboration

Partner closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring target deployments to production

What We're Looking For

2+ years optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain

ML frameworks: deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation

Kernel development and tooling: proven experience building custom GPU kernels, with deep familiarity with optimization libraries and compilers (Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) and profiling tools (Nsight Systems, Nsight Compute)

GPU architecture: deep expertise in GPU microarchitecture, including SM execution, warp scheduling, and full memory hierarchy optimization from registers to HBM

Prior experience in financial trading is not required.

Nice to Have

Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs

Why Join?

This is a role with genuine decision-making scope. Rather than optimizing code for whatever hardware happens to be available, you will determine which hardware the workload should run on in the first place, prove it with data, and then build for it. That combination of architectural judgment and hands-on kernel, and the results are measurable in production almost immediately.

You'll work on inference at microsecond latency, where the constraints are physical rather than theoretical, and where memory hierarchy, interconnect behavior, thermal envelopes, and fleet utilization all shape the answer.

Benefits include generous paid time off, regional savings and financial wellness plans, hybrid working options, free breakfast, lunch, and snacks daily, in-office wellness experiences and reimbursement for select wellness expenses, company-sponsored sports teams and fitness events, volunteer and charitable giving opportunities, regular social events, and ongoing workshops and learning opportunities.

The culture is collaborative and low on hierarchy, smart, driven people, an open-plan workspace, casual dress, and an environment where the best idea wins.

#J-18808-Ljbffr

Worksite address

new york, NY, 10261, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Space Systems Mechanical Engineer

laurel, MD

See pay details in description

Description Do you want to design and build unique space structures and spacecraft for NASA missions that enable groundbreaking scientific disco…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Security Engineer

laurel, MD

See pay details in description

Description Are you looking for an opportunity to utilize your technical skills to solve complex, real-world problems? If so, we're looking …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Advanced Reentry Mission Engineer

laurel, MD

See pay details in description

Description Are you interested in hypersonic and reentry system design and prototyping? Do you want to make contributions to next generation …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Thermal and EO/IR Modeling and Simulation Engineer

laurel, MD

See pay details in description

Description Are you looking for a unique opportunity to impact significant advances to the nation's groundbreaking integrated air and missile de…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Network Effects Engineer

laurel, MD

See pay details in description

Description Do you want to perform advanced research, development, and test & evaluation of communications systems and network technologies that…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Realization and Resilience Engineer

laurel, MD

See pay details in description

Description Are you passionate about applying system engineering principles to influence the development and resilience of future strategic weap…

Listing review due 2026-10-06View job