waypointjobs

Career Techniques

Machine Learning Performance Engineer (Inference)

new york, NY

Check who can apply and the requirements below before continuing.

About this opportunity

Career Techniques lists this Machine Learning Performance Engineer (Inference) opportunity in new york, New York. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Responsibilities

Benchmarking & Strategy: Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide infrastructure deployment decisions.

System Architecture Optimization: Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle.

Infrastructure & Deployment Feasibility: Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes.

GPU Kernel Development: Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon.

Model Optimization & Deployment: Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements.

Cross-Functional Collaboration: Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments.

Qualifications

2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain.

ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation.

Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute).

GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM).

Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures.

Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs.

Prior experience in financial trading is not required.

Comp: $200-300K + Bonus

#J-18808-Ljbffr

Worksite address

new york, NY, 10261, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

American Honda Motor Co., Inc.

WhatJobs

Senior Product Quality & Root Cause Engineer

haw river, NC

Salary not specified

What Makes a Honda, is Who makes a Honda Honda has a clear vision for the future, and it’s a joyful one.  We are looking for individuals with t…

Listing review due 2026-10-06View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing…

Listing review due 2026-10-06View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive busine…

Listing review due 2026-10-06View job