waypointjobs

OP Recruiting

AI Research Engineer, Inference

chicago, IL

Check who can apply and the requirements below before continuing.

About this opportunity

OP Recruiting lists this AI Research Engineer, Inference opportunity in chicago, Illinois. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Job Title:

AI Research Engineer – Deep Learning Inference

Location:

New York, NY | London, UK

About The Opportunity

Join an industry-leading global quantitative technology organization at the forefront of machine learning innovation. We are seeking a High-Performance AI Research Engineer to drive speed and scalability across our real-time predictive modeling infrastructure. In this role, you will bridge the gap between machine learning research and hardware acceleration, building low-latency inference systems that process continuous global data streams to drive core business decisions.

Responsibilities

Drive performance optimizations across all aspects of large-scale model execution, including custom kernel creation, data streaming pipelines, and novel hardware integration.

Partner directly with machine learning researchers to co-design neural architectures optimized for extreme real-time execution.

Author low-level primitives and custom operations to extract maximum compute throughput from modern hardware architectures.

Evaluate, benchmark, and deploy novel hardware acceleration tech, ranging from off-the-shelf accelerators to custom specialized silicon.

Formulate and execute engineering initiatives that address complex, non-obvious bottlenecks in ultra-low-latency deep learning inference.

Requirements (Must-Have)

At least two years of hands-on experience engineering production-grade deep learning systems within any complex domain (such as robotics, computer vision, audio, NLP, physics, or recommender platforms).

Strong lower-level engineering foundation, including experience writing custom compute kernels (e.g., CUDA, Triton, Pallas, or CuTe DSLs).

Proficiency with framework compilation internals and low-level runtime environments (PyTorch, JAX, XLA, or CUDA Graphs).

Practical exposure to hardware acceleration tech, such as FPGAs, ASICs, or specialized AI processors.

Proven ability to adapt algorithms and technical concepts across different domain applications.

Preferred Qualifications

Experience optimizing or serving Large Language Models (LLMs) and foundation architectures.

Note: Prior background in quantitative finance or trading is explicitly NOT required.

Compensation & Benefits

Highly competitive base salary, performance-based bonus incentive, and premium health/wellness benefits package. Equal Opportunity Employer.

Accepting Candidates

#J-18808-Ljbffr

Worksite address

chicago, IL, 60290, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

American Honda Motor Co., Inc.

WhatJobs

Senior Product Quality & Root Cause Engineer

haw river, NC

Salary not specified

What Makes a Honda, is Who makes a Honda Honda has a clear vision for the future, and it’s a joyful one.  We are looking for individuals with t…

Listing review due 2026-10-06View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing…

Listing review due 2026-10-06View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive busine…

Listing review due 2026-10-06View job