About this opportunity
OP Recruiting lists this AI Research Engineer, Inference opportunity in chicago, Illinois. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Job Title:
AI Research Engineer – Deep Learning Inference
Location:
New York, NY | London, UK
About The Opportunity
Join an industry-leading global quantitative technology organization at the forefront of machine learning innovation. We are seeking a High-Performance AI Research Engineer to drive speed and scalability across our real-time predictive modeling infrastructure. In this role, you will bridge the gap between machine learning research and hardware acceleration, building low-latency inference systems that process continuous global data streams to drive core business decisions.
Responsibilities
Drive performance optimizations across all aspects of large-scale model execution, including custom kernel creation, data streaming pipelines, and novel hardware integration.
Partner directly with machine learning researchers to co-design neural architectures optimized for extreme real-time execution.
Author low-level primitives and custom operations to extract maximum compute throughput from modern hardware architectures.
Evaluate, benchmark, and deploy novel hardware acceleration tech, ranging from off-the-shelf accelerators to custom specialized silicon.
Formulate and execute engineering initiatives that address complex, non-obvious bottlenecks in ultra-low-latency deep learning inference.
Requirements (Must-Have)
At least two years of hands-on experience engineering production-grade deep learning systems within any complex domain (such as robotics, computer vision, audio, NLP, physics, or recommender platforms).
Strong lower-level engineering foundation, including experience writing custom compute kernels (e.g., CUDA, Triton, Pallas, or CuTe DSLs).
Proficiency with framework compilation internals and low-level runtime environments (PyTorch, JAX, XLA, or CUDA Graphs).
Practical exposure to hardware acceleration tech, such as FPGAs, ASICs, or specialized AI processors.
Proven ability to adapt algorithms and technical concepts across different domain applications.
Preferred Qualifications
Experience optimizing or serving Large Language Models (LLMs) and foundation architectures.
Note: Prior background in quantitative finance or trading is explicitly NOT required.
Compensation & Benefits
Highly competitive base salary, performance-based bonus incentive, and premium health/wellness benefits package. Equal Opportunity Employer.
Accepting Candidates
#J-18808-Ljbffr
Worksite address
chicago, IL, 60290, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.