About this opportunity
iFrame Corporation lists this Inference Runtime Performance Engineer — GPU Kernels opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
iFrame Corporation is hiring for a role on the runtime team to own end-to-end performance for multiple model families on our managed inference runtime. You will write fused CUDA/Triton kernels and work across tokenizer through KV cache and decoding, targeting H100/H200 and other accelerators.
With 5+ years in systems-level performance and a track record of speed-ups, you will design long-context primitives and publish external write-ups quarterly while sharing on-call duties with SRE and
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.