About this opportunity
Plaud Inc. lists this Real-Time LLM Inference & Serving Engineer opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Plaud Inc. is building the next generation real-world AI interface for professionals.
We seek engineers who can push high-throughput, ultra-low-latency inference for large language and speech models, balancing latency and throughput in live streaming contexts, and optimizing multi-GPU deployments. Join a fast-moving team with cross-functional collaboration between ML training and backend infrastructure to deliver hardware-software AI systems that amplify human productivity, with hybrid office
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.