Availability awaiting confirmation
We are waiting for a fresh update from the source. This page preserves the last received job details; current availability is not confirmed.
Job description
Thinking Machines is hiring an engineer to ensure the reliability of our GPU supercomputing fleet, owning the seam between hardware, firmware, and operating system. You will diagnose hardware anomalies, track root causes to the hardware, and coordinate fixes with vendors so researchers can run at scale.
Based in San Francisco, this full-time role requires owning drivers, kernel surfaces, and diagnostics, plus automating fleet monitoring and reliability improvements across multi-disciplinary
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.