Member of Technical Staff - Inference Research

ML EngineerFull-timeLead · Any experienceNew York, San FranciscoOn-site
LLM InferenceCUDAQuantizationDistributed SystemsPython

Modal is seeking a technical staff member to lead LLM inference research. A deep understanding of the LLM serving stack and experience in inference optimization are essential. You will research and deploy high-performance techniques like quantization, scheduling, and kernel optimization.

Be the first to hear about new roles at Modal

Follow this company and we'll notify you as soon as new roles are posted.

Start free
Open