Algorithm - Serving System Engineer

Software EngineerFull-timeAll levels · 3+ yearsSeoulHybrid
LLM inferencevLLMSGLangTensorRT-LLMCUDATriton

FuriosaAI is looking for an engineer to research and implement next-generation NPU-based serving systems. Candidates must have a deep understanding of LLM inference and experience in accelerator programming using CUDA or Triton. Experience with vLLM or similar inference systems i

Be the first to hear about new roles at FuriosaAI

Follow this company and we'll notify you as soon as new roles are posted.

Start free
Open