Sr. Inference Optimization Engineer (local / edge runtime)

AI EngineerFull-timeSenior · 8+ yearsSanta Clara, California, USAHybrid
C++PythonLLM inferencePerformance profilingLinuxSystems-level programming

Intel is seeking a Senior Engineer to lead inference engine optimization for local and edge environments. You must be proficient in C++ and Python with deep experience in LLM inference and performance profiling. You will optimize engines like llama.cpp and vLLM to maximize hardware e

Be the first to hear about new roles at Intel

Follow this company and we'll notify you as soon as new roles are posted.

Start free
Open