오픈 소스 AI 프레임워크의 기능을 설계하고 개발하며, AI 가속기 및 차세대 GPU를 위한 소프트웨어 스택을 최적화해요. CS, ECE 관련 학사 이상의 학위와 6년에서 12년 사이의 경력이 필수예요. C++ 14/17, 파이썬, 병렬 프로그래밍에 능숙해야 하며, 머신러닝 커널 개발 경험이 필요해요. 현장 근무가 필수이며, 다국적 팀과 협업하며 복잡한 소프트웨어 시스템을 디버깅하는 업무를 수행해요.
We are looking for a dynamic and passionate senior contributor to work in Intel's Data Center and AI group (DCAI). Day-to-day work involves working on Open source AI Frameworks such as PyTorch, Tensorflow, JAX etc. The job role involves design and developing features for Intel' AI frameworks software stack. You will be participating in enabling and optimizing state of the art Software stack for Intel's AI accelerators and next generation GPUs.
The roles and responsibilities that you would need to performance may include the following:• Design and develop SW features for AI frameworks - both HW-agnostic and HW-aware, especially in ML kernel development.
Enhance and extend the Deep learning training, and Inference capabilities in the Software stack.
Identifying optimization opportunities in the software stack to enhance performance of Deep learning workloads
Participate in discussions with Open-source community, involve in development, adopting upstream and Upstream software.
BTech or MS/MTech in CS, ECE or related fields with an overall experience of 6 to 12 years. Proficient in Advanced C++ (C++ 14/17) and Intermediate skills of Python and parallel programming.Experience in developing machine learning kernels such as GEMM, Convolution, Flash attention etcIn depth and hands on experience in one of the frameworks such as PyTorch, Tensorflow or JAXPractical knowledge of Deep Learning models/LLMs for Vision / NLPAbility to debug complex issues in multi layered SW systems. Understanding of SW integration in large open-source frameworks.
Strong understanding of computer architecture and HW-SW optimization techniques.
Experience in working on frameworks/platforms that have gone to production.Effective communication skills and experience with working in a cross-geo teams.
Preferable Experience in developing and integrating CUTLASS or Triton based kernels in Large language models (LLMs).
Knowledge of compiler algorithms for heterogeneous system and Fuser optimizations.
.
Experienced Hire
Shift 1 (India)
India, Bangalore
All qualified applicants will receive consideration for employment without regard to race, color, religion, religious creed, sex, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, military and veteran status, marital status, pregnancy, gender, gender expression, gender identity, sexual orientation, or any other characteristic protected by local law, regulation, or ordinance.
N/A
This role will require an on-site presence. * Job posting details (such as work model, location or time type) are subject to change.
ADDITIONAL INFORMATION: Intel is committed to Responsible Business Alliance (RBA) compliance and ethical hiring practices. We do not charge any fees during our hiring process. Candidates should never be required to pay recruitment fees, medical examination fees, or any other charges as a condition of employment. If you are asked to pay any fees during our hiring process, please report this immediately to your recruiter.