프리퍼드네트웍스 · 채용 중 73건
LLM Serving Engine Engineer / LLMサービングエンジンエンジニア
LLM Serving Engine Engineer / LLMサービングエンジンエンジニア
AI 엔지니어정규직전체 · 경력 무관
프리퍼드네트웍스에서 MN-Core L 시리즈를 위한 LLM 서빙 엔진 개발자를 모집합니다. LLM 추론에 대한 이론적 지식과 vLLM, SGLang 등 오픈소스 엔진에 대한 이해가 필수입니다. 하드웨어 사양과 성능의 관계를 파악하고, 실행 제어 및 KV 캐시 최적화를 수행합니다. 하드웨어 개발자와 협업하며 최첨단 컴퓨팅 기술을 실무에 적용하는 도전적인 환경입니다.
We are looking for an engineer to join the development of the LLM inference serving engine for the MN-Core L series.
At PFN, we are developing and utilizing the MN-Core™ architecture along with its recent generations MN-Core L series. The MN-Core architecture employs compiler-based pre-scheduling to optimize the majority of processor control, meaning the compiler's quality directly determines the performance of the accelerator. The L series leverages novel 3D stacked DRAM technology to achieve high bandwidth required for LLM inference. The serving engine aims to enable ultra-low latency LLM inference by tightly integrating the compiler-generated programs, runtime environment, and data processing on DRAM.
In this position, you will be involved in developing this serving engine. Your specific software development responsibilities will include:
In addition to these core responsibilities, depending on your interests and aptitudes, you may also participate in broader projects related to high-performance computing system design and utilization, compression such as quantization, testing with downstream tasks, and architectural considerations for the MN-Core itself.
MN-Core Lシリーズ向けLLMサービングエンジンの開発チームに加わるエンジニアを募集しています。
PFNでは、最新世代のMN-Core™アーキテクチャおよびMN-Core Lシリーズの開発・活用を進めています。MN-Coreアーキテクチャはコンパイラによる事前スケジューリングを採用しており、プロセッサ制御の大部分を最適化しています。このため、コンパイラの品質がアクセラレータの性能に直接影響します。Lシリーズでは、LLM推論に必要な高帯域幅を実現するため、3D積層DRAM技術を採用しています。本LLMサービングエンジンは、コンパイラが生成したプログラム、ランタイム環境、およびDRAM上でのデータ処理を緊密に協調させることで、超低遅延なLLM推論を実現することを目的としています。
本ポジションでは、このサービスエンジンの開発に携わっていただきます。具体的なソフトウェア開発業務としては以下が含まれます:
これらの主要業務に加え、ご自身の興味や適性に応じて、高性能コンピューティングシステムの設計・活用、量子化などの圧縮技術、下流タスクとの連携テスト、MN-Coreアーキテクチャ自体の設計検討など、より広範なプロジェクトにも関与していただくことも可能です。
At PFN, you'll have the opportunity to work on the development of the serving engine for cutting-edge LLM inference accelerators. You'll work in an environment where you can collaborate closely with users as well as hardware developers, making this an ideal opportunity for those passionate about applying world-class computing technology in real-world applications.
PFNでは、最先端のLLM推論アクセラレータ向けソフトウェアの開発に携わることができます。ハードウェア開発者と密接に連携しながら開発を進めることができる環境であり、世界最先端のコンピューティング技術を実社会で活用することに情熱を持つ方にとって、理想的な環境です。
Individuals with broad interests and a desire to acquire knowledge in new technical domains
Colleagues who can respect and work well together, regardless of job role or background
Those who can leverage their strengths and support team members
Individuals who can approach problem-solving as their own responsibility, regardless of ownership
People who can absorb new knowledge and enjoy working in environments with diverse expertise
様々な分野への関心、新たな技術領域の知見獲得の意欲のある方
同職種・他職種に関わらずリスペクトして一緒に楽しく働ける方
強みを活かして、チームメンバと助け合える方
周りの課題に対しても自分事として捉え課題解決を推進できる方
様々な専門性を持つ人がいる環境で新しいことを吸収し、楽しめる方
[About the inference chip MN-Core L1000](https://mn-core.com/)
[推論チップ MN-Core L1000について](https://mn-core.com/ja)
Basic understanding of computer science at university undergraduate level
Interest in emerging LLM inference accelerators.
Theoretical and practical familiarity with LLM inference.
Understand how hardware specifications influence performance.
Understanding of different workload scenarios and how they relate to performance.
Understanding of the different tradeoffs involved in inference optimization.
Familiar with open source LLM inference engines (such as vLLM, SGLang, dynamo, llama.cpp).
Openness to work in an English-Japanese-mixed environment
大学学部レベルのコンピュータサイエンスについての基礎的な理解
新興のLLM推論アクセラレータ技術に対する関心
LLM推論に関する理論的・実践的な知識を有していること
ハードウェア仕様が性能に与える影響を理解できること
様々なワークロードシナリオとそれらが性能に及ぼす影響についての理解
推論最適化に伴う各種トレードオフについての理解
オープンソースのLLM推論エンジン(vLLM、SGLang、dynamo、llama.cppなど)に精通していること
英語と日本語が混在する環境での業務に柔軟に対応できること
Contributions to open source LLM inference engines (vLLM, SGLang, llama.cpp, dynamo, …).
Experience with LLM serving (either running locally or at scale).
Familiar with core parts of LLM serving (such as request scheduling, KV cache management or prefix caching).
Knowledge about LLM serving APIs (/chat/completion, /responses, /messages, …).
Interest in LLM applications (such as coding agents).
Knowledge about ascertaining LLM inference performance through benchmarks and simulations.
System engineering experience (understanding of performance optimization, profiling, memory management, storage, concurrency, etc.).
Experience with system programming languages (C/C++, Rust, Go, …).
Experience with emerging LLM inference accelerators.
Ability to read technical documentation/discussions in Japanese.
OSSのLLM推論エンジン(vLLM、SGLang、llama.cpp、dynamoなど)への貢献経験
LLMのサービス提供に関する実務経験(ローカル環境および大規模環境での運用双方)
LLMサービス提供の主要コンポーネントに関する知識(リクエストスケジューリング、KVキャッシュ管理、プレフィックスキャッシュ処理など)
LLMサービス提供APIに関する知見(/chat/completion、/responses、/messagesなど)
LLM応用技術への関心(コーディング支援エージェントなどの事例)
ベンチマークテストやシミュレーションを通じたLLM推論性能評価に関する知識
システムエンジニアリングの実務経験(パフォーマンス最適化、プロファイリング、メモリ管理、ストレージ、並行処理などの理解)
システムプログラミング言語の使用経験(C/C++、Rust、Goなど)
最新のLLM推論アクセラレータ技術に関する知見
日本語の技術文書/ディスカッションを読み解く能力
経験、業績、能力、貢献に応じて、当社規定により優遇
Experience, performance, skills, contribution are taken into consideration.
東京都千代田区大手町1−6−1 大手町ビル / Otemachi Bldg., 1-6-1 Otemachi, Chiyoda-ku, Tokyo, Japan 100-0004
リモート勤務制度あり (日本国内に限る) / Remote work system available (limited to work in Japan)