👋 About the Team
The NetsPresso Platform Team researches AI model compression and optimization technologies at Nota AI, and designs and implements these technologies into real products for users. The team consists of several specialized areas, including Model Representation, Quantization, Graph Optimization, Model Engineering, and SW Engineering.
Among them, the Graph Optimization team focuses on discovering ways to execute AI models more efficiently on real hardware environments.
Our work involves analyzing both AI models and target hardware to identify optimization opportunities, then experimentally validating and implementing optimization techniques that can be practically applied. This role requires working across both the theoretical and practical aspects of AI models and AI hardware, while actively leveraging AI development tools throughout the process.
We’re looking for someone ready to explore this space with us.
📌 What You’ll Do at This Position
In this role, you will gain hands-on experience across the entire lifecycle of AI model execution on real hardware—from graph-level transformations, to lowering into target backends and runtimes, all the way to analyzing actual inference behavior on hardware.
Rather than simply implementing predefined optimization tasks, you will investigate why performance bottlenecks occur in specific model and hardware combinations, and take part in defining entirely new optimization opportunities yourself.
You will work within an environment equipped with optimization-heavy codebases and AI agent-driven automation workflows, allowing you to rapidly iterate through implementation, validation, and experimentation cycles to produce meaningful results.
If you are already comfortable using AI tools not merely as assistants, but as core productivity accelerators, this environment will enable you to grow quickly.
💡 We’d Love to Work With Someone Who...
- Has personally investigated why an AI model failed to run correctly on a specific runtime or hardware target
- Can clearly explain why a particular optimization strategy or technical decision was chosen
- Actively embraces AI tools in their workflow and can concretely describe how they have leveraged them in practice
✅ Key Responsibilities
- Analyze and resolve mismatches between AI models and target hardware environments such as NPUs, GPUs, and CPUs
- Design and implement graph-level optimization passes, including op fusion, folding, decomposition, and replacement
- Explore and experiment with optimization opportunities by jointly considering model architecture and hardware characteristics
- Perform model transformation and optimization using frameworks such as ExecuTorch, ONNX, and PyTorch
- Design and utilize AI agent-based automation workflows
✅ Requirements
- At least 2 years of industry experience in a related field after a bachelor’s degree, or a master’s degree or higher in a relevant field
- Practical experience with PyTorch, ONNX, Python, Linux, and Git
- Experience following the full lifecycle of AI models, from training to execution on real devices
- Strong foundational understanding in at least one relevant area, such as AI model optimization, compression, compilers, or kernel optimization—beyond simply using existing libraries
- Ability to clearly articulate the reasoning behind technical decisions
- No restrictions on overseas travel
✅ Pluses
- Experience analyzing and transforming models at the Graph IR level (e.g., TensorRT, TFLite, ExecuTorch)
- Experience optimizing models for NPUs or custom accelerators
- Deep understanding of GPU or NPU kernels
- Strong understanding of optimization techniques such as quantization and pruning
- Understanding of Generative AI architectures, including LLMs, VLMs, and diffusion models
- Experience integrating AI agent tools (e.g., Claude Code, Codex) into real engineering workflows
- Experience publishing related papers or contributing to open source projects
✅ Hiring Process
- Document Screening → 1st Interview → Online Assignment → 2nd Interview → 3rd Interview → Offer → Hire
- The process may be partially adjusted depending on circumstances, with prior notice provided.
- Additional assignments may be included during the process.
🤓 A Message from the Team
There’s a difference between running a model and understanding why it behaves the way it does. We want to work with people who are interested in the latter.
We’re looking for someone who can explain the reasoning behind technical decisions, who has experience tracing problems to their root causes when things break, and who already treats AI tools as genuine productivity multipliers in their workflow.
If that sounds like you, we believe you’ll find meaningful work and growth with our team.
Please Check Before Applying! 👀
- This job posting is open continuously, and it may close early upon completion of the hiring process.
- Resumes that include sensitive personal information, such as salary details, may be excluded from the review process.
- Providing false information in the submitted materials may result in the cancellation of the application.
- Please be aware that references will be checked before finalizing the hiring decision.
- Compensation will be discussed separately upon successful completion of the final interview.
- There will be a probationary period after joining, and there will be no discrimination in the treatment during this period.
- To support the employment of persons with disabilities, you may optionally submit a copy of your disability registration certificate under “Additional Documents,” if administrative verification is required. Submission is optional and does not affect the evaluation process.
- Veterans and individuals with disabilities will receive preferential treatment in accordance with relevant regulations.
🔎 Helpful materials