👋 About the Team
The NetsPresso Platform Team designs and develops the core platforms and software behind NetsPresso, transforming Nota AI’s research in AI model compression and optimization into real-world products.
Working alongside the Model Representation, Quantization, Graph Optimization, and SW Engineering teams, the Model Engineering team serves as the bridge that enables NetsPresso’s optimization technologies to be applied reliably across a wide range of AI models and device environments.
Our team:
- Optimizes diverse AI models for real-world deployment on target devices
- Builds agent-driven experimentation and automation workflows that enable optimization processes to scale beyond manual operations
- Analyzes customer requirements and emerging technologies to help shape the future direction of NetsPresso’s core technologies
We operate at the intersection of research and engineering, connecting innovation to real-world applications.
📌 What You’ll Do at This Position
Can this LLM run on a mobile device?
Can we optimize hundreds of models quickly and automatically, instead of handling each one manually?
The NetsPresso Model Engineering Part creates practical answers to these questions. Using optimization engines such as NetsPresso IR, Graph Optimization, and Quantization, you will turn models from research papers or customer-developed models into models that actually run on customer devices.
This is a senior engineering position. Beyond executing model optimization tasks, we are looking for someone who can lead technical decision-making across diverse models, devices, and customer scenarios, and help shape the technical direction of the part.
- Design and execute optimization strategies for various AI models and devices.
- Design agent-based automation scenarios to systemize repetitive tasks.
✅ Key Responsibilities
- AI model optimization and management for target devices
- Optimize various AI models, including Computer Vision and Generative AI models such as LLMs and VLMs, for target devices such as CPU, GPU, and NPU using NetsPresso’s model compression and optimization modules
- Quantitatively analyze trade-offs among performance, accuracy, memory, and latency, and derive optimal solutions for each model and device
- Manage device-specific optimized models and turn experimental results into reusable assets
- Agent-based optimization automation and scenario development
- Design agent-based model optimization experiment structures that do not rely on manual configuration
- Implement and improve automated optimization scenarios tailored to different device characteristics
- Turn repetitive tasks into reusable assets through automation tools, internal libraries, and CI/CD pipelines
- Customer model support and technical research
- Establish model optimization strategies and provide technical support based on customer device environments
- Research and apply new AI models, devices, and optimization techniques
- Identify and propose improvements to NetsPresso Core technologies from a practical deployment perspective
- Technical leadership and collaboration
- Define technical interfaces with other parts and lead cross-functional collaboration
- Review code and experimental results from fellow engineers and provide technical mentoring
- Contribute to establishing technical standards and decision-making within the part
✅ Requirements
- Bachelor’s degree or higher in Computer Science, Electrical Engineering, AI/ML, or a related field
- 7+ years of experience in a relevant field
- Proficiency in Python-based software development, with experience writing readable code and collaborating through Git
- Experience developing and debugging models using deep learning frameworks such as PyTorch
- Hands-on experience with AI model inference, deployment, and optimization
- Practical experience using inference engines such as ExecuTorch, ONNX Runtime, or LiteRT, serving tools such as vLLM, or optimization tools such as Compression or Quantization
- Ability to design inference pipelines and analyze performance, including accuracy, latency, and memory
- A problem-solving mindset for turning repetitive tasks into automatable structures
- Experience leading technical decision-making or technical execution at the project or small-team level
- No restrictions on overseas travel
✅ Pluses
- Master’s or Ph.D. degree in AI/ML, Computer Science, Electrical Engineering, or a related field
- Deep experience with model optimization and compression, such as Quantization, Pruning, Distillation, or Graph Optimization
- Experience optimizing models across various devices, including CPU, GPU, and NPU
- Experience with on-device optimization for Generative AI models such as LLMs and VLMs
- Experience with ML pipeline automation, including CI/CD and experiment management
- Experience building agent-based or AutoML-based automation systems
- Experience providing customer technical support, conducting PoCs, or delivering technical consulting
- Experience leading a team or part of 5+ members, or mentoring multiple junior engineers
- Ability to communicate in English in a business setting
✅ Hiring Process
- Document Screening → 1st Interview → Assignment → 2nd Interview → 3rd Interview → Offer → Hire
(Additional assignments may be included during the process.)
🤓 A Message from the Team
In this role, strong interest in new AI models and device environments is important, along with the ability to turn ideas into working results. This position goes beyond simple experimentation or analysis. You will implement and advance model optimization and engineering technologies that are directly connected to the NetsPresso service.
Because various models, devices, and automation scenarios are closely connected, we highly value active communication within the team and a proactive approach to problem-solving. We are looking for someone who can bring deep expertise as a senior engineer while also working with teammates to shape the technical direction of the part.
Please Check Before Applying! 👀
- This job posting is open continuously, and it may close early upon completion of the hiring process.
- Resumes that include sensitive personal information, such as salary details, may be excluded from the review process.
- Providing false information in the submitted materials may result in the cancellation of the application.
- Please be aware that references will be checked before finalizing the hiring decision.
- Compensation will be discussed separately upon successful completion of the final interview.
- There will be a probationary period after joining, and there will be no discrimination in the treatment during this period.
- To support the employment of persons with disabilities, you may optionally submit a copy of your disability registration certificate under “Additional Documents,” if administrative verification is required. Submission is optional and does not affect the evaluation process.
- Veterans and individuals with disabilities will receive preferential treatment in accordance with relevant regulations.
🔎 Helpful materials