👋 About the Team
The NetsPresso Platform Team designs and develops the core platforms and software behind NetsPresso, transforming Nota AI’s research in AI model compression and optimization into real-world products.
Working alongside the Model Representation, Quantization, Graph Optimization, and SW Engineering teams, the Model Engineering team serves as the bridge that enables NetsPresso’s optimization technologies to be applied reliably across a wide range of AI models and device environments.
Our team:
- Optimizes diverse AI models for deployment on real-world devices
- Builds agent-driven experimentation and automation workflows that enable optimization processes to scale beyond manual operations
- Analyzes customer requirements and emerging technologies to help shape the future direction of NetsPresso’s core technologies
We operate at the intersection of research and engineering, connecting innovation to real-world applications.
📌 What You’ll Do at This Position
"Can this LLM run efficiently on a mobile device?"
"Can we optimize hundreds of models automatically instead of tuning each one manually?"
The Model Engineering team exists to answer questions like these.
Using NetsPresso's optimization technologies—including its IR, Graph Optimization, and Quantization engines—you will transform research models and customer-developed models into production-ready AI models that run efficiently on real devices.
This role focuses on bridging cutting-edge AI technologies with practical deployment by combining model optimization expertise, automation, and engineering execution.
✅ Key Responsibilities
AI Model Optimization & Deployment
- Optimize AI models for target devices using NetsPresso’s compression and optimization technologies
- Tune Computer Vision and Generative AI models (including LLMs and VLMs) with a focus on performance, accuracy, and memory efficiency
- Manage optimized models across multiple device environments and maintain experiment results and benchmarking data
Agent-Based Optimization Automation
- Design agent-driven optimization workflows that reduce dependency on manual configuration and experimentation
- Develop and improve automated optimization scenarios tailored to specific device characteristics
Customer Enablement & Technology Research
- Support customer model optimization projects and provide technical assistance for deployment environments
- Research emerging AI models, devices, and optimization techniques
- Identify opportunities to improve NetsPresso technologies from a real-world deployment perspective and propose enhancements to the platform
✅ Requirements
- Master’s degree or higher in AI, Machine Learning, or a related field, or equivalent industry experience
- Strong software development skills in Python and experience working in collaborative development environments
- Ability to write maintainable, readable code and use Git-based version control workflows
- Experience developing and debugging models using deep learning frameworks such as PyTorch
- Hands-on experience with AI model Inference, Deployment, and Optimization
- Practical experience with at least one inference engine (e.g., ONNX Runtime), serving framework (e.g., TorchServe), or optimization tool (e.g., Quantization)
- Experience designing inference pipelines and analyzing performance metrics, including accuracy, latency, and memory usage
- Strong problem-solving mindset with a focus on transforming repetitive tasks into scalable and automated workflows
- No restrictions on overseas travel
✅ Pluses
- Experience documenting and sharing complex experimental findings
- Experience with model optimization and compression techniques such as Quantization, Pruning, and Graph Optimization
- Experience optimizing AI models across various hardware platforms, including CPUs, GPUs, and NPUs
- Experience using model conversion and deployment tools such as ExecuTorch, ONNX, TFLite, and TensorRT
- Experience with ML pipeline automation, including CI/CD and experiment management systems
- Experience building Agent-based or AutoML-driven automation systems
- Experience supporting customer technical engagements or conducting Proof-of-Concept (PoC) projects
✅ Hiring Process
- Document Screening → 1st Interview → Assignment Presentation → 2nd Interview → 3rd Interview → Offer → Hire
- (Additional assignments may be included during the process.)
🤓 A Message from the Team
We value people who are genuinely curious about new AI models and device environments, and who have the ability to turn ideas into working solutions.
This role goes beyond experimentation and analysis. You will build model optimization and engineering technologies that are directly integrated into the NetsPresso product and used by real customers.
Because our work connects diverse models, hardware platforms, and automation workflows, we place a strong emphasis on open communication, collaboration, and proactive problem-solving.
If you enjoy diving deep into the constraints and challenges of real-world AI deployment, growing alongside talented teammates, and building technology that people can actually use, we believe you will find meaningful opportunities to make an impact here.
Please Check Before Applying! 👀
- This job posting is open continuously, and it may close early upon completion of the hiring process.
- Resumes that include sensitive personal information, such as salary details, may be excluded from the review process.
- Providing false information in the submitted materials may result in the cancellation of the application.
- Please be aware that references will be checked before finalizing the hiring decision.
- Compensation will be discussed separately upon successful completion of the final interview.
- There will be a probationary period after joining, and there will be no discrimination in the treatment during this period.
- To support the employment of persons with disabilities, you may optionally submit a copy of your disability registration certificate under “Additional Documents,” if administrative verification is required. Submission is optional and does not affect the evaluation process.
- Veterans and individuals with disabilities will receive preferential treatment in accordance with relevant regulations.
🔎 Helpful materials