LLM 추론 최적화를 위한 하드웨어 인지형 양자화 연구

대상 하드웨어의 특성을 고려한 양자화 기법을 통해 거대 언어 모델(LLM)의 추론 성능을 최적화하는 과제이다. 모델 정확도를 유지하면서 연산량과 메모리 사용량을 줄여, 자원 제약이 있는 플랫폼에서도 LLM을 효율적으로 배포하는 것을 목표로 한다.

  • This project is funded by Jeonbuk National University.
  • Period: Apr. 2026 – Feb. 2027.

The goal of this project is to develop hardware-aware quantization techniques that optimize the inference performance of large language models (LLMs). By taking into account the characteristics of the target hardware, the proposed quantization methods aim to reduce computational overhead and memory footprint while preserving model accuracy, enabling efficient deployment of LLMs on resource-constrained platforms.