Journal of Jilin University(Engineering and Technology Edition) ›› 2026, Vol. 56 ›› Issue (3): 793-801.doi: 10.13229/j.cnki.jdxbgxb.20240809

Previous Articles    

Mixed-precision quantization of post-trained conversation summarization based on sensitivity analysis

Yu-peng LIU(),Yu-hao ZHANG,Xin MENG   

  1. School of Computer Science and Technology,Harbin University of Science and Technology,Harbin 150080,China
  • Received:2024-07-19 Online:2026-03-01 Published:2026-03-31

Abstract:

This paper proposes a mixed precision integer quantization method for the layer of a trained dialogue summarization model, which uses a more accurate method based on augmented Hessian matrix to evaluate layer sensitivity, considering inter-layer correlation while effectively preserving intra-layer information. Due to the large number of outliers in the conversation summarization model activation, quantifying them directly will result in significant decrease in model accuracy. In this paper, we propose a method of smoothing first and then mixed precision quantization. Activation outliers are first translated to weights using smoothing factors, and then weights and activations are quantified with mixed precision based on sensitivity assessment scores. This method can quantize the model from 16 bit or 32 bit floating-point to 4-16 bit mixed precision, which reduces the storage requirement of the model and accelerates the inference speed. On the benchmark dataset SAMSum, a significant performance improvement is achieved compared with the classical baseline system, which is almost the same as the performance of the non-quantized model.

Key words: sensitivity, mixed precision, outliers, conversation summarization

CLC Number: 

  • TP312

Fig.1

Performance comparison before and after quantization"

Fig.2

Structure figure of the hybrid-precision quantization method"

Fig.3

Comparison of three search schemes"

Fig.4

Sensitivity comparison"

Table 1

Introduction to the experiment model"

模型名称介 绍
LBART22基于BART23的预训练模型,在对话摘要任务上微调模型,共1亿4千万参数
DSMC 24基于Transformer的非预训练模型,参数量约7千万

Table 2

Description the comparison models"

量化方法主要思想量化位宽(A为激活,W为权重)
QBert25基于分组的量化方案,使用二阶泰勒展开来评估权重A8W8
MREM-P3将模型分成多个模块,采用并行策略同时为每个模块最小化量化带来的重构误差A8W8
BRECQ4利用神经网络的基本构建块,逐个进行重构,同时在跨层依赖和泛化误差之间取得了良好平衡A8W8
QDrop26通过随机丢弃来实现激活量化,并首次将激活的量化推进到2 bitA4W8
MrBiQ27利用多级二值化表示权重并允许不同数据表示激活A8W8
AdaQuant13通过分别优化每个层或块的参数,并使用校准集来设置激活动态范围,以最小化量化误差A4W8

Fig.5

Comparison of other studies on DSMC model"

Fig.6

Comparing other studies on LBART model"

Fig.7

Per-layer bit-width settings under 99.99% and 98% scoring searches"

Table 3

Quantization number of the different bit widths"

评分4 bit8 bit16 bit
99.99%311278
98%181741

Table 4

Evaluation number of LBART at the different accuracies"

方法LBART 99%LBART 99.9%
8 bit4 bit8 bit4 bit
Hessian6667
AugHessian6867
[1] Dettmers T, Lewis M, Belkada Y, et al. GPT3. int8 (): 8-bit matrix multiplication for transformers at scale[J]. Advances in Neural Information Processing Systems,2022,35: 30318-30332.
[2] Zafrir O, Boudoukh G, Izsak P, et al. Q8BERT: Quantized 8bit bert[C]∥Proceedings of the Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing-NeurIPS Edition, Vancouver, Canada, 2019: 36-39.
[3] Bai H L, Hou L, Shang L F, et al. Towards efficient post-training quantization of pre-trained language models[J]. Advances in Neural Information Processing Systems, 2022, 35: 1405-1418.
[4] Li Y H, Gong R H, Tan X, et al. BRECQ: pushing the limit of post-training quantization by block reconstruction[C]∥International Conference on Learning Representations, Vienna, Austria, 2021, 2: 1-17.
[5] Wang K, Liu Z, Lin Y, et al. HAQ: hardware-aware automated quantization with mixed precision[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 8612-8620.
[6] Schaefer C J S, Joshi S, Li S,et al. Edge inference with fully differentiable quantized mixed precision neural networks[C]∥Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, USA, 2024: 8460-8469.
[7] Park E, Yoo S. PROFIT: a novel training method for sub-4-bit mobilenet models[C]∥Proceedings of the Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, 2020: 430-446.
[8] Yao Z W, Dong Z, Zheng Z C, et al. HAWQ-V3: dyadic neural network quantization[C]∥Proceedings of the International Conference on Machine Learning, Shenzhen, China, 2021: 11875-11886.
[9] Nahshan Y, Chmiel B, Baskin C, et al. Loss aware post-training quantization[J]. Machine Learning, 2021, 110(11-12): 3245-3262.
[10] Nagel M, Amjad R A, Van B M, et al. Up or down? Adaptive rounding for post-training quantization[C]∥Proceedings of the International Conference on Machine Learning, Vienna, Austria, 2020: 7197-7206.
[11] Yao Z W, Aminabadi R Y, Zhang M J, et al. ZeroQuant: efficient and affordable post-training quantization for large-scale transformers[J]. Advances in Neural Information Processing Systems, 2022, 35: 27168-27183.
[12] Cai Y, Yao Z, Dong Z, et al. Zeroq: a novel zero shot quantization framework[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2020: 13169-13178.
[13] Hubara I, Nahshan Y, Hanani Y, et al. Accurate post training quantization with small calibration sets[C]∥Proceedings of the International Conference on Machine Learning, Online, 2021: 4466-4475.
[14] Frantar E, Alistarh D. Optimal brain compression: a framework for accurate post-training quantization and pruning[J]. Advances in Neural Information Processing Systems, 2022, 35: 4475-4488.
[15] Demidovskij A, Smirnov E.Effective post-training quantization of neural networks for inference on low power neural accelerator[C]∥International Joint Conference on Neural Networks, Glasgow, UK, 2020: 1-7.
[16] Zandonati B, Pol A A, Pierini M, et al. Fit: a metric for model sensitivity[C]∥Proceedings of ICLR, Addis Ababa, Ethiopia, 2022: 1-20.
[17] Zheng D, Liu Y, Li L. Leveraging inter-layer dependency for post-training quantization[J]. Advances in Neural Information Processing Systems, 2022, 35: 6666-6679.
[18] Wei X Y, Zhang Y C, Zhang X G, et al. Outlier suppression: pushing the limit of low-bit transformer language models[J]. Advances in Neural Information Processing Systems, 2022, 35: 17402-17414.
[19] Bondarenko Y, Nagel M, Blankevoort T. Understanding and overcoming the challenges of efficient transformer quantization[C]∥Proceedings of the Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 2021: 7947-7969.
[20] Zeng A, Liu X, Du Z, et al. GLM-130B: an open bilingual pre-trained model[C]∥Proceedings of ICLR, Kigali, Rwanda, 2023: 1-56.
[21] Wu H, Judd P, Zhang X, et al. Integer quantization for deep learning inference: principles and empirical evaluation[J/OL].[2024-07-06]. arXiv Preprint arXiv:.
[22] Fabbri A, Rahman F, Rizvi I, et al. ConvoSumm: conversation summarization benchmark and improved abstractive summarization with argument mining[C]∥Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Bangkok, Thailand, 2021: 6866-6880.
[23] Lewis M, Liu Y, Goyal N, et al. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension[C]∥Proceedings of the 58th Annual Meeting of the Association for Computational Linguistic, Florence, Italy, 2019: 7871-7880.
[24] Liu Y P, Zhang Y H, Liu G. A conversation summary generation method for medical consultations[P]. China Patent: ZL115964475A, 2023-04-14.
[25] Shen S, Dong Z, Ye J Y, et al. Q-BERT: hessian based ultra low precision quantization of bert[C]∥Proceedings of the AAAI Conference on Artificial Intelligence, New York, USA, 2020: 8815-8821.
[26] Wei X Y, Gong R H, Li Y H, et al. QDrop: randomly dropping quantization for extremely low-bit post-training quantization[C]∥Proceedings of ICLR, Online, 2022: 1-19.
[27] Jeon Y, Lee C, Cho E, et al. Mr. BIQ: post-training non-uniform quantization based on minimizing the reconstruction error[C]∥Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 12319-12328.
[28] Goo C W, Chen Y N. Abstractive dialogue summarization with sentence-gated modeling optimized by dialogue acts[C]∥Proceedings of the IEEE Spoken Language Technology Workshop, Athens, Greece, 2018: 735-742.
[1] Jun ZHANG,Yu-he WANG,Nan JIANG,Jia-le CAI,Xuan-de TENG,Peng ZHANG. Modeling and accuracy allocation of inter-dimensional coupling errors in multi pivot force measurement platforms [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(8): 2501-2510.
[2] Xian-zhen HUANG,Ming-fei MA,Chao LI,Xu WANG,Zhi-ming RONG. Global reliability sensitivity analysis of deformation for machine tool rotary table [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(6): 1906-1914.
[3] Kai MA,Jian-hang SUN,Sen-kang YAN,Yan TAO,Wen-tao WANG,Gui-kai GUO. Multi-objective optimization method of structural static displacement based on projection priority selection method [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(1): 74-83.
[4] Wan-feng WEI,Ling-yun KONG,Wei-an XUAN,Fan YANG,Peng GUO. Review of characteristics of asphalt foaming and moisture sensitivity of warm mix mixtures [J]. Journal of Jilin University(Engineering and Technology Edition), 2025, 55(1): 20-35.
[5] Guo-lin YANG,Yi-fan YANG,Hao-dong XU,Gui-jun LUO,Hong-bo XIAO. Calculation method and influencing factors of surface displacement during construction of curved shield tunnel [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(7): 1997-2008.
[6] Tian-hao WANG,Bo LI,Quan-yi YU,Lin-lin XU,Guo-qiang JIA,Shan-shan GUAN. Uncertainty quantification of electric vehicle's wireless power transfer efficiency based on sparse polynomial chaos expansion method [J]. Journal of Jilin University(Engineering and Technology Edition), 2024, 54(12): 3433-3442.
[7] Xian-zhen HUANG,Kai-bo SUN,Xiao-gang LUAN,Bing HU. Reliability sensitivity analysis of bolt pre-tightening connection [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(8): 2219-2226.
[8] Yao-long KANG,Li-lu FENG,Jing-an ZHANG,Su-e CAO. Fast outlier mining algorithm in uncertain data set based on spectral clustering [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(4): 1181-1186.
[9] Zhi-qiang HAN,Gang XIE,Yong-jun ZHOU,Shi-zhong LIU,Min-jie JIN. Numerical analysis method of vehicle⁃bridge coupling vibration of curved bridge [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(2): 515-522.
[10] Xu CHEN,Chao-fei CAO,Jing SHANG,Ming-xing HUANG,Chang-fa AI,Dong-ya Ren. Evaluation of influence of gradation segregation on pavement moisture damage under action of dynamic and static water environment [J]. Journal of Jilin University(Engineering and Technology Edition), 2023, 53(1): 210-219.
[11] Zi-rong YANG,Yan LI,Xue-feng JI,Fang LIU,Dong HAO. Sensitivity analysis of operating parameters for proton exchange membrane fuel cells [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(9): 1971-1981.
[12] Zi-ling ZHANG,Xiong HU,Yin QI,Wei WANG,Zhi-qiang TAO,Zhi-feng LIU. An approach for error allocation of machine tool based on vector projection response surface method [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(2): 384-391.
[13] Yu-xuan WEI,Ming ZHANG,Jia LIU,Shuo LIU,Ming-yu LU,Hong-yu WANG. Buckling performance of variable stiffness composite cylindrical shells based on mode imperfections [J]. Journal of Jilin University(Engineering and Technology Edition), 2022, 52(1): 91-100.
[14] Min-da REN,Lin CONG,Si-lin SUN,Han-qing FENG. Experiment of performance evolution of asphalt mixtures under multiple pore water pressure cycles [J]. Journal of Jilin University(Engineering and Technology Edition), 2021, 51(4): 1277-1286.
[15] Xue-zhi YAN,Zi-ting WANG,Xin WANG. Analysis of relationship between baseline length and error transfer in ultrasonic 3D positioning system [J]. Journal of Jilin University(Engineering and Technology Edition), 2021, 51(4): 1461-1469.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!