Sociedade Brasileira de Telecomunicações · desde 1983 secretaria@sbrt.org.br
← SBrT2021

Codificação de Redes Neurais sem Retreino

Marcos Vinicius Tonin, Ricardo L de Queiroz
Compressão de redes neuraisambientes de recursos limitadosDistribuições de pesosOpen Neural Network Exchange (ONNX)

Resumo

Machine learning is used in many areas on many problems. Hence, neural network development continuously grows and their sizes steadily increase. This work intends to reduce the networks file sizes without re-training. We relate weight entropy and accuracy for several quantization methods. It is not apparent any correlation among the weights that would enable a more sophisticated coding than aggressive scalar quantization followed by entropy coding of the weights. Our studies indicate that it is possible to reduce fourfold the network size without significantly impacting its performance. The recommended quantization and encoding may be incorporated into a format for the deployment of neural networks.