Journal of Shanghai University(Natural Science Edition) ›› 2026, Vol. 32 ›› Issue (4): 665-677.doi: 10.12066/j.issn.1007-2861.2555

• Communication and Information Engineering • Previous Articles    

Translating classical Chinese with pre-trained center embedding model

JIN Yanliang1,2, GAO Zhifeng1,2, GAO Yuan1,2   

  1. 1. School of Communication and Information Engineering, Shanghai University, Shanghai 200444, China;
    2. Shanghai Institute for Advanced Communication and Data Science, Shanghai University, Shanghai 200444, China
  • Received:2023-12-14 Published:2026-09-04

Abstract: To solve the low accuracy of model translation words owing to the lack of parallel corpora and variable meanings in the current ancient Chinese machine translation model, a Transformer-based pre-trained feature center algorithm model is proposed. This method first trains the language model to extract pretrained feature centers to improve generalization ability when parallel corpora are insufficient. Second, the Transformer generates queries in the feature center space and determines the generation based on the distance between the queries in the feature center to improve the model’s understanding of polysemy. The experimental results indicate that the generation accuracy of the proposed method is generally higher than that of the benchmark model. In the XiHan Datasets, BLEU 1~4 scores were 6.1, 4.9, 3.7, and 2.8 higher than those of the Transformers. In the WuDaiShiGuo Datasets, the BLEU 1~4 scores were 4.2, 4.2, 3.3, and 2.6 higher than those of the Transformers. In the Tang Datasets, the BLEU 1~4 scores were the best. Simultaneously, compared with the Transformer, the parameter count of the proposed method decreased by 5.73 M, demonstrating its advantages.

Key words: natural language processing, Transformer, machine translation, pre-trained feature center, classical Chinese translation

CLC Number: