Author’s invariants in English literary texts
https://doi.org/10.25205/1818-7935-2026-24-1-102-112
Abstract
The paper proposes and tests a method for detecting stylistic change points in English-language fiction texts based on statistical analysis of a binary sequence taking into account the words belonging to the author’s invariant. A set of function words was considered as the author’s invariant; texts were displayed as sequences of zeros and ones used to study the stylistic features of texts. An empirical bridge was constructed based on this sequence. The maximum deviation of the empirical bridge served as statistical norm to detect possible change points. The method was tested on a corpus of 100 novels by British and American authors of the 19th–21st centuries, as well as on 9,000 pairwise combinations of concatenated texts. Threshold statistical values are identified that make it possible to distinguish homogeneous texts and texts with a possible stylistic change point. An approximation of the Kolmogorov distribution is used to assess statistical significance. P-values are calculated on the basis of the limiting Kolmogorov distribution. The results of the experiments show a stable ability of the method to identify a stylistic change point between the novels by different authors; the efficiency of the algorithm in distinguishing one text by one author and a combination of two texts by one author; the efficiency of the algorithm in distinguishing one text by one author and a pair of texts by different authors; independence of the method of the text length, that is, high and low values of statistics and p-values are observed for different text lengths. The analysis of empirical bridges and the behavior of individual function words show that stylistic features of the text can manifest themselves both on the scale of the entire work and within its individual parts. The method effectively records such changes. The developed approach can be used in the framework of corpus linguistics, applied stylistics, natural language processing, as well as in the creation of intelligent text analysis systems.
About the Authors
A. P. KovalevskiiRussian Federation
Artyom P. Kovalevskii, DSc., Associate Professor, Leading Researcher
Yu. V. Pavlova
Russian Federation
Yulia V. Pavlova, Student
References
1. Gusarova G. V., Kovalevskii A. P., Makarenko A. G. Criteria for the existence of a change point.
2. Sibirskii Zhurnal Industrial’noi Matematiki, 2005, vol. 8, no. 4, pp. 18–33. (in Russ.)
3. Abebe B., Chebunin M., Kovalevskii A. Text Segmentation Via Processes that Count the Number of Different Words Forward and Backward. Journal of Quantitative Linguistics, vol. 31, no. 1, 2024, pp. 1–18. DOI: https://doi.org/10.1080/09296174.2023.2275342
4. Abebe B., Chebunin M., Kovalevskii A., Zakrevskaya N. Statistical tests for text homogeneity: Using forward and backward processes of numbers of different words. Glottometrics, vol. 53, no. 1, 2022, pp. 42–58. DOI: https://doi.org/10.53482/2022_53_401
5. Beeferman D., Berger A., Lafferty J. D. Statistical models for text segmentation. Machine Learning, vol. 34, no. 1–3, 1999, pp. 177–210. DOI: https://doi.org/10.1023/A:1007506220214
6. Choi F. Y. Y. Advances in domain independent linear text segmentation. 1st Meeting of the North American Chapter of the Association for Computational Linguistics. URL: https://aclanthology.org/A00-2004.pdf
7. Dembele S., Lo G. S. Probabilistic, Statistical and Algorithmic Aspects of the Similarity of Texts and Application to Gospels Comparison. Journal of Data Analysis and Information Processing, 2015, vol. 3, pp. 112–127. DOI: https://doi.org/10.4236/jdaip.2015.34012
8. Hearst M. A. Text tiling: Segmenting text into multi-paragraph subtopic passages. Computational Linguistics, 1997, vol. 23, no. 1, pp. 33–64.
9. Itoh N., Kurths J. Change-point detection of climate time series by nonparametric method. Proceedings of the World Congress on Engineering and Computer Science, 2010, vol. 1, pp. 445–448.
10. Mikolov T., Sutskever I., Chen K., Corrado G. S., Dean J. Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems, 2013, vol. 26, pp. 1–9.
11. Nanni G., Glavaš F., Ponzetto S. P. Unsupervised text segmentation using semantic relatedness graphs. In: Proceedings of the Fifth Joint Conference on Lexical and Computational Semantics (*SEM). Berlin, Germany, 2016, pp. 125–130.
12. Németh G., Zainkó C. Multilingual Statistical Text Analysis, Zipf’s Law and Hungarian Speech Generation. Acta Linguistica Hungarica, 2002, vol. 49, no. 3–4, pp. 385–405.
13. Pevzner L., Hearst M. A. A critique and improvement of an evaluation metric for text segmentation. Computational Linguistics, 2002, vol. 28, no. 1, pp. 19–36. DOI: https://doi.org/10.1162/089120102317341756
Review
For citations:
Kovalevskii A.P., Pavlova Yu.V. Author’s invariants in English literary texts. NSU Vestnik. Series: Linguistics and Intercultural Communication. 2026;24(1):102-112. (In Russ.) https://doi.org/10.25205/1818-7935-2026-24-1-102-112
JATS XML




















