Authors: Assistant Professor S.Venkateswara Rao, G. Varalaxmi
Abstract: Significantly transformed the way scientific documents are generated, summarized, and paraphrased. Modern AI systems are capable of producing highly fluent and contextually meaningful scientific abstracts that often resemble human-written content,Natural Language Processing (NLP), plagiarism detection, academic integrity, and automated content verification. text similarity methods for evaluating A large dataset consisting of approximately 36,500 scientific abstracts is categorized into three groups to facilitate detailed similarity analysis. The proposed evaluation framework incorporates multiple categories of similarity techniques, including lexical similarity methods, character-based algorithms, statistical text representation approaches, semantic embedding models, and transformer-based language models. Traditional , Dice Coefficient, Smith-Waterman Algorithm, Levenshtein Distance, and TF-IDF are evaluated alongside modern embedding techniques including Word2Vec, FastText, GloVe, and BERT. To measure the effectiveness of these approaches, statistical indicators Experimental analysis demonstrates that contextual embedding models consistently achieve superior semantic similarity performance compared with conventional lexical methods. The findings also indicate that AI-generated abstracts preserve semantic meaning more effectively than lexical structure, highlighting the importance of contextual language models in scientific text evaluation. The proposed comparative analysis provides valuable insights for plagiarism detection systems, AI content verification, scientific publishing, and future NLP applications involving automated text similarity assessment.
