Explorerโ€บData Scienceโ€บMachine Learning
Research PaperResearchia:202609.24005

Contrastive Learning for Authorship Verification

Peter Kirby

Abstract

Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors of model performance. Based on these considerations, we develop a ModernBERT Bi-Encoder model that achieves 98.4% accuracy on the PAN21 authorship verification task. --- Source: arXiv:2609.28471v1 - h...

Submitted: September 24, 2026Subjects: Machine Learning; Data Science

Description / Details

Our results show that contrastive learning outperforms a classification-based approach to authorship verification under the tested settings. We identify loss function, batch size, training duration, pre-trained model, input context length, and random text span data augmentation as important factors of model performance. Based on these considerations, we develop a ModernBERT Bi-Encoder model that achieves 98.4% accuracy on the PAN21 authorship verification task.


Source: arXiv:2609.28471v1 - http://arxiv.org/abs/2609.28471v1 PDF: https://arxiv.org/pdf/2609.28471v1 Original Link: http://arxiv.org/abs/2609.28471v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 24, 2026
Topic:
Data Science
Area:
Machine Learning
Comments:
0
Bookmark