A Function-level Dataset of Vulnerable and Fixed Source Code in JavaScript and TypeScript
Abstract
JavaScript and TypeScript are widely used in modern web development, making their security critical; however, automated vulnerability detection is often constrained by the availability of high-quality training data. Here we present JsVul, a dataset curated from seven major sources. Unlike generic multi-language datasets that may retain noise -- such as minified code and cosmetic edits -- JsVul utilizes a language-specific pipeline. We collected pre-fix and post-fix versions of files around secur...
Description / Details
JavaScript and TypeScript are widely used in modern web development, making their security critical; however, automated vulnerability detection is often constrained by the availability of high-quality training data. Here we present JsVul, a dataset curated from seven major sources. Unlike generic multi-language datasets that may retain noise -- such as minified code and cosmetic edits -- JsVul utilizes a language-specific pipeline. We collected pre-fix and post-fix versions of files around security fixes and, by filtering irrelevant artifacts and applying automated syntax normalization, isolated security-related changes. We ensured data integrity through multi-stage deduplication and heuristic-based labeling. Provided in a time-ordered JSONL format, JsVul supports robust model training in the JavaScript and TypeScript ecosystem and demonstrates the importance of language-aware preprocessing in building vulnerability datasets.
Source: arXiv:2609.38012v1 - http://arxiv.org/abs/2609.38012v1 PDF: https://arxiv.org/pdf/2609.38012v1 Original Link: http://arxiv.org/abs/2609.38012v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Sep 30, 2026
Computer Science
Cybersecurity
0