A police sniffer dog in a warehouse walks past boxes of earlier releases of a software library, each marked with a green tick, and noses the latest release, labelled awesome-lib v1.2.3 with a red warning mark; the open box holds a device and a tag reading: added: 230 statements
Measurement Paper · Supply chain · October 2026

The change is the threat: 14 hijacked packages, and why you scan what changed

Fourteen real releases of popular npm and PyPI packages were hijacked and published with malicious code inside. Scanning each whole package for dangerous patterns found almost nothing. Measuring what changed from the release before put the attack in a single file in 7 of the 12 pairs that could be compared.

First release 6 October 2026 · CodeDelta engine 2.2.2 (build 52)

Abstract

A supply-chain attack on a package registry usually works the same way: an attacker gains publishing rights to a trusted package and releases a new version with a small addition. Fourteen such releases were taken from Datadog’s public dataset of malicious packages, each checked against the file hash the dataset publishes, and set beside the clean release that preceded it, downloaded from the official registry. CodeDelta’s Agent Scan, which looks for AI-agent patterns across a whole package, separated a hijacked release from its clean predecessor in one case. CodeDelta’s churn measurement, which compares the two releases statement by statement, showed the added code sitting in one file — 230 added statements in the only source file of five widely used npm libraries, a new 545-statement file in a sixth, 11 added statements in a Python package’s start-up file in a seventh. Churn says where the change is, not whether it is malicious. That is exactly the question a reviewer needs answered first.

In plain terms

When a trusted library is hijacked, the bad code arrives as an update. Searching the whole library for known dangerous patterns is looking for a needle in a haystack that is mostly the library’s own legitimate code. Comparing the update with the version before it removes the haystack: what is left is what changed. In the npm attack of September 2025, a routine patch release added 230 logical statements to a library whose code had otherwise not changed at all. That single number is the attack.

01Fourteen real hijacks

Datadog maintains a public research collection of malicious packages found on npm and PyPI, separated into packages built to be malicious and compromised releases of legitimate ones. The second group is the subject here: genuine, widely installed libraries whose publishing was taken over. Fourteen compromised releases were chosen, among them the September 2025 npm attack on debug, ansi-styles and their neighbours, the December 2024 @solana/web3.js key theft, and the self-spreading worm found in @ctrl/tinycolor. Each sample’s hash was checked against the one Datadog’s repository publishes; all fourteen matched.

For each, the comparison point is the release immediately before it, taken from the official npm or PyPI registry and checked against Datadog’s list of known-bad versions so that it is not itself compromised. Where the version just before was also bad or had been withdrawn, the last good release was used.

PackageRegistryClean releaseHijacked releaseFound (dataset date)
debugnpm4.4.14.4.28 Sep 2025
ansi-stylesnpm6.2.16.2.28 Sep 2025
color-convertnpm3.1.03.1.18 Sep 2025
error-exnpm1.3.21.3.38 Sep 2025
ansi-regexnpm6.2.06.2.18 Sep 2025
@solana/web3.jsnpm1.95.51.95.73 Dec 2024
@ctrl/tinycolornpm4.1.04.1.115 Sep 2025
ultralyticsPyPI8.3.408.3.414 Dec 2024
litellmPyPI1.82.61.82.724 Mar 2026
num2wordsPyPI0.5.140.5.1528 Jul 2025
mistralaiPyPI2.4.52.4.612 May 2026
telnyxPyPI4.87.04.87.127 Mar 2026
guardrails-aiPyPI0.10.00.10.112 May 2026
lightningPyPI2.6.12.6.230 Apr 2026

The dates are the discovery dates in the dataset’s file names, not necessarily the publication dates.

02Scanning the whole package found almost nothing

The first test was the obvious one: scan each hijacked release for dangerous code. CodeDelta’s Agent Scan looks for code that drives AI agents — model calls, a program running what a model tells it to, prompt-injection inputs — along with committed credentials and install hooks. It was run on every hijacked release and on its clean predecessor.

For the seven npm packages, where the two releases compare like for like, the scan flagged nothing in either version of six of them. In the seventh, @ctrl/tinycolor, it found an install hook and two credential-access matches that the clean release did not have, and rated one file ELEVATED. For the PyPI packages the result is no result: several of the dataset’s samples carry far more files than the release itself (git history, images, tests), so the two scans were not measuring the same thing and no conclusion is drawn from them.

That outcome is not a malfunction. These attacks steal cryptocurrency and credentials; they contain no AI agent, which is what the scan was built to find. But it illustrates the general problem with scanning a whole package: the malicious addition is a few hundred statements inside tens of thousands of legitimate ones, written to look ordinary.

03Looking at the change

The second test asked a different question: what is in the hijacked release that was not in the clean one? CodeDelta measures churn — the code changed, deleted and added between two versions — in logical statements (LLOC: a statement counts once however it is spread across lines; see the companion paper on why the unit matters). The released engine was run on each pair, clean release first.

PackageWhere the change isStatements addedFile after (statements)
debug 4.4.1 → 4.4.2src/index.js, the library’s entry point+230232
ansi-styles 6.2.1 → 6.2.2index.js, its only source file+230287
color-convert 3.1.0 → 3.1.1index.js+230261
error-ex 1.3.2 → 1.3.3index.js, its only source file+230285
ansi-regex 6.2.0 → 6.2.1index.js+230236
@ctrl/tinycolor 4.1.0 → 4.1.1a new file, bundle.js+545545
guardrails-ai 0.10.0 → 0.10.1guardrails/__init__.py, which Python runs when the package is imported+1124

Why five different libraries show the same number. The five September 2025 releases were not five attacks but one, copied: the attacker appended the same payload to each library’s main file. The proof is in the files. In every one of the five, the added code is a single line of 76,438 bytes, and its SHA-256 fingerprint is identical in all five: b54086257d7c8f87a652d53b7207ef040c84c9059839e008977944da30034f6a. The only other added lines are blank. The files themselves differ — 232, 287, 261, 285 and 236 statements after the attack, because each library’s own code is a different size — but the same 76,438-byte line counts as the same 230 logical statements wherever it lands.

In each of these seven, apart from package.json (version and publishing fields, and in tinycolor the install hook), that one file is the only code that changed. The five September 2025 releases make the point most sharply: the same 230 statements were added to five different libraries, and nothing else in their code moved. A patch-level release of a small, stable library that suddenly grows by hundreds of statements in its entry file is not a subtle signal. It needs no pattern, no signature and no knowledge of the attacker’s technique to see — only a measurement of the change.

For guardrails-ai the churn report also lists packaging files (PKG-INFO, README.md and others) and a metadata file the dataset adds to each sample. Those rows come from the two releases being stored in different download formats and are not code; the only code change is the start-up file.

04Where the picture is less clean

Five pairs are reported as they came out, without tidying.

05Method — reproduce it

Safety first. The samples are live malware. Every step here reads files; nothing is installed or executed. Even so, use a disposable machine or a temporary cloud runner, never a working computer, and never run npm install or pip install on a sample.

For each pair: fetch the hijacked release from the dataset (each sample is a zip encrypted with the password infected; check its hash against the one the repository publishes), fetch the clean release from the registry, unpack both, and compare clean against hijacked:

unzip -P infected 2025-09-08-debug-v4.4.2.zip -d hijacked
curl -sL https://registry.npmjs.org/debug/-/debug-4.4.1.tgz | tar xz -C clean

codedelta clean hijacked -o report.html --csv metrics.csv

The CSV lists every file with its changed, deleted and added statements. Measured with CodeDelta engine 2.2.2 (build 52), the public release, using its free evaluation licence. Where the dataset stores extra material beside a release, only the files the two releases share were compared, plus any additional file in a folder the clean release also has — which is where an injected file would sit; the extra material was left out and is not counted above.

The identical-payload check takes the clean and hijacked versions of each main file, lists the lines present only in the hijacked one, and fingerprints the longest of them with SHA-256 — a step any reader can repeat with diff and sha256sum.

06What to do with it

The honest limit. Churn shows where a change is and how large it is. It does not say whether the change is malicious: a legitimate feature release also adds code to entry files. What it does is cut the search from a whole package to a handful of files, which is the difference between a review that happens and one that does not. Seven of the twelve comparable pairs reduced to a single file; five did not reduce cleanly, for the reasons given above.

References

  1. Datadog, Malicious software packages dataset — github.com/DataDog/malicious-software-packages-dataset (samples under samples/npm/compromised_lib and samples/pypi/compromised_lib)
  2. npm registry — npmjs.com; Python Package Index — pypi.org (clean releases)
  3. Companion paper, Why LLOC is what really counts — paper-lloc.html (why the unit is the logical statement)
  4. Companion paper, True churn — paper-true-churn.html
  5. CodeDelta user guide, Agent Scan — user-guide.html