Fourteen real releases of popular npm and PyPI packages were hijacked and published with malicious code inside. Scanning each whole package for dangerous patterns found almost nothing. Measuring what changed from the release before put the attack in a single file in 7 of the 12 pairs that could be compared.
A supply-chain attack on a package registry usually works the same way: an attacker gains publishing rights to a trusted package and releases a new version with a small addition. Fourteen such releases were taken from Datadog’s public dataset of malicious packages, each checked against the file hash the dataset publishes, and set beside the clean release that preceded it, downloaded from the official registry. CodeDelta’s Agent Scan, which looks for AI-agent patterns across a whole package, separated a hijacked release from its clean predecessor in one case. CodeDelta’s churn measurement, which compares the two releases statement by statement, showed the added code sitting in one file — 230 added statements in the only source file of five widely used npm libraries, a new 545-statement file in a sixth, 11 added statements in a Python package’s start-up file in a seventh. Churn says where the change is, not whether it is malicious. That is exactly the question a reviewer needs answered first.
When a trusted library is hijacked, the bad code arrives as an update. Searching the whole library for known dangerous patterns is looking for a needle in a haystack that is mostly the library’s own legitimate code. Comparing the update with the version before it removes the haystack: what is left is what changed. In the npm attack of September 2025, a routine patch release added 230 logical statements to a library whose code had otherwise not changed at all. That single number is the attack.
Datadog maintains a public research collection of malicious packages found on npm and PyPI, separated into packages built to be malicious and compromised releases of legitimate ones. The second group is the subject here: genuine, widely installed libraries whose publishing was taken over. Fourteen compromised releases were chosen, among them the September 2025 npm attack on debug, ansi-styles and their neighbours, the December 2024 @solana/web3.js key theft, and the self-spreading worm found in @ctrl/tinycolor. Each sample’s hash was checked against the one Datadog’s repository publishes; all fourteen matched.
For each, the comparison point is the release immediately before it, taken from the official npm or PyPI registry and checked against Datadog’s list of known-bad versions so that it is not itself compromised. Where the version just before was also bad or had been withdrawn, the last good release was used.
| Package | Registry | Clean release | Hijacked release | Found (dataset date) |
|---|---|---|---|---|
debug | npm | 4.4.1 | 4.4.2 | 8 Sep 2025 |
ansi-styles | npm | 6.2.1 | 6.2.2 | 8 Sep 2025 |
color-convert | npm | 3.1.0 | 3.1.1 | 8 Sep 2025 |
error-ex | npm | 1.3.2 | 1.3.3 | 8 Sep 2025 |
ansi-regex | npm | 6.2.0 | 6.2.1 | 8 Sep 2025 |
@solana/web3.js | npm | 1.95.5 | 1.95.7 | 3 Dec 2024 |
@ctrl/tinycolor | npm | 4.1.0 | 4.1.1 | 15 Sep 2025 |
ultralytics | PyPI | 8.3.40 | 8.3.41 | 4 Dec 2024 |
litellm | PyPI | 1.82.6 | 1.82.7 | 24 Mar 2026 |
num2words | PyPI | 0.5.14 | 0.5.15 | 28 Jul 2025 |
mistralai | PyPI | 2.4.5 | 2.4.6 | 12 May 2026 |
telnyx | PyPI | 4.87.0 | 4.87.1 | 27 Mar 2026 |
guardrails-ai | PyPI | 0.10.0 | 0.10.1 | 12 May 2026 |
lightning | PyPI | 2.6.1 | 2.6.2 | 30 Apr 2026 |
The dates are the discovery dates in the dataset’s file names, not necessarily the publication dates.
The first test was the obvious one: scan each hijacked release for dangerous code. CodeDelta’s Agent Scan looks for code that drives AI agents — model calls, a program running what a model tells it to, prompt-injection inputs — along with committed credentials and install hooks. It was run on every hijacked release and on its clean predecessor.
For the seven npm packages, where the two releases compare like for like, the scan flagged nothing in either version of six of them. In the seventh, @ctrl/tinycolor, it found an install hook and two credential-access matches that the clean release did not have, and rated one file ELEVATED. For the PyPI packages the result is no result: several of the dataset’s samples carry far more files than the release itself (git history, images, tests), so the two scans were not measuring the same thing and no conclusion is drawn from them.
That outcome is not a malfunction. These attacks steal cryptocurrency and credentials; they contain no AI agent, which is what the scan was built to find. But it illustrates the general problem with scanning a whole package: the malicious addition is a few hundred statements inside tens of thousands of legitimate ones, written to look ordinary.
The second test asked a different question: what is in the hijacked release that was not in the clean one? CodeDelta measures churn — the code changed, deleted and added between two versions — in logical statements (LLOC: a statement counts once however it is spread across lines; see the companion paper on why the unit matters). The released engine was run on each pair, clean release first.
| Package | Where the change is | Statements added | File after (statements) |
|---|---|---|---|
debug 4.4.1 → 4.4.2 | src/index.js, the library’s entry point | +230 | 232 |
ansi-styles 6.2.1 → 6.2.2 | index.js, its only source file | +230 | 287 |
color-convert 3.1.0 → 3.1.1 | index.js | +230 | 261 |
error-ex 1.3.2 → 1.3.3 | index.js, its only source file | +230 | 285 |
ansi-regex 6.2.0 → 6.2.1 | index.js | +230 | 236 |
@ctrl/tinycolor 4.1.0 → 4.1.1 | a new file, bundle.js | +545 | 545 |
guardrails-ai 0.10.0 → 0.10.1 | guardrails/__init__.py, which Python runs when the package is imported | +11 | 24 |
Why five different libraries show the same number. The five September 2025 releases were not five attacks but one, copied: the attacker appended the same payload to each library’s main file. The proof is in the files. In every one of the five, the added code is a single line of 76,438 bytes, and its SHA-256 fingerprint is identical in all five: b54086257d7c8f87a652d53b7207ef040c84c9059839e008977944da30034f6a. The only other added lines are blank. The files themselves differ — 232, 287, 261, 285 and 236 statements after the attack, because each library’s own code is a different size — but the same 76,438-byte line counts as the same 230 logical statements wherever it lands.
In each of these seven, apart from package.json (version and publishing fields, and in tinycolor the install hook), that one file is the only code that changed. The five September 2025 releases make the point most sharply: the same 230 statements were added to five different libraries, and nothing else in their code moved. A patch-level release of a small, stable library that suddenly grows by hundreds of statements in its entry file is not a subtle signal. It needs no pattern, no signature and no knowledge of the attacker’s technique to see — only a measurement of the change.
For guardrails-ai the churn report also lists packaging files (PKG-INFO, README.md and others) and a metadata file the dataset adds to each sample. Those rows come from the two releases being stored in different download formats and are not code; the only code change is the start-up file.
Five pairs are reported as they came out, without tidying.
@solana/web3.js — the change is there, small and identical across each compiled bundle of the library, plus a one-statement change to src/connection.ts. But the release also ships source maps, which the comparison counts, and they make up most of the 402 churned statements. A reviewer would find the attack, but not at a glance.litellm — about fifteen source files changed, most plausibly ordinary release work. One of them, litellm/proxy/proxy_server.py, grew by roughly 142 KB, yet the churn report counts only 7 added statements, because the addition is packed into twelve very long lines. Statement counts measure logic, not bytes; a payload packed into a few lines can be large while counting small. Byte and line-length growth would show what statement counts hide.mistralai and telnyx — the dataset stores these with much of the project’s repository alongside the release, and the differences that creates outweigh the change itself. The added code is present (17 added statements in mistralai) but buried.ultralytics, num2words and lightning — could not be compared: the samples are stored as whole source repositories and no files lined up with the published release.npm install or pip install on a sample.
For each pair: fetch the hijacked release from the dataset (each sample is a zip encrypted with the password infected; check its hash against the one the repository publishes), fetch the clean release from the registry, unpack both, and compare clean against hijacked:
unzip -P infected 2025-09-08-debug-v4.4.2.zip -d hijacked
curl -sL https://registry.npmjs.org/debug/-/debug-4.4.1.tgz | tar xz -C clean
codedelta clean hijacked -o report.html --csv metrics.csv
The CSV lists every file with its changed, deleted and added statements. Measured with CodeDelta engine 2.2.2 (build 52), the public release, using its free evaluation licence. Where the dataset stores extra material beside a release, only the files the two releases share were compared, plus any additional file in a folder the clean release also has — which is where an injected file would sit; the extra material was left out and is not counted above.
The identical-payload check takes the clean and hijacked versions of each main file, lists the lines present only in the hijacked one, and fingerprints the longest of them with SHA-256 — a step any reader can repeat with diff and sha256sum.
__init__.py or an npm package’s entry file runs the moment the package is imported. Three of the plainly written attacks here sat in exactly those files.tinycolor. The rest went past it.samples/npm/compromised_lib and samples/pypi/compromised_lib)