Security Finding: Unbounded edge allocation: cross-file import resolution creates FCK edges from a tiny corpus
Severity: HIGH (CVSS 7.1)
CWE: CWE-770
Repository: Graphify-Labs/graphify
Affected: graphify/extract.py:2235
Description
extract._resolve_cross_file_imports (extract.py:2123-2253) iterates every imported name and, for each, adds an edge from EVERY local class in the importing file (lines 2235-2247). Cost is O(files x imported_names x local_classes) with no cap on edges, nodes, or edges-per-node anywhere in extract()/build(). A 5.26 MB generated corpus (401 files: one module defining 2000 classes, 400 files each defining 8 classes and importing all 2000 names) produced 6,405,600 edges and 5,601 nodes, driving peak RSS to ~2.1-2.7 GB and 6.6-21s of work. The subsequent watch rebuild and to_json/to_cypher then materialize all of it again.
Impact
A small, easily authored source tree forces multi-gigabyte memory use and multi-million-edge graphs. Because graphify runs this rebuild automatically via the watch loop and the post-commit git hook, simply committing the crafted files triggers the exhaustion. graph.json/GRAPH_REPORT.md/observed outputs grow proportionally, exhausting disk as well.
Remediation
Cross-file import resolution no longer emits a cartesian product of every imported name by every class in the importing file. Each class node is now mapped to the identifiers its own body references, and a uses edge is emitted only for classes that actually mention the imported name (or its alias), with (source, target) pairs deduplicated so repeated import statements cannot multiply edges. A hard global budget (_MAX_CROSS_FILE_EDGES = 100,000) aborts with a clear ValueError before more memory is allocated, so a small crafted corpus can no longer drive multi-gigabyte extraction/build memory or multi-million-edge graph.json artifacts.
Affected Code
graphify-main/graphify/extract.py:2235
line = node.start_point[0] + 1
for name in imported_names:
tgt_nid = stem_to_entities[target_stem].get(name)
if tgt_nid:
for src_class_nid in local_classes:
new_edges.append({
"source": src_class_nid,
"target": tgt_nid,
"relation": "uses",
"confidence": "INFERRED",
"source_file": str_path,
"source_location": f"L{line}",
"weight": 0.8,
})
Verification
Adversarially verified (GLM) — passed.
Reported by OpenClaw BountyBot via Failsafe Nexus (Pandora) automated security analysis. Please review carefully before acting.
Security Finding: Unbounded edge allocation: cross-file import resolution creates FCK edges from a tiny corpus
Severity: HIGH (CVSS 7.1)
CWE: CWE-770
Repository: Graphify-Labs/graphify
Affected:
graphify/extract.py:2235Description
extract._resolve_cross_file_imports (extract.py:2123-2253) iterates every imported name and, for each, adds an edge from EVERY local class in the importing file (lines 2235-2247). Cost is O(files x imported_names x local_classes) with no cap on edges, nodes, or edges-per-node anywhere in extract()/build(). A 5.26 MB generated corpus (401 files: one module defining 2000 classes, 400 files each defining 8 classes and importing all 2000 names) produced 6,405,600 edges and 5,601 nodes, driving peak RSS to ~2.1-2.7 GB and 6.6-21s of work. The subsequent watch rebuild and to_json/to_cypher then materialize all of it again.
Impact
A small, easily authored source tree forces multi-gigabyte memory use and multi-million-edge graphs. Because graphify runs this rebuild automatically via the watch loop and the post-commit git hook, simply committing the crafted files triggers the exhaustion. graph.json/GRAPH_REPORT.md/observed outputs grow proportionally, exhausting disk as well.
Remediation
Cross-file import resolution no longer emits a cartesian product of every imported name by every class in the importing file. Each class node is now mapped to the identifiers its own body references, and a
usesedge is emitted only for classes that actually mention the imported name (or its alias), with (source, target) pairs deduplicated so repeated import statements cannot multiply edges. A hard global budget (_MAX_CROSS_FILE_EDGES = 100,000) aborts with a clear ValueError before more memory is allocated, so a small crafted corpus can no longer drive multi-gigabyte extraction/build memory or multi-million-edge graph.json artifacts.Affected Code
Verification
Adversarially verified (GLM) — passed.
Reported by OpenClaw BountyBot via Failsafe Nexus (Pandora) automated security analysis. Please review carefully before acting.