Skip to content

Add a Salam (.salam) language extractor - #3862

Open
BaseMax wants to merge 7 commits into
Graphify-Labs:v8from
MaxFork:feat/salam-language-support
Open

BaseMax wants to merge 7 commits into
Graphify-Labs:v8from
MaxFork:feat/salam-language-support

Conversation

@BaseMax

@BaseMax BaseMax commented Sep 26, 2026 •

Copy link
Copy Markdown

What does this PR do?

Adds an extractor for Salam (.salam) source files, so a Salam codebase gets extracted into the graph instead of being skipped. Salam has no tree-sitter grammar, so this is a small token-based scanner instead (like the COBOL/Apex extractors), no new dependency.

Salam can be written in English or Persian (every keyword has both spellings), so the scanner recognizes both. It pulls out packages, imports, functions, structs + methods, enums, interfaces, impl, type aliases, constants, extern C functions, link libraries, layout blocks/components, and switch/match. @en/@fa annotations become name aliases so a Persian call resolves to its English definition and back. Cross-file resolution covers package-qualified calls, string-import aliases, and receiver-typed method calls.

Closes #3860

Type of change

  • New feature

Verification & Invariants

  • Read the CONTRIBUTING.md guide.
  • Reproduced the issue and identified the invariant.
  • Made the smallest fix necessary.
  • Added a regression test (if bug fix) or isolated boundary test.
  • Kept the PR description synchronized with the final implementation.
  • Documented any limitations / unsupported cases explicitly.

How was this tested?

Ran the extractor over the ~2600 .salam files in the Salam repo itself (English and Persian sources, stdlib and compiler) and checked every block opens/closes correctly. Added tests/test_salam_extractor.py (16 tests) covering declarations, cross-file calls, switch/match, and Persian keywords. Ran the full test suite.

uv run pytest tests/test_salam_extractor.py -q
uv run pytest tests -q

Limitations: a chained call on an unresolved return type (f(x).method()) doesn't resolve, since the extractor doesn't track return types through call chains.

Graphify-specific checklist

[x] I updated generated skill artifacts (uv run python -m tools.skillgen --bless) when changing their source fragments.

uv run python -m tools.skillgen --check # 134 artifacts, no drift

[x] I confirmed that AST/structural extraction remains deterministic (no ambient state dependencies like ENV variables).
[x] I reviewed changes for security implications (no unsafe interpolation into shell/Python).
[x] I confirmed no API keys or local-only graph data are included.
[x] I disclosed AI authorship in my commit messages.

Token-based scanner with English and Persian keywords. Extracts packages,
imports, functions, structs/methods, enums, interfaces, impl, extern, layout
and components, with call and type-reference edges and a cross-file resolver
for package calls, receiver types and @fa/@en aliases.
Drop removed English/Persian spellings (while, old Persian words), add
واردسازی, برابر/نابرابر, the repeat-step 'each'/هر rule, and Persian comma
and question mark.
Match the current compiler: add switch/ترابرد statement parsing (multi-value,
range and relational cases, else), recognize «» guillemet string literals,
and stop a keyword reused as a struct-literal field name (Type { if = true })
from being misread as a control block opener.
@github-actions

Copy link
Copy Markdown

Thanks for the pull request, @BaseMax. A maintainer will review it soon.

Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions.

A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic.

@graphify-labs graphify-labs Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 5 advisory finding(s) below merit a look before merge.


Graphify review — findings

Adds a Salam (.salam) language extractor implemented as a standalone token scanner with no third-party parser, wiring it into detection, analysis family mapping, the extractor registry, dispatch, and the reference resolver via extract_salam and resolve_salam_references. The scanner handles English and Persian keyword spellings, switch/case, guillemet «» strings, and folds Arabic yeh/kaf and ZWNJ so Persian and English identifiers canonicalize, emitting packages, imports, functions, structs/methods, enums, interfaces, impl, extern C, layout/components, plus call and type-reference edges with @en/@fa name aliases so cross-language call sites resolve to their definitions. Documents the new language in the README support table.

Worth a look

  • handle_impl returns None despite int cursor contract — graphify/extractors/salam.py · Escalate · high
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Resolver raises on null arity metadata — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Variadic marker is counted as an extra parameter — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Trailing comma counted as an extra call argument — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Top-level variable type references are parsed then discarded — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 3211 functions depend on the 555 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 715 callers, 48 callees
  • new: _rebuild_code() — 144 callers, 55 callees
  • new: detect() — 112 callers, 15 callees
  • new: to_obsidian() — 41 callers, 14 callees
  • new: _extract_generic() — 18 callers, 29 callees
  • new: save_manifest() — 42 callers, 12 callees
  • new: to_json() — 58 callers, 7 callees
  • new: extract_js() — 87 callers, 4 callees
  • …and 81 more — each is listed as a finding

Verification — 3211 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 2891 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

303 of 303 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — impact, full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — impact, full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — impact, full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — impact, full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — impact, full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — impact, full-run-safety
  • tests/test_chunking.py — impact, full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — impact, full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — impact, full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — impact, full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — impact, full-run-safety
  • tests/test_cross_language_call_resolution.py — impact, full-run-safety
  • tests/test_cross_repo_external_call_guards.py — impact, full-run-safety
  • tests/test_cross_repo_member_calls.py — impact, full-run-safety
  • … and 253 more

non-code file(s) changed (README.md, tests/fixtures/sample.salam, tests/fixtures/sample_fa.salam) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (README.md, graphify/extractors/__init__.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

· 89 more finding(s) on lines outside this diff (see the check run).

Handles makeCircle(r).area(), a local bound from a call, a struct literal,
and an 'as Type' cast, using the callee's declared return type across
files where needed. Deeper (2+ hop) chains stay unresolved rather than
guessed.
# Conflicts:
#	CHANGELOG.md

@graphify-labs graphify-labs Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 5 advisory finding(s) below merit a look before merge.


Graphify review — findings

Adds a Salam (.salam) language extractor via extract_salam, a dependency-free token scanner that emits packages, imports, functions, structs with methods, enums, interfaces, impl blocks, type aliases, constants, extern: C functions, link libraries, and layout blocks/components, along with calls and references edges, and registers it across detection, dispatch, and the analyze language-family map. Wires up resolve_salam_references as a cross-file resolver that binds pkg.Func(), string-import aliases, receiver-typed method calls, and type references — reading a callee's declared return type across files to resolve member calls chained off another call, a local variable, a struct literal, or an as Type cast, while leaving 2+-hop chains unresolved rather than guessing. Handles both English and Persian keyword spellings (including switch/case, the «» guillemet string form, and the repeat ... each N in i step spelling) and avoids misreading a keyword reused as a struct-literal field name as a control block.

Worth a look

  • Arity filter falls back to wrong-arity candidates — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Unqualified type refs resolve to unique type in unrelated package — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Qualified interface name is discarded in impl declarations — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Builtin casts overwrite earlier user-type casts — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Multi-token type names are parsed as only the first token — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 3446 functions depend on the 790 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 721 callers, 48 callees
  • new: _rebuild_code() — 144 callers, 55 callees
  • new: detect() — 112 callers, 15 callees
  • new: to_obsidian() — 41 callers, 14 callees
  • new: _extract_generic() — 18 callers, 29 callees
  • new: save_manifest() — 42 callers, 12 callees
  • new: to_json() — 58 callers, 7 callees
  • new: extract_js() — 87 callers, 4 callees
  • …and 81 more — each is listed as a finding

Verification — 3446 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 3126 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

303 of 303 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — impact, full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — impact, full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — impact, full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — impact, full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — impact, full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — impact, full-run-safety
  • tests/test_chunking.py — impact, full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — impact, full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — impact, full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — impact, full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — impact, full-run-safety
  • tests/test_cross_language_call_resolution.py — impact, full-run-safety
  • tests/test_cross_repo_external_call_guards.py — impact, full-run-safety
  • tests/test_cross_repo_member_calls.py — impact, full-run-safety
  • … and 253 more

non-code file(s) changed (CHANGELOG.md, README.md, tests/fixtures/sample.salam, tests/fixtures/sample_fa.salam) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (CHANGELOG.md, README.md, graphify/extractors/__init__.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

· 89 more finding(s) on lines outside this diff (see the check run).

Bare function names passed as arguments (register("/", home)) now resolve
to indirect_call edges, same-file or package-qualified across files, with
shadowing by a same-named param/local correctly excluded.

Also: this/این was missing from the keyword table (receiver chains for
this.field could merge wrong), and a multi-word method name was truncated
to its last word on call.

@graphify-labs graphify-labs Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graphify reviewed this change.

Worth a look — the grounded gate found no coupling regressions or blocking issues, but 2 advisory finding(s) below merit a look before merge.


Graphify review — findings

Adds a Salam (.salam) language extractor — a dependency-free token scanner (extract_salam) that emits packages, imports, functions, structs with methods, enums, interfaces, impl blocks, type aliases, constants, extern C functions, and layout/components across the full English and Persian keyword set (including switch/case, guillemet «» strings, and @en/@fa name aliases), wired into detection, dispatch, and the language-family maps. Its cross-file resolver (resolve_salam_references) binds pkg.Func() calls, string-import aliases, receiver-typed method calls, call-chain and cast return types, and struct-literal targets, resolving a bare function reference passed as an argument to an indirect_call edge (cross-file only when package-qualified, same-file otherwise to avoid wrong guesses). Fixes a param/local shadowing an import alias being misread as a package-qualified reference, restores multi-word method-name merging, and registers the omitted this/این keyword so this.field receiver chains resolve.

Worth a look

  • Call resolution links functions with too many required parameters — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
  • Public component flag is ignored — graphify/extractors/salam.py · Escalate · medium
    • agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Analysis details — impact, health, verification

Impact & health

Graphify review

Impact — 3464 functions depend on the 808 functions this change touches.

Health — this change adds coupling hotspots:

  • new: extract() — 723 callers, 48 callees
  • new: _rebuild_code() — 144 callers, 55 callees
  • new: detect() — 112 callers, 15 callees
  • new: to_obsidian() — 41 callers, 14 callees
  • new: _extract_generic() — 18 callers, 29 callees
  • new: save_manifest() — 42 callers, 12 callees
  • new: to_json() — 58 callers, 7 callees
  • new: extract_js() — 87 callers, 4 callees
  • …and 81 more — each is listed as a finding

Verification — 3464 functions in the blast radius were not formally verified this run (proofs are advisory here).

Gate & verification

graphify gate

PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.

Advisory (not blocking):

  • verification_scope: 3144 function(s) in the blast radius were not formally verified this run

Test selection

Test selection

303 of 303 test file(s) selected (100%) via static blast radius.

Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.

  • tests/test_affected_cli.py — full-run-safety
  • tests/test_affected_member_seed.py — full-run-safety
  • tests/test_agents_platform.py — full-run-safety
  • tests/test_analyze.py — impact, full-run-safety
  • tests/test_anthropic_custom_endpoint.py — full-run-safety
  • tests/test_antigravity_install.py — full-run-safety
  • tests/test_apm_fallback_version.py — full-run-safety
  • tests/test_architecture_doc.py — full-run-safety
  • tests/test_astro_extraction.py — impact, full-run-safety
  • tests/test_astro_import_ids.py — impact, full-run-safety
  • tests/test_atomic_canvas_export.py — impact, full-run-safety
  • tests/test_atomic_version_stamp.py — full-run-safety
  • tests/test_atomic_writes.py — impact, full-run-safety
  • tests/test_backend_env_isolation.py — full-run-safety
  • tests/test_backend_extras.py — full-run-safety
  • tests/test_benchmark.py — full-run-safety
  • tests/test_benchmark_raw_graph.py — full-run-safety
  • tests/test_build.py — impact, full-run-safety
  • tests/test_build_merge_dedup_scope.py — full-run-safety
  • tests/test_build_merge_hyperedges_and_prune.py — full-run-safety
  • tests/test_build_merge_shrink_guard.py — full-run-safety
  • tests/test_builtin_global_type_refs.py — impact, full-run-safety
  • tests/test_cache.py — full-run-safety
  • tests/test_callflow_html.py — full-run-safety
  • tests/test_cargo_introspect.py — impact, full-run-safety
  • tests/test_cargo_missing_manifest.py — full-run-safety
  • tests/test_carried_hyperedge_remap.py — impact, full-run-safety
  • tests/test_case_sensitive_resolution.py — impact, full-run-safety
  • tests/test_charmap_encoding.py — impact, full-run-safety
  • tests/test_chunking.py — impact, full-run-safety
  • tests/test_cjs_module_extension.py — impact, full-run-safety
  • tests/test_claude_cli_backend.py — impact, full-run-safety
  • tests/test_claude_md.py — full-run-safety
  • tests/test_cli_broken_pipe.py — full-run-safety
  • tests/test_cli_export.py — impact, full-run-safety
  • tests/test_cli_help.py — full-run-safety
  • tests/test_cluster.py — full-run-safety
  • tests/test_cobol_extractor.py — impact, full-run-safety
  • tests/test_codebuddy.py — full-run-safety
  • tests/test_community_hub_labels.py — full-run-safety
  • tests/test_community_labels_skill.py — impact, full-run-safety
  • tests/test_confidence.py — impact, full-run-safety
  • tests/test_corrupt_graph_json.py — full-run-safety
  • tests/test_cpp_nested_and_cli.py — impact, full-run-safety
  • tests/test_cpp_objc_cross_file_calls.py — impact, full-run-safety
  • tests/test_cpp_preprocess.py — full-run-safety
  • tests/test_cross_extension_reexport_self_cycle.py — impact, full-run-safety
  • tests/test_cross_language_call_resolution.py — impact, full-run-safety
  • tests/test_cross_repo_external_call_guards.py — impact, full-run-safety
  • tests/test_cross_repo_member_calls.py — impact, full-run-safety
  • … and 253 more

non-code file(s) changed (CHANGELOG.md, README.md, tests/fixtures/sample.salam, tests/fixtures/sample_fa.salam) → running the full suite for safety (a code graph can't see config/fixture/data deps)

changed code file(s) with no mapped test (CHANGELOG.md, README.md, graphify/extractors/__init__.py) — a coverage gap or a missing link — running the full suite rather than only the selected tests

Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.

· 89 more finding(s) on lines outside this diff (see the check run).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a Salam (.salam) language extractor

1 participant