From e5f37685dabb32f05ab8f78af0eb42d7f7404d11 Mon Sep 17 00:00:00 2001 From: Aman kumar <168410465+ghostiee-11@users.noreply.github.com> Date: Wed, 22 Jul 2026 01:05:46 +0530 Subject: [PATCH 1/8] docs: serve an llms.txt index of the documentation Zensical copies static files from docs/ to the site root, so this needs no build changes. A test keeps the link list in sync with the pages that actually exist. --- docs/llms.txt | 80 ++++++++++++++++++++++++++++++++++++++++ lumen/tests/test_docs.py | 32 ++++++++++++++++ 2 files changed, 112 insertions(+) create mode 100644 docs/llms.txt create mode 100644 lumen/tests/test_docs.py diff --git a/docs/llms.txt b/docs/llms.txt new file mode 100644 index 000000000..260c2e955 --- /dev/null +++ b/docs/llms.txt @@ -0,0 +1,80 @@ +# Lumen + +> Lumen is an open-source and extensible agent framework for chatting with data and for retrieval augmented generation. It turns natural language into SQL, data transformation pipelines, visualizations and dashboards, and every step stays inspectable, editable and reproducible. + +Lumen is built on Panel and the wider HoloViz stack. Its declarative data model means anything the language model produces can be serialized to YAML, reopened in a notebook, or composed into a dashboard. It connects to local files, SQL databases, data lakes and multidimensional array data, and it can be extended with custom sources, agents, tools and analyses. + +Install with `pip install 'lumen[ai-openai]'` and start with `lumen-ai serve data.csv`. + +## Getting started + +- [Quick Start](https://lumen.holoviz.org/quick_start/): install, serve a dataset and ask the first question +- [Installation](https://lumen.holoviz.org/installation/): install options for each LLM provider +- [Launching Lumen](https://lumen.holoviz.org/getting_started/launching_lumen/): the `lumen-ai` command line interface +- [Using Lumen AI](https://lumen.holoviz.org/getting_started/using_lumen_ai/): the conversational workflow end to end +- [Navigating the UI](https://lumen.holoviz.org/getting_started/navigating_the_ui/): explorations, tables, charts and reports +- [Building Lumen Apps](https://lumen.holoviz.org/getting_started/building_lumen_apps/): embedding Lumen in your own application +- [FAQ](https://lumen.holoviz.org/faq/) + +## Configuration + +- [CLI](https://lumen.holoviz.org/configuration/cli/): all `lumen-ai serve` flags +- [Data Sources](https://lumen.holoviz.org/configuration/sources/): DuckDB, Snowflake, BigQuery, SQLAlchemy, Intake, Prometheus, xarray +- [LLM Providers](https://lumen.holoviz.org/configuration/llm_providers/): OpenAI, Anthropic, Google, Mistral, Bedrock, Ollama, LlamaCpp and others +- [Agents](https://lumen.holoviz.org/configuration/agents/): the agents that generate SQL, charts and prose +- [Tools](https://lumen.holoviz.org/configuration/tools/): lookup tools and MCP server integration +- [Analyses](https://lumen.holoviz.org/configuration/analyses/): packaging domain logic as a reusable analysis +- [Coordinator](https://lumen.holoviz.org/configuration/coordinators/): how a request is planned and dispatched +- [Prompts](https://lumen.holoviz.org/configuration/prompts/): overriding the prompt templates +- [Context](https://lumen.holoviz.org/configuration/context/): what the model is told about your data +- [Reports](https://lumen.holoviz.org/configuration/reports/): assembling and exporting reports +- [UI](https://lumen.holoviz.org/configuration/ui/) and [Controls](https://lumen.holoviz.org/configuration/controls/): customizing the interface +- [Embeddings](https://lumen.holoviz.org/configuration/embeddings/) and [Vector Stores](https://lumen.holoviz.org/configuration/vector_stores/): retrieval configuration +- [Extending Lumen](https://lumen.holoviz.org/extending/): custom sources, agents, tools and views + +## Declarative specification + +- [Overview](https://lumen.holoviz.org/configuration/spec/): the YAML specification +- [Core Concepts](https://lumen.holoviz.org/configuration/spec/concepts/) +- [Loading Data](https://lumen.holoviz.org/configuration/spec/sources/) +- [Transforming Data](https://lumen.holoviz.org/configuration/spec/pipelines/) +- [Visualizing Data](https://lumen.holoviz.org/configuration/spec/views/) +- [Variables and References](https://lumen.holoviz.org/configuration/spec/variables/) +- [Custom Components](https://lumen.holoviz.org/configuration/spec/customization/) +- [Python API](https://lumen.holoviz.org/configuration/spec/python-api/) +- [Data Downloads](https://lumen.holoviz.org/configuration/spec/downloads/) +- [Authentication](https://lumen.holoviz.org/configuration/spec/authentication/) +- [Deployment](https://lumen.holoviz.org/configuration/spec/deployment/) + +## Examples + +- [Tutorials](https://lumen.holoviz.org/examples/tutorials/): worked end-to-end walkthroughs +- [Weather Data AI Explorer](https://lumen.holoviz.org/examples/tutorials/weather_data_ai_explorer/) +- [SaaS Company Report Dashboard](https://lumen.holoviz.org/examples/tutorials/saas_company_report_dashboard/) +- [Census Data AI Explorer](https://lumen.holoviz.org/examples/tutorials/census_data_ai_explorer/) +- [Mesonet Weather Explorer](https://lumen.holoviz.org/examples/tutorials/mesonet_weather_explorer/) +- [Gallery](https://lumen.holoviz.org/examples/gallery/): short declarative specification examples + +## API reference + +- [AI Agents](https://lumen.holoviz.org/reference/api/ai/agents/) +- [AI Coordinator](https://lumen.holoviz.org/reference/api/ai/coordinator/) +- [AI Tools](https://lumen.holoviz.org/reference/api/ai/tools/) +- [AI Models](https://lumen.holoviz.org/reference/api/ai/models/) +- [AI Core](https://lumen.holoviz.org/reference/api/ai/core/) +- [Pipeline](https://lumen.holoviz.org/reference/api/pipeline/) +- [Sources](https://lumen.holoviz.org/reference/api/sources/) +- [Transforms](https://lumen.holoviz.org/reference/api/transforms/) +- [Views](https://lumen.holoviz.org/reference/api/views/) + +## Optional + +- [Penguins](https://lumen.holoviz.org/examples/gallery/penguins/) +- [London Bike Points](https://lumen.holoviz.org/examples/gallery/bikes/) +- [Seattle Weather](https://lumen.holoviz.org/examples/gallery/seattle/) +- [NYC Taxi](https://lumen.holoviz.org/examples/gallery/nyc_taxi/) +- [Precipitation](https://lumen.holoviz.org/examples/gallery/precip/) +- [Earthquakes](https://lumen.holoviz.org/examples/gallery/earthquakes/) +- [Wind Turbines](https://lumen.holoviz.org/examples/gallery/windturbines/) +- [Releases](https://lumen.holoviz.org/releases/) +- [Contributing](https://lumen.holoviz.org/contributing/) diff --git a/lumen/tests/test_docs.py b/lumen/tests/test_docs.py new file mode 100644 index 000000000..df6dd9dbd --- /dev/null +++ b/lumen/tests/test_docs.py @@ -0,0 +1,32 @@ +import re + +from pathlib import Path + +import pytest + +DOCS = Path(__file__).parents[2] / "docs" +LLMS_TXT = DOCS / "llms.txt" +SITE_URL = "https://lumen.holoviz.org/" + +pytestmark = pytest.mark.skipif( + not LLMS_TXT.is_file(), reason="docs directory is not available" +) + + +def _page(url): + """Map a documentation URL back to the Markdown file that builds it.""" + path = url.removeprefix(SITE_URL).strip("/") + return DOCS / (f"{path}.md" if path else "index.md"), DOCS / path / "index.md" + + +def test_llms_txt_links_resolve(): + urls = re.findall(rf"\]\(({re.escape(SITE_URL)}[^)]*)\)", LLMS_TXT.read_text()) + assert urls, "llms.txt lists no documentation pages" + missing = [url for url in urls if not any(p.is_file() for p in _page(url))] + assert not missing, f"llms.txt links to pages that do not exist: {missing}" + + +def test_llms_txt_follows_spec(): + lines = [line for line in LLMS_TXT.read_text().splitlines() if line.strip()] + assert lines[0].startswith("# "), "llms.txt must open with an H1 project name" + assert lines[1].startswith("> "), "llms.txt must follow the H1 with a summary blockquote" From c2091ee5f8ec3f1d4ed9b07775299152b4b3e2a7 Mon Sep 17 00:00:00 2001 From: ghostiee-11 Date: Fri, 24 Jul 2026 22:48:49 +0530 Subject: [PATCH 2/8] docs: generate llms.txt and a markdown mirror at build time Replaces the hand-written docs/llms.txt with a build step, following the approach in holoviz/hvplot#1732. The zensical nav is the single source of truth, so pages and labels can never drift from the site, and the linked markdown is mirrored under /markdown so an LLM can fetch the prose without parsing the rendered HTML. --- docs/llms.txt | 80 ----------------------------------- lumen/tests/test_docs.py | 53 +++++++++++++++++------- pixi.toml | 4 +- scripts/build_llms_txt.py | 87 +++++++++++++++++++++++++++++++++++++++ 4 files changed, 127 insertions(+), 97 deletions(-) delete mode 100644 docs/llms.txt create mode 100644 scripts/build_llms_txt.py diff --git a/docs/llms.txt b/docs/llms.txt deleted file mode 100644 index 260c2e955..000000000 --- a/docs/llms.txt +++ /dev/null @@ -1,80 +0,0 @@ -# Lumen - -> Lumen is an open-source and extensible agent framework for chatting with data and for retrieval augmented generation. It turns natural language into SQL, data transformation pipelines, visualizations and dashboards, and every step stays inspectable, editable and reproducible. - -Lumen is built on Panel and the wider HoloViz stack. Its declarative data model means anything the language model produces can be serialized to YAML, reopened in a notebook, or composed into a dashboard. It connects to local files, SQL databases, data lakes and multidimensional array data, and it can be extended with custom sources, agents, tools and analyses. - -Install with `pip install 'lumen[ai-openai]'` and start with `lumen-ai serve data.csv`. - -## Getting started - -- [Quick Start](https://lumen.holoviz.org/quick_start/): install, serve a dataset and ask the first question -- [Installation](https://lumen.holoviz.org/installation/): install options for each LLM provider -- [Launching Lumen](https://lumen.holoviz.org/getting_started/launching_lumen/): the `lumen-ai` command line interface -- [Using Lumen AI](https://lumen.holoviz.org/getting_started/using_lumen_ai/): the conversational workflow end to end -- [Navigating the UI](https://lumen.holoviz.org/getting_started/navigating_the_ui/): explorations, tables, charts and reports -- [Building Lumen Apps](https://lumen.holoviz.org/getting_started/building_lumen_apps/): embedding Lumen in your own application -- [FAQ](https://lumen.holoviz.org/faq/) - -## Configuration - -- [CLI](https://lumen.holoviz.org/configuration/cli/): all `lumen-ai serve` flags -- [Data Sources](https://lumen.holoviz.org/configuration/sources/): DuckDB, Snowflake, BigQuery, SQLAlchemy, Intake, Prometheus, xarray -- [LLM Providers](https://lumen.holoviz.org/configuration/llm_providers/): OpenAI, Anthropic, Google, Mistral, Bedrock, Ollama, LlamaCpp and others -- [Agents](https://lumen.holoviz.org/configuration/agents/): the agents that generate SQL, charts and prose -- [Tools](https://lumen.holoviz.org/configuration/tools/): lookup tools and MCP server integration -- [Analyses](https://lumen.holoviz.org/configuration/analyses/): packaging domain logic as a reusable analysis -- [Coordinator](https://lumen.holoviz.org/configuration/coordinators/): how a request is planned and dispatched -- [Prompts](https://lumen.holoviz.org/configuration/prompts/): overriding the prompt templates -- [Context](https://lumen.holoviz.org/configuration/context/): what the model is told about your data -- [Reports](https://lumen.holoviz.org/configuration/reports/): assembling and exporting reports -- [UI](https://lumen.holoviz.org/configuration/ui/) and [Controls](https://lumen.holoviz.org/configuration/controls/): customizing the interface -- [Embeddings](https://lumen.holoviz.org/configuration/embeddings/) and [Vector Stores](https://lumen.holoviz.org/configuration/vector_stores/): retrieval configuration -- [Extending Lumen](https://lumen.holoviz.org/extending/): custom sources, agents, tools and views - -## Declarative specification - -- [Overview](https://lumen.holoviz.org/configuration/spec/): the YAML specification -- [Core Concepts](https://lumen.holoviz.org/configuration/spec/concepts/) -- [Loading Data](https://lumen.holoviz.org/configuration/spec/sources/) -- [Transforming Data](https://lumen.holoviz.org/configuration/spec/pipelines/) -- [Visualizing Data](https://lumen.holoviz.org/configuration/spec/views/) -- [Variables and References](https://lumen.holoviz.org/configuration/spec/variables/) -- [Custom Components](https://lumen.holoviz.org/configuration/spec/customization/) -- [Python API](https://lumen.holoviz.org/configuration/spec/python-api/) -- [Data Downloads](https://lumen.holoviz.org/configuration/spec/downloads/) -- [Authentication](https://lumen.holoviz.org/configuration/spec/authentication/) -- [Deployment](https://lumen.holoviz.org/configuration/spec/deployment/) - -## Examples - -- [Tutorials](https://lumen.holoviz.org/examples/tutorials/): worked end-to-end walkthroughs -- [Weather Data AI Explorer](https://lumen.holoviz.org/examples/tutorials/weather_data_ai_explorer/) -- [SaaS Company Report Dashboard](https://lumen.holoviz.org/examples/tutorials/saas_company_report_dashboard/) -- [Census Data AI Explorer](https://lumen.holoviz.org/examples/tutorials/census_data_ai_explorer/) -- [Mesonet Weather Explorer](https://lumen.holoviz.org/examples/tutorials/mesonet_weather_explorer/) -- [Gallery](https://lumen.holoviz.org/examples/gallery/): short declarative specification examples - -## API reference - -- [AI Agents](https://lumen.holoviz.org/reference/api/ai/agents/) -- [AI Coordinator](https://lumen.holoviz.org/reference/api/ai/coordinator/) -- [AI Tools](https://lumen.holoviz.org/reference/api/ai/tools/) -- [AI Models](https://lumen.holoviz.org/reference/api/ai/models/) -- [AI Core](https://lumen.holoviz.org/reference/api/ai/core/) -- [Pipeline](https://lumen.holoviz.org/reference/api/pipeline/) -- [Sources](https://lumen.holoviz.org/reference/api/sources/) -- [Transforms](https://lumen.holoviz.org/reference/api/transforms/) -- [Views](https://lumen.holoviz.org/reference/api/views/) - -## Optional - -- [Penguins](https://lumen.holoviz.org/examples/gallery/penguins/) -- [London Bike Points](https://lumen.holoviz.org/examples/gallery/bikes/) -- [Seattle Weather](https://lumen.holoviz.org/examples/gallery/seattle/) -- [NYC Taxi](https://lumen.holoviz.org/examples/gallery/nyc_taxi/) -- [Precipitation](https://lumen.holoviz.org/examples/gallery/precip/) -- [Earthquakes](https://lumen.holoviz.org/examples/gallery/earthquakes/) -- [Wind Turbines](https://lumen.holoviz.org/examples/gallery/windturbines/) -- [Releases](https://lumen.holoviz.org/releases/) -- [Contributing](https://lumen.holoviz.org/contributing/) diff --git a/lumen/tests/test_docs.py b/lumen/tests/test_docs.py index df6dd9dbd..a95f666dd 100644 --- a/lumen/tests/test_docs.py +++ b/lumen/tests/test_docs.py @@ -1,32 +1,53 @@ -import re +import sys +import tomllib from pathlib import Path import pytest -DOCS = Path(__file__).parents[2] / "docs" -LLMS_TXT = DOCS / "llms.txt" -SITE_URL = "https://lumen.holoviz.org/" +ROOT = Path(__file__).parents[2] +DOCS = ROOT / "docs" +CONFIG_FILE = ROOT / "zensical.toml" +SCRIPTS = ROOT / "scripts" pytestmark = pytest.mark.skipif( - not LLMS_TXT.is_file(), reason="docs directory is not available" + not CONFIG_FILE.is_file(), reason="docs directory is not available" ) -def _page(url): - """Map a documentation URL back to the Markdown file that builds it.""" - path = url.removeprefix(SITE_URL).strip("/") - return DOCS / (f"{path}.md" if path else "index.md"), DOCS / path / "index.md" +@pytest.fixture(scope="module") +def build_llms_txt(): + """Import the build script, which lives outside the installed package.""" + sys.path.insert(0, str(SCRIPTS)) + try: + import build_llms_txt + finally: + sys.path.remove(str(SCRIPTS)) + return build_llms_txt -def test_llms_txt_links_resolve(): - urls = re.findall(rf"\]\(({re.escape(SITE_URL)}[^)]*)\)", LLMS_TXT.read_text()) - assert urls, "llms.txt lists no documentation pages" - missing = [url for url in urls if not any(p.is_file() for p in _page(url))] - assert not missing, f"llms.txt links to pages that do not exist: {missing}" +@pytest.fixture(scope="module") +def pages(build_llms_txt): + nav = tomllib.loads(CONFIG_FILE.read_text(encoding="utf-8"))["project"]["nav"] + return build_llms_txt.iter_nav_pages(nav) -def test_llms_txt_follows_spec(): - lines = [line for line in LLMS_TXT.read_text().splitlines() if line.strip()] +def test_nav_pages_exist(pages): + """Every page llms.txt advertises has to be a file the build can copy.""" + assert pages, "zensical nav lists no documentation pages" + missing = [path for _, _, path in pages if not (DOCS / path).is_file()] + assert not missing, f"nav lists pages that do not exist: {missing}" + + +def test_llms_txt_follows_spec(build_llms_txt, pages): + rendered = build_llms_txt.render_llms_txt(pages) + lines = [line for line in rendered.splitlines() if line.strip()] assert lines[0].startswith("# "), "llms.txt must open with an H1 project name" assert lines[1].startswith("> "), "llms.txt must follow the H1 with a summary blockquote" + + +def test_llms_txt_links_every_nav_page(build_llms_txt, pages): + rendered = build_llms_txt.render_llms_txt(pages) + url = build_llms_txt.MARKDOWN_URL + unlinked = [path for _, _, path in pages if f"({url}/{path})" not in rendered] + assert not unlinked, f"llms.txt omits nav pages: {unlinked}" diff --git a/pixi.toml b/pixi.toml index f348e64be..5d3b1bfd8 100644 --- a/pixi.toml +++ b/pixi.toml @@ -117,7 +117,9 @@ mkdocs-literate-nav = "*" [feature.doc.tasks] # Updated tasks to use zensical instead of mkdocs -docs-build = { cmd = 'zensical build --clean', depends-on = ['install'] } +_docs-build-site = { cmd = 'zensical build --clean', depends-on = ['install'] } +# llms.txt is written into builtdocs, so it has to run after --clean has wiped it. +docs-build = { cmd = 'python scripts/build_llms_txt.py', depends-on = ['_docs-build-site'] } docs-serve = { cmd = 'zensical serve', depends-on = ['install'] } docs-clean = { cmd = 'rm -rf builtdocs', depends-on = ['install'] } diff --git a/scripts/build_llms_txt.py b/scripts/build_llms_txt.py new file mode 100644 index 000000000..6de77d008 --- /dev/null +++ b/scripts/build_llms_txt.py @@ -0,0 +1,87 @@ +""" +Build the markdown mirror and llms.txt that Lumen serves for LLM consumers. + +Runs as the last step of the docs build: + pixi run -e docs docs-build +""" + +import shutil +import tomllib + +from pathlib import Path + +ROOT = Path(__file__).parent.parent +CONFIG_FILE = ROOT / "zensical.toml" +DOCS_DIR = ROOT / "docs" +MARKDOWN_URL = "/markdown" + +SUMMARY = ( + "Lumen is an open-source and extensible agent framework for chatting with data " + "and for retrieval augmented generation. It turns natural language into SQL, data " + "transformation pipelines, visualizations and dashboards, and every step stays " + "inspectable, editable and reproducible." +) + +PREAMBLE = [ + "Lumen is built on Panel and the wider HoloViz stack. Its declarative data model means " + "anything the language model produces can be serialized to YAML, reopened in a notebook, " + "or composed into a dashboard.", + "", + "Install with `pip install 'lumen[ai-openai]'` and start with `lumen-ai serve data.csv`.", +] + +# Nav entries that sit outside any group, e.g. Quick Start and Installation. +UNGROUPED_SECTION = "Overview" + + +def iter_nav_pages(nav: list, trail: tuple[str, ...] = ()) -> list[tuple[str, str, str]]: + """ + Flatten zensical's nav into (section, label, path) triples. + + The nav is the single source of truth here: it decides which pages reach + llms.txt and what they are called, so no label or grouping is inferred + from file paths. + """ + pages = [] + for entry in nav: + for label, target in entry.items(): + if isinstance(target, list): + pages.extend(iter_nav_pages(target, trail + (label,))) + else: + pages.append((" / ".join(trail) or UNGROUPED_SECTION, label, target)) + return pages + + +def copy_markdown(pages: list[tuple[str, str, str]], output_dir: Path) -> None: + """Mirror every listed page into the built site so the links resolve.""" + for _, _, path in pages: + destination = output_dir / path + destination.parent.mkdir(parents=True, exist_ok=True) + shutil.copy2(DOCS_DIR / path, destination) + + +def render_llms_txt(pages: list[tuple[str, str, str]]) -> str: + lines = ["# Lumen", "", f"> {SUMMARY}", "", *PREAMBLE, ""] + for section in dict.fromkeys(section for section, _, _ in pages): + lines.extend([f"## {section}", ""]) + lines.extend( + f"- [{label}]({MARKDOWN_URL}/{path})" + for page_section, label, path in pages + if page_section == section + ) + lines.append("") + return "\n".join(lines) + + +def main() -> None: + config = tomllib.loads(CONFIG_FILE.read_text(encoding="utf-8"))["project"] + site_dir = ROOT / config["site_dir"] + pages = iter_nav_pages(config["nav"]) + + copy_markdown(pages, site_dir / MARKDOWN_URL.strip("/")) + (site_dir / "llms.txt").write_text(render_llms_txt(pages), encoding="utf-8") + print(f"Wrote llms.txt and {len(pages)} markdown pages to {site_dir}") + + +if __name__ == "__main__": + main() From c97a323b8cebf6e65cf8bbc333ad98102492ce99 Mon Sep 17 00:00:00 2001 From: ghostiee-11 Date: Fri, 24 Jul 2026 22:55:44 +0530 Subject: [PATCH 3/8] docs: add the four unlisted tutorials to the nav The tutorials index links seven tutorials but the nav only carried four, so the rest were published without a sidebar entry and were skipped by llms.txt, which is generated from the nav. Labels match the index page. --- zensical.toml | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/zensical.toml b/zensical.toml index e9dd8803b..9196f7887 100644 --- a/zensical.toml +++ b/zensical.toml @@ -26,9 +26,10 @@ nav = [ {"SaaS Company Report Dashboard" = "examples/tutorials/saas_company_report_dashboard.md"}, {"Census Data AI Explorer" = "examples/tutorials/census_data_ai_explorer.md"}, {"Mesonet Weather Explorer" = "examples/tutorials/mesonet_weather_explorer.md"}, - {"Weather API Explorer" = "examples/tutorials/weather_openapi_explorer.md"}, - {"Massive Stock Explorer" = "examples/tutorials/massive_stock_explorer.md"}, - {"Build Dashboard with Spec" = "examples/tutorials/penguins_dashboard_spec.md"} + {"Weather API Explorer (OpenAPI)" = "examples/tutorials/weather_openapi_explorer.md"}, + {"Stock Market Explorer" = "examples/tutorials/massive_stock_explorer.md"}, + {"Build Dashboard with Spec" = "examples/tutorials/penguins_dashboard_spec.md"}, + {"SaaS Executive Dashboard" = "examples/tutorials/saas_metrics_reports.md"} ]}, {"Gallery" = [ {"Overview" = "examples/gallery/index.md"}, From fbdd7a2ac72250eefc87a709ae2d8d5a83b054f2 Mon Sep 17 00:00:00 2001 From: Aman Kumar Date: Mon, 24 Aug 2026 22:48:43 +0530 Subject: [PATCH 4/8] docs: build llms.txt with nbsite's reusable builder --- lumen/tests/test_docs.py | 53 ----------------- pixi.toml | 4 +- scripts/build_llms_txt.py | 87 ---------------------------- scripts/llms_config.py | 118 ++++++++++++++++++++++++++++++++++++++ 4 files changed, 121 insertions(+), 141 deletions(-) delete mode 100644 lumen/tests/test_docs.py delete mode 100644 scripts/build_llms_txt.py create mode 100644 scripts/llms_config.py diff --git a/lumen/tests/test_docs.py b/lumen/tests/test_docs.py deleted file mode 100644 index a95f666dd..000000000 --- a/lumen/tests/test_docs.py +++ /dev/null @@ -1,53 +0,0 @@ -import sys -import tomllib - -from pathlib import Path - -import pytest - -ROOT = Path(__file__).parents[2] -DOCS = ROOT / "docs" -CONFIG_FILE = ROOT / "zensical.toml" -SCRIPTS = ROOT / "scripts" - -pytestmark = pytest.mark.skipif( - not CONFIG_FILE.is_file(), reason="docs directory is not available" -) - - -@pytest.fixture(scope="module") -def build_llms_txt(): - """Import the build script, which lives outside the installed package.""" - sys.path.insert(0, str(SCRIPTS)) - try: - import build_llms_txt - finally: - sys.path.remove(str(SCRIPTS)) - return build_llms_txt - - -@pytest.fixture(scope="module") -def pages(build_llms_txt): - nav = tomllib.loads(CONFIG_FILE.read_text(encoding="utf-8"))["project"]["nav"] - return build_llms_txt.iter_nav_pages(nav) - - -def test_nav_pages_exist(pages): - """Every page llms.txt advertises has to be a file the build can copy.""" - assert pages, "zensical nav lists no documentation pages" - missing = [path for _, _, path in pages if not (DOCS / path).is_file()] - assert not missing, f"nav lists pages that do not exist: {missing}" - - -def test_llms_txt_follows_spec(build_llms_txt, pages): - rendered = build_llms_txt.render_llms_txt(pages) - lines = [line for line in rendered.splitlines() if line.strip()] - assert lines[0].startswith("# "), "llms.txt must open with an H1 project name" - assert lines[1].startswith("> "), "llms.txt must follow the H1 with a summary blockquote" - - -def test_llms_txt_links_every_nav_page(build_llms_txt, pages): - rendered = build_llms_txt.render_llms_txt(pages) - url = build_llms_txt.MARKDOWN_URL - unlinked = [path for _, _, path in pages if f"({url}/{path})" not in rendered] - assert not unlinked, f"llms.txt omits nav pages: {unlinked}" diff --git a/pixi.toml b/pixi.toml index 5d3b1bfd8..99c401c05 100644 --- a/pixi.toml +++ b/pixi.toml @@ -114,12 +114,14 @@ mkdocstrings = "*" mkdocstrings-python = "*" mkdocs-gen-files = "*" mkdocs-literate-nav = "*" +nbsite = ">=0.10.0a0" [feature.doc.tasks] # Updated tasks to use zensical instead of mkdocs _docs-build-site = { cmd = 'zensical build --clean', depends-on = ['install'] } # llms.txt is written into builtdocs, so it has to run after --clean has wiped it. -docs-build = { cmd = 'python scripts/build_llms_txt.py', depends-on = ['_docs-build-site'] } +_docs-markdown = 'python -m nbsite build-llms --config scripts/llms_config.py' +docs-build = { depends-on = ['_docs-build-site', '_docs-markdown'] } docs-serve = { cmd = 'zensical serve', depends-on = ['install'] } docs-clean = { cmd = 'rm -rf builtdocs', depends-on = ['install'] } diff --git a/scripts/build_llms_txt.py b/scripts/build_llms_txt.py deleted file mode 100644 index 6de77d008..000000000 --- a/scripts/build_llms_txt.py +++ /dev/null @@ -1,87 +0,0 @@ -""" -Build the markdown mirror and llms.txt that Lumen serves for LLM consumers. - -Runs as the last step of the docs build: - pixi run -e docs docs-build -""" - -import shutil -import tomllib - -from pathlib import Path - -ROOT = Path(__file__).parent.parent -CONFIG_FILE = ROOT / "zensical.toml" -DOCS_DIR = ROOT / "docs" -MARKDOWN_URL = "/markdown" - -SUMMARY = ( - "Lumen is an open-source and extensible agent framework for chatting with data " - "and for retrieval augmented generation. It turns natural language into SQL, data " - "transformation pipelines, visualizations and dashboards, and every step stays " - "inspectable, editable and reproducible." -) - -PREAMBLE = [ - "Lumen is built on Panel and the wider HoloViz stack. Its declarative data model means " - "anything the language model produces can be serialized to YAML, reopened in a notebook, " - "or composed into a dashboard.", - "", - "Install with `pip install 'lumen[ai-openai]'` and start with `lumen-ai serve data.csv`.", -] - -# Nav entries that sit outside any group, e.g. Quick Start and Installation. -UNGROUPED_SECTION = "Overview" - - -def iter_nav_pages(nav: list, trail: tuple[str, ...] = ()) -> list[tuple[str, str, str]]: - """ - Flatten zensical's nav into (section, label, path) triples. - - The nav is the single source of truth here: it decides which pages reach - llms.txt and what they are called, so no label or grouping is inferred - from file paths. - """ - pages = [] - for entry in nav: - for label, target in entry.items(): - if isinstance(target, list): - pages.extend(iter_nav_pages(target, trail + (label,))) - else: - pages.append((" / ".join(trail) or UNGROUPED_SECTION, label, target)) - return pages - - -def copy_markdown(pages: list[tuple[str, str, str]], output_dir: Path) -> None: - """Mirror every listed page into the built site so the links resolve.""" - for _, _, path in pages: - destination = output_dir / path - destination.parent.mkdir(parents=True, exist_ok=True) - shutil.copy2(DOCS_DIR / path, destination) - - -def render_llms_txt(pages: list[tuple[str, str, str]]) -> str: - lines = ["# Lumen", "", f"> {SUMMARY}", "", *PREAMBLE, ""] - for section in dict.fromkeys(section for section, _, _ in pages): - lines.extend([f"## {section}", ""]) - lines.extend( - f"- [{label}]({MARKDOWN_URL}/{path})" - for page_section, label, path in pages - if page_section == section - ) - lines.append("") - return "\n".join(lines) - - -def main() -> None: - config = tomllib.loads(CONFIG_FILE.read_text(encoding="utf-8"))["project"] - site_dir = ROOT / config["site_dir"] - pages = iter_nav_pages(config["nav"]) - - copy_markdown(pages, site_dir / MARKDOWN_URL.strip("/")) - (site_dir / "llms.txt").write_text(render_llms_txt(pages), encoding="utf-8") - print(f"Wrote llms.txt and {len(pages)} markdown pages to {site_dir}") - - -if __name__ == "__main__": - main() diff --git a/scripts/llms_config.py b/scripts/llms_config.py new file mode 100644 index 000000000..1ed3e1370 --- /dev/null +++ b/scripts/llms_config.py @@ -0,0 +1,118 @@ +"""Config for building Lumen markdown docs and llms.txt. + +The zensical nav is the single source of truth for which pages ship and what +they are called: this reads it once and reuses it to build the llms.txt +sections, instead of maintaining a second, parallel page list. +""" + +import tomllib + +from pathlib import Path + +from nbsite.scripts import LlmsBuildConfig, LlmsSection, MarkdownSource + +ROOT = Path(__file__).parent.parent +DOCS_DIR = ROOT / "docs" +BUILTDOCS_DIR = ROOT / "builtdocs" +OUTPUT_DIR = BUILTDOCS_DIR / "markdown" + + +def _flatten_nav(nav: list, trail: tuple[str, ...] = ()) -> list[tuple[tuple[str, ...], str, Path]]: + """Flatten zensical's nav into (group trail, label, path) triples.""" + pages = [] + for entry in nav: + for label, target in entry.items(): + if isinstance(target, list): + pages.extend(_flatten_nav(target, trail + (label,))) + else: + pages.append((trail, label, Path(target))) + return pages + + +_config = tomllib.loads((ROOT / "zensical.toml").read_text(encoding="utf-8"))["project"] +NAV_PAGES = _flatten_nav(_config["nav"]) +LABELS = {path: label for _, label, path in NAV_PAGES} +TRAILS = {path: trail for trail, _, path in NAV_PAGES} + + +def _label(path: Path) -> str: + return LABELS.get(path, path.stem.replace("_", " ")) + + +def _at(*trail: str): + """Pages whose nav group is exactly *trail*, e.g. Examples/Gallery.""" + return lambda path: TRAILS.get(path) == trail + + +def _under(*trail: str): + """Pages whose nav group starts with *trail*, e.g. everything under Reference.""" + return lambda path: TRAILS.get(path, ())[: len(trail)] == trail + + +CONFIG = LlmsBuildConfig( + project_title="Lumen", + project_description=( + "Lumen is an open-source and extensible agent framework for chatting with data " + "and for retrieval augmented generation. It turns natural language into SQL, data " + "transformation pipelines, visualizations and dashboards, and every step stays " + "inspectable, editable and reproducible." + ), + markdown_root=OUTPUT_DIR, + llms_output_path=BUILTDOCS_DIR / "llms.txt", + markdown_base_url="/markdown", + sources=(MarkdownSource(source_dir=DOCS_DIR, output_dir=OUTPUT_DIR),), + sections=( + LlmsSection( + title="Overview", + description="Quick start, installation, and top-level pages.", + path_prefix=Path("."), + path_filter=_at(), + label_builder=_label, + ), + LlmsSection( + title="Getting Started", + description="Launching Lumen, navigating the UI, and building your first app.", + path_prefix=Path("."), + path_filter=_at("Getting Started"), + label_builder=_label, + ), + LlmsSection( + title="Tutorials", + description="Full walkthroughs building an AI-driven data exploration app end to end.", + path_prefix=Path("."), + path_filter=_at("Examples", "Tutorials"), + label_builder=_label, + group="Examples", + ), + LlmsSection( + title="Gallery", + description="Short example specs demonstrating individual sources, transforms, and views.", + path_prefix=Path("."), + path_filter=_at("Examples", "Gallery"), + label_builder=_label, + group="Examples", + ), + LlmsSection( + title="Configuration", + description="Top-level spec reference: sources, transforms, views, agents, and more.", + path_prefix=Path("."), + path_filter=_at("Configuration"), + label_builder=_label, + ), + LlmsSection( + title="YAML Spec", + description="Detailed reference for writing a Lumen dashboard spec by hand.", + path_prefix=Path("."), + path_filter=_at("Configuration", "Specs"), + label_builder=_label, + group="Configuration", + ), + LlmsSection( + title="API Reference", + description="Python API reference for Lumen's pipeline, sources, transforms, views, and AI components.", + path_prefix=Path("."), + path_filter=_under("Reference"), + label_builder=_label, + ), + ), +) From a9e68e05467f168752658db7b52acc37d49b4950 Mon Sep 17 00:00:00 2001 From: Aman Kumar Date: Mon, 24 Aug 2026 22:55:28 +0530 Subject: [PATCH 5/8] fix: stop nesting YAML Spec under a duplicate Configuration heading --- scripts/llms_config.py | 1 - 1 file changed, 1 deletion(-) diff --git a/scripts/llms_config.py b/scripts/llms_config.py index 1ed3e1370..092b2a46f 100644 --- a/scripts/llms_config.py +++ b/scripts/llms_config.py @@ -105,7 +105,6 @@ def _under(*trail: str): path_prefix=Path("."), path_filter=_at("Configuration", "Specs"), label_builder=_label, - group="Configuration", ), LlmsSection( title="API Reference", From 6393faa71359f0b0b6783e62670f977ee1c68646 Mon Sep 17 00:00:00 2001 From: Aman Kumar Date: Tue, 25 Aug 2026 02:33:12 +0530 Subject: [PATCH 6/8] refactor: collapse repeated LlmsSection boilerplate into a _section() builder --- scripts/llms_config.py | 77 ++++++++++-------------------------------- 1 file changed, 18 insertions(+), 59 deletions(-) diff --git a/scripts/llms_config.py b/scripts/llms_config.py index 092b2a46f..91472b257 100644 --- a/scripts/llms_config.py +++ b/scripts/llms_config.py @@ -39,14 +39,17 @@ def _label(path: Path) -> str: return LABELS.get(path, path.stem.replace("_", " ")) -def _at(*trail: str): - """Pages whose nav group is exactly *trail*, e.g. Examples/Gallery.""" - return lambda path: TRAILS.get(path) == trail - - -def _under(*trail: str): - """Pages whose nav group starts with *trail*, e.g. everything under Reference.""" - return lambda path: TRAILS.get(path, ())[: len(trail)] == trail +def _section(title: str, description: str, *trail: str, group: str | None = None, under: bool = False) -> LlmsSection: + """One LlmsSection for a nav group, matched either exactly (*trail*) or by prefix (*under*).""" + matches_trail = (lambda path: TRAILS.get(path, ())[: len(trail)] == trail) if under else (lambda path: TRAILS.get(path) == trail) + return LlmsSection( + title=title, + description=description, + path_prefix=Path("."), + path_filter=matches_trail, + label_builder=_label, + group=group, + ) CONFIG = LlmsBuildConfig( @@ -62,56 +65,12 @@ def _under(*trail: str): markdown_base_url="/markdown", sources=(MarkdownSource(source_dir=DOCS_DIR, output_dir=OUTPUT_DIR),), sections=( - LlmsSection( - title="Overview", - description="Quick start, installation, and top-level pages.", - path_prefix=Path("."), - path_filter=_at(), - label_builder=_label, - ), - LlmsSection( - title="Getting Started", - description="Launching Lumen, navigating the UI, and building your first app.", - path_prefix=Path("."), - path_filter=_at("Getting Started"), - label_builder=_label, - ), - LlmsSection( - title="Tutorials", - description="Full walkthroughs building an AI-driven data exploration app end to end.", - path_prefix=Path("."), - path_filter=_at("Examples", "Tutorials"), - label_builder=_label, - group="Examples", - ), - LlmsSection( - title="Gallery", - description="Short example specs demonstrating individual sources, transforms, and views.", - path_prefix=Path("."), - path_filter=_at("Examples", "Gallery"), - label_builder=_label, - group="Examples", - ), - LlmsSection( - title="Configuration", - description="Top-level spec reference: sources, transforms, views, agents, and more.", - path_prefix=Path("."), - path_filter=_at("Configuration"), - label_builder=_label, - ), - LlmsSection( - title="YAML Spec", - description="Detailed reference for writing a Lumen dashboard spec by hand.", - path_prefix=Path("."), - path_filter=_at("Configuration", "Specs"), - label_builder=_label, - ), - LlmsSection( - title="API Reference", - description="Python API reference for Lumen's pipeline, sources, transforms, views, and AI components.", - path_prefix=Path("."), - path_filter=_under("Reference"), - label_builder=_label, - ), + _section("Overview", "Quick start, installation, and top-level pages."), + _section("Getting Started", "Launching Lumen, navigating the UI, and building your first app.", "Getting Started"), + _section("Tutorials", "Full walkthroughs building an AI-driven data exploration app end to end.", "Examples", "Tutorials", group="Examples"), + _section("Gallery", "Short example specs demonstrating individual sources, transforms, and views.", "Examples", "Gallery", group="Examples"), + _section("Configuration", "Top-level spec reference: sources, transforms, views, agents, and more.", "Configuration"), + _section("YAML Spec", "Detailed reference for writing a Lumen dashboard spec by hand.", "Configuration", "Specs"), + _section("API Reference", "Python API reference for Lumen's pipeline, sources, transforms, views, and AI components.", "Reference", under=True), ), ) From 547a70baf8c5a18f2ba14b114a23da2f77dca3a2 Mon Sep 17 00:00:00 2001 From: Aman Kumar Date: Tue, 25 Aug 2026 02:36:06 +0530 Subject: [PATCH 7/8] docs: exclude releases and contributing pages from llms.txt --- scripts/llms_config.py | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/scripts/llms_config.py b/scripts/llms_config.py index 91472b257..9929a0f41 100644 --- a/scripts/llms_config.py +++ b/scripts/llms_config.py @@ -16,6 +16,12 @@ BUILTDOCS_DIR = ROOT / "builtdocs" OUTPUT_DIR = BUILTDOCS_DIR / "markdown" +# Pages that carry no LLM code-gen value and should be excluded from the build. +EXCLUDE_FILES = ( + Path("releases.md"), + Path("contributing.md"), +) + def _flatten_nav(nav: list, trail: tuple[str, ...] = ()) -> list[tuple[tuple[str, ...], str, Path]]: """Flatten zensical's nav into (group trail, label, path) triples.""" @@ -63,7 +69,7 @@ def _section(title: str, description: str, *trail: str, group: str | None = None markdown_root=OUTPUT_DIR, llms_output_path=BUILTDOCS_DIR / "llms.txt", markdown_base_url="/markdown", - sources=(MarkdownSource(source_dir=DOCS_DIR, output_dir=OUTPUT_DIR),), + sources=(MarkdownSource(source_dir=DOCS_DIR, output_dir=OUTPUT_DIR, exclude_files=EXCLUDE_FILES),), sections=( _section("Overview", "Quick start, installation, and top-level pages."), _section("Getting Started", "Launching Lumen, navigating the UI, and building your first app.", "Getting Started"), From 8f337a3ce1e1ff47d61b691c5e7a82fd04c3bd97 Mon Sep 17 00:00:00 2001 From: ghostiee-11 Date: Tue, 8 Sep 2026 19:32:00 +0530 Subject: [PATCH 8/8] docs: curate llms index for development --- scripts/llms_config.py | 107 ++++++++++++++++++++--------------------- 1 file changed, 52 insertions(+), 55 deletions(-) diff --git a/scripts/llms_config.py b/scripts/llms_config.py index 9929a0f41..bac866d39 100644 --- a/scripts/llms_config.py +++ b/scripts/llms_config.py @@ -1,11 +1,4 @@ -"""Config for building Lumen markdown docs and llms.txt. - -The zensical nav is the single source of truth for which pages ship and what -they are called: this reads it once and reuses it to build the llms.txt -sections, instead of maintaining a second, parallel page list. -""" - -import tomllib +"""Config for building Lumen markdown docs and its developer-facing llms.txt.""" from pathlib import Path @@ -16,67 +9,71 @@ BUILTDOCS_DIR = ROOT / "builtdocs" OUTPUT_DIR = BUILTDOCS_DIR / "markdown" -# Pages that carry no LLM code-gen value and should be excluded from the build. -EXCLUDE_FILES = ( - Path("releases.md"), - Path("contributing.md"), -) - - -def _flatten_nav(nav: list, trail: tuple[str, ...] = ()) -> list[tuple[tuple[str, ...], str, Path]]: - """Flatten zensical's nav into (group trail, label, path) triples.""" - pages = [] - for entry in nav: - for label, target in entry.items(): - if isinstance(target, list): - pages.extend(_flatten_nav(target, trail + (label,))) - else: - pages.append((trail, label, Path(target))) - return pages - - -_config = tomllib.loads((ROOT / "zensical.toml").read_text(encoding="utf-8"))["project"] -NAV_PAGES = _flatten_nav(_config["nav"]) -LABELS = {path: label for _, label, path in NAV_PAGES} -TRAILS = {path: trail for trail, _, path in NAV_PAGES} - - -def _label(path: Path) -> str: - return LABELS.get(path, path.stem.replace("_", " ")) - - -def _section(title: str, description: str, *trail: str, group: str | None = None, under: bool = False) -> LlmsSection: - """One LlmsSection for a nav group, matched either exactly (*trail*) or by prefix (*under*).""" - matches_trail = (lambda path: TRAILS.get(path, ())[: len(trail)] == trail) if under else (lambda path: TRAILS.get(path) == trail) +DEVELOPMENT_PAGES = { + Path("contributing.md"): "Contributing", + Path("extending.md"): "Extending Lumen", +} +ARCHITECTURE_PAGES = { + Path("configuration/context.md"): "Context", + Path("configuration/agents.md"): "Agents", + Path("configuration/coordinators.md"): "Coordinators", + Path("configuration/tools.md"): "Tools", + Path("configuration/spec/customization.md"): "Custom Components", +} +API_PAGES = { + Path("reference/api.md"): "Overview", + Path("reference/api/ai.md"): "AI", + Path("reference/api/ai/agents.md"): "AI Agents", + Path("reference/api/ai/coordinator.md"): "AI Coordinator", + Path("reference/api/ai/core.md"): "AI Core", + Path("reference/api/ai/models.md"): "AI Models", + Path("reference/api/ai/tools.md"): "AI Tools", + Path("reference/api/pipeline.md"): "Pipeline", + Path("reference/api/sources.md"): "Sources", + Path("reference/api/transforms.md"): "Transforms", + Path("reference/api/views.md"): "Views", +} + + +def _section(title: str, description: str, pages: dict[Path, str]) -> LlmsSection: return LlmsSection( title=title, description=description, path_prefix=Path("."), - path_filter=matches_trail, - label_builder=_label, - group=group, + path_filter=pages.__contains__, + label_builder=pages.__getitem__, ) CONFIG = LlmsBuildConfig( project_title="Lumen", project_description=( - "Lumen is an open-source and extensible agent framework for chatting with data " - "and for retrieval augmented generation. It turns natural language into SQL, data " - "transformation pipelines, visualizations and dashboards, and every step stays " - "inspectable, editable and reproducible." + "Developer documentation for contributing to and extending Lumen, an extensible " + "framework for building data applications and AI-powered data workflows." ), markdown_root=OUTPUT_DIR, llms_output_path=BUILTDOCS_DIR / "llms.txt", markdown_base_url="/markdown", - sources=(MarkdownSource(source_dir=DOCS_DIR, output_dir=OUTPUT_DIR, exclude_files=EXCLUDE_FILES),), + sources=(MarkdownSource( + source_dir=DOCS_DIR, + output_dir=OUTPUT_DIR, + exclude_files=(Path("releases.md"),), + ),), sections=( - _section("Overview", "Quick start, installation, and top-level pages."), - _section("Getting Started", "Launching Lumen, navigating the UI, and building your first app.", "Getting Started"), - _section("Tutorials", "Full walkthroughs building an AI-driven data exploration app end to end.", "Examples", "Tutorials", group="Examples"), - _section("Gallery", "Short example specs demonstrating individual sources, transforms, and views.", "Examples", "Gallery", group="Examples"), - _section("Configuration", "Top-level spec reference: sources, transforms, views, agents, and more.", "Configuration"), - _section("YAML Spec", "Detailed reference for writing a Lumen dashboard spec by hand.", "Configuration", "Specs"), - _section("API Reference", "Python API reference for Lumen's pipeline, sources, transforms, views, and AI components.", "Reference", under=True), + _section( + "Development", + "Repository setup, contribution workflow, and extension points.", + DEVELOPMENT_PAGES, + ), + _section( + "Architecture", + "Core concepts for Lumen's agent and component architecture.", + ARCHITECTURE_PAGES, + ), + _section( + "API Reference", + "Python APIs for Lumen pipelines, sources, transforms, views, and AI components.", + API_PAGES, + ), ), )