Skip to content

Latest commit

 

History

History
135 lines (98 loc) · 9.81 KB

File metadata and controls

135 lines (98 loc) · 9.81 KB

Embedded runtime

Data Liberation exposes its existing product operations through a generic Node API:

import {
  inspectSource,
  captureWebsite,
  checkFidelity,
  serveCapture,
  publishSite,
  registerPlatform,
  registerPublishTarget,
} from 'data-liberation/runtime';

data-liberation and data-liberation/runtime resolve to the same module, including the same platform and publish registries. The runtime uses the implementations used by the CLI and MCP, with no separate pipeline, destination policy or sandbox configuration.

Standalone distribution

await serveCapture(directory) serves the owned portable website/ using the same clean-URL and asset resolver as fidelity comparison. The directory may be the capture root or its website/ directory. The returned CapturePreviewServer provides url, port, urlForPage, and close(); the caller must close it in a finally block. This helper makes no live-origin request and requires no browser, so consumers can reuse source rendering without implementing another server.

Shared parts and serving

The canonical website/ may contain parts/header-<sha256>.html and parts/footer-<sha256>.html. Each is an exact byte slice of a repeated site-level semantic landmark. The filename hash is a stable initial identity, not an edit lock: consumers may edit a part in place. Routes reference the part in position:

<!--#include virtual="/parts/header-<sha256>.html" -->

serveCapture resolves real HTML comment nodes on the server for every HTML request, preserving all other bytes. The browser receives a complete document, without client fetching or an assembly script. Offline self-consistency uses the same resolver, so shared IDs, links and remote assets are checked on every route. Capture evidence and geometry source hashes are projected before extraction and describe the expanded document. Parts are resources and do not enlarge the receipt route table. Existing tree/ZIP reporting includes their stored bytes.

The resolver accepts exactly the comment syntax above, with a root-relative /parts/ path, a filename containing ASCII letters, digits, _ or -, and an .html suffix. Include-looking strings in scripts, attributes and raw text are left untouched. Malformed directives, missing files, cycles, traversal and symlinks fail: preview returns HTTP 500 with a resolution error and offline comparison reports include-unresolved. Reads are bounded before allocation: 16 MiB per file, 64 MiB cumulative read/expanded document budget, 16 nesting levels and 1,024 file reads per request. No expanded copy is written to disk.

A raw file server or file:// does not resolve these comments. To serve without DLA, use an include-aware server or compile the same grammar into expanded HTML for deployment. Compilation must walk HTML comment nodes, preserve surrounding bytes, and resolve bounded paths inside the website root. Consumer platform support is maintained separately; the export is destination-neutral.

The committed dist/capture-engine.bundle.mjs now exports the full runtime. Its historical filename is retained for consumers that already pin that artifact. An embedded runner can import the file directly:

import { inspectSource, captureWebsite, checkFidelity } from './capture-engine.bundle.mjs';

Node 22 or later is required. The bundle includes ordinary JavaScript dependencies. Importing it, registering platforms/publish targets, HTTP-only inspection and publishing to an in-process target work without Playwright or the source checkout.

Browser operations require separately provisioned Playwright and Chromium. Provision them in an ancestor node_modules visible to the bundle, and make the installed browser cache available to the runtime user. The installed-package gate copies only playwright and its playwright-core dependency beside a relocated bundle before exercising the browser workflow. single-file-cli remains an optional external dependency of the existing freeze path. No runtime operation installs dependencies.

Use a tested immutable DLA revision and provision browser dependencies during environment construction. A dependency pin baked into an existing environment must be rebuilt to receive a newer bundle.

Release asset

Each GitHub Release attaches the package's npm pack tarball, data-liberation-<version>.tgz. It contains exactly the published package files, including dist/capture-engine.bundle.mjs and the src/ runtime assets its modules resolve relative to themselves. Consumers pin a release by URL and digest instead of cloning a commit:

  • Resolve the newest stable release with GET /repos/Automattic/data-liberation-agent/releases/latest and select the .tgz asset. GitHub reports its SHA-256 in the asset's digest field (sha256:<hex>).
  • Download browser_download_url, verify it against that digest, and extract with tar -xzf <asset> --strip-components=1 so dist/capture-engine.bundle.mjs lands at the root of the target directory.
  • Provision Playwright beside it as described above; the tarball does not include dependencies.

Operations and failure behavior

Operation Inputs Result and failure contract
inspectSource(url, options?) Bounded discovery/rendering options; rendered: false selects HTTP-only SourceInspection, including complexity factors, coverage, unknowns and issues. Missing browser support becomes browser-unavailable and unknown complexity. Invalid input or an unrecoverable entry request rejects.
captureWebsite(options) url, outputDir, optional resume, captureImages, learnFluid, strict, onProgress CaptureResult with receipt path, route counts, complete, and unresolvedAnchors. A partial site still resolves with complete: false unless strict: true, which rejects with IncompleteCaptureError. Callers must inspect complete rather than inferring coverage from counters. Setup errors reject. Export publication rejects while another export holds the run lock; it does not rematerialize over that writer. Pre-commit failures restore the previous public generation; a post-commit cleanup failure retains the new generation and recovery state. Preview and fidelity-reference updates run after that publication. Output follows the existing cwd-local path contract. See export publication.
checkFidelity(options) directory, optional stage, routes, widths, states, screenshots/settling/log callback, optional candidateUrl FidelityReport. Defaults to frozen source → capture at 390/768/1440; with candidateUrl, portable capture → candidate. Missing/stale evidence stays pending/unproven and cannot pass. Explicit stage: 'drift' preserves live-source sampling, observation injection and motion contracts. See reference contract and evidence.
publishSite(options) directory, target, optional credentials/log callback PublishResult from the selected target. Publishing is an explicit operation. Target and setup failures reject; optional attribution runs in disposable staging.

Types are exported for all options/results, including InspectOptions, CaptureOptions, CaptureResult, FidelityCheckOptions, FidelityReport, PublishSiteOptions, and PublishResult.

An application can compose these operations while keeping its acceptance policy explicit:

export async function prepareSource(url, outputDir, acceptSource) {
  const inspection = await inspectSource(url, { sampleLimit: 5 });
  if (!acceptSource(inspection)) throw new Error('Source needs review');

  const capture = await captureWebsite({ url, outputDir });
  if (!capture.complete) throw new Error('Capture is incomplete');

  const comparison = await checkFidelity({ directory: outputDir });
  if (!comparison.pass) throw new Error('Captured site failed fidelity checks');

  return { inspection, capture, comparison };
}

Inspection and comparison are advisory/results APIs, not automatic gates inside capture. Capture can emit existing diagnostic logging; a host that reserves stdout for its own protocol must account for that output. All asynchronous operations should be awaited so browser/context cleanup completes.

Extension points

Import registration and operations from this runtime entry. Custom platforms contribute discovery, inspection signals and cleanup rules through registerPlatform. Destinations contribute publishing and optional attribution through registerPublishTarget. The bundle and package share the exact public contract; callers need no internal src/ imports or MCP transport to use it.

Verification

Tracking: #215

npm ci
npm run setup:browser
npm run build
npm run test:package
npm test -- --maxWorkers=2 --testTimeout=45000

test:package installs the packed package into a separate consumer and verifies runtime/root module identity and TypeScript imports. It then copies just the committed bundle into two standalone directories:

  1. No dependencies: real HTTP inspection, graceful missing-browser inspection/comparison failure, custom platform registration, and in-process publishing with destination attribution.
  2. Only Playwright/core provisioned: rendered inspection of JS-created application content, custom-platform capture with cleanup, comparison of the portable artifact, rejection of deliberately removed owner content, and publishing without modifying the canonical artifact.

The workflow is implemented in scripts/test-runtime.mjs and runs against a local HTTP fixture. It asserts emitted artifacts and outcomes, rather than only export names. It performs no external publish.

AI assistance: OpenAI gpt-6-astra via OpenCode implemented and verified this generic distribution change directly in an isolated worktree under Chris Huber's direction.