Skip to content

[bug] batch worker writes extraction.json to {DATA_DIR}/{slug}/ instead of nested product path #26

Description

@kiannidev

Duplicate check

Searched open/closed issues (May 2026):

Item Relation
#22 (open) Not duplicate — commit-API path traversal on product MD creation
#10 (open) Not duplicate — manifest/source prefix boundary checks
#24 (open) Not duplicate — orphan sweep deletes source files
No open issue Batch poll worker output path mismatch

Summary

The Anthropic batch extraction flow stores custom_id as the bare product slug (e.g. r770), but the poll worker reconstructs the output directory as path.resolve(PRODUCT_MCP_DATA_DIR, slug). For nested products this resolves to data/sample/r770/ instead of data/sample/server/dell/poweredge/r770/. Sync extraction via scripts/extract-one.ts correctly uses ctx.productDir.

Batch extraction via the UI/API therefore appears to succeed in the DB while extraction.json is written to the wrong path or fails with ENOENT.


Steps to reproduce

  1. Start app + worker:

    npm run dev:all
    # or: npm run dev  +  npm run worker
  2. Submit a batch for the sample product:

    curl -sS -X POST http://localhost:3210/api/pipeline/extract \
      -H 'Content-Type: application/json' \
      --data '{"category":"server","products":["r770"]}'
  3. Wait for the worker to poll until the batch completes.

  4. Compare paths:

    # Wrong path (worker target)
    ls -la data/sample/r770/extraction.json 2>&1
    
    # Correct path (where viewer reads)
    ls -la data/sample/server/dell/poweredge/r770/extraction.json
  5. Contrast with sync mode:

    npm run extract-one -- --product server/dell/poweredge/r770
    # writes to ctx.productDir — correct nested path

Expected behavior

When a batch result arrives, extraction.json is written under the same product directory used to build the extraction prompt (ctx.productDir), e.g. server/dell/poweredge/r770/extraction.json.


Actual behavior

Worker resolves output to {PRODUCT_MCP_DATA_DIR}/{slug}/:

  • submitBatchForCategory() sets custom_id: ctx.slug (last path segment only)
  • handleAnthropicBatchPoll() does path.resolve(PRODUCT_MCP_DATA_DIR, productSlug)

Root cause

    requests.push({ custom_id: ctx.slug, params: built.params });
        const productDir = path.resolve(
          env.PRODUCT_MCP_DATA_DIR,
          productSlug
        );
        const outPath = writeExtractionJson(productDir, parsed);

ProductContext has productDir at submit time but it is not encoded in the batch custom_id or persisted on extraction_results for the worker to use.


Suggested fix

  1. Use full relative path as custom_id, e.g. path.relative(DATA_DIR, ctx.productDir).
  2. Or store output_path / product_path_rel on extraction_results at insert time and read it in the worker instead of reconstructing from slug.
  3. Add unit test: custom_id → resolved write path stays under nested product directory.
  4. Call invalidateExtractionCache() after write (see issue 02).

PR scope

~40–80 lines across extract.ts, anthropic-batch.ts, and one worker test.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions