Duplicate check
Searched open/closed issues (May 2026):
| Item |
Relation |
| #22 (open) |
Not duplicate — commit-API path traversal on product MD creation |
| #10 (open) |
Not duplicate — manifest/source prefix boundary checks |
| #24 (open) |
Not duplicate — orphan sweep deletes source files |
| No open issue |
Batch poll worker output path mismatch |
Summary
The Anthropic batch extraction flow stores custom_id as the bare product slug (e.g. r770), but the poll worker reconstructs the output directory as path.resolve(PRODUCT_MCP_DATA_DIR, slug). For nested products this resolves to data/sample/r770/ instead of data/sample/server/dell/poweredge/r770/. Sync extraction via scripts/extract-one.ts correctly uses ctx.productDir.
Batch extraction via the UI/API therefore appears to succeed in the DB while extraction.json is written to the wrong path or fails with ENOENT.
Steps to reproduce
-
Start app + worker:
npm run dev:all
# or: npm run dev + npm run worker
-
Submit a batch for the sample product:
curl -sS -X POST http://localhost:3210/api/pipeline/extract \
-H 'Content-Type: application/json' \
--data '{"category":"server","products":["r770"]}'
-
Wait for the worker to poll until the batch completes.
-
Compare paths:
# Wrong path (worker target)
ls -la data/sample/r770/extraction.json 2>&1
# Correct path (where viewer reads)
ls -la data/sample/server/dell/poweredge/r770/extraction.json
-
Contrast with sync mode:
npm run extract-one -- --product server/dell/poweredge/r770
# writes to ctx.productDir — correct nested path
Expected behavior
When a batch result arrives, extraction.json is written under the same product directory used to build the extraction prompt (ctx.productDir), e.g. server/dell/poweredge/r770/extraction.json.
Actual behavior
Worker resolves output to {PRODUCT_MCP_DATA_DIR}/{slug}/:
submitBatchForCategory() sets custom_id: ctx.slug (last path segment only)
handleAnthropicBatchPoll() does path.resolve(PRODUCT_MCP_DATA_DIR, productSlug)
Root cause
requests.push({ custom_id: ctx.slug, params: built.params });
const productDir = path.resolve(
env.PRODUCT_MCP_DATA_DIR,
productSlug
);
const outPath = writeExtractionJson(productDir, parsed);
ProductContext has productDir at submit time but it is not encoded in the batch custom_id or persisted on extraction_results for the worker to use.
Suggested fix
- Use full relative path as
custom_id, e.g. path.relative(DATA_DIR, ctx.productDir).
- Or store
output_path / product_path_rel on extraction_results at insert time and read it in the worker instead of reconstructing from slug.
- Add unit test:
custom_id → resolved write path stays under nested product directory.
- Call
invalidateExtractionCache() after write (see issue 02).
PR scope
~40–80 lines across extract.ts, anthropic-batch.ts, and one worker test.
Duplicate check
Searched open/closed issues (May 2026):
Summary
The Anthropic batch extraction flow stores
custom_idas the bare product slug (e.g.r770), but the poll worker reconstructs the output directory aspath.resolve(PRODUCT_MCP_DATA_DIR, slug). For nested products this resolves todata/sample/r770/instead ofdata/sample/server/dell/poweredge/r770/. Sync extraction viascripts/extract-one.tscorrectly usesctx.productDir.Batch extraction via the UI/API therefore appears to succeed in the DB while
extraction.jsonis written to the wrong path or fails with ENOENT.Steps to reproduce
Start app + worker:
npm run dev:all # or: npm run dev + npm run workerSubmit a batch for the sample product:
Wait for the worker to poll until the batch completes.
Compare paths:
Contrast with sync mode:
npm run extract-one -- --product server/dell/poweredge/r770 # writes to ctx.productDir — correct nested pathExpected behavior
When a batch result arrives,
extraction.jsonis written under the same product directory used to build the extraction prompt (ctx.productDir), e.g.server/dell/poweredge/r770/extraction.json.Actual behavior
Worker resolves output to
{PRODUCT_MCP_DATA_DIR}/{slug}/:submitBatchForCategory()setscustom_id: ctx.slug(last path segment only)handleAnthropicBatchPoll()doespath.resolve(PRODUCT_MCP_DATA_DIR, productSlug)Root cause
ProductContexthasproductDirat submit time but it is not encoded in the batchcustom_idor persisted onextraction_resultsfor the worker to use.Suggested fix
custom_id, e.g.path.relative(DATA_DIR, ctx.productDir).output_path/product_path_relonextraction_resultsat insert time and read it in the worker instead of reconstructing from slug.custom_id→ resolved write path stays under nested product directory.invalidateExtractionCache()after write (see issue 02).PR scope
~40–80 lines across
extract.ts,anthropic-batch.ts, and one worker test.