Generated by Co-Pilot:
Proposal: Ingest repository "non-commit" data (issues, PRs, PR reviews, discussions, wikis, project boards) into our knowledge and search infrastructure to improve developer productivity, triage, provenance, and RAG-based QA.
Rationale / Benefits
- Developer context & rationale: Issues, PR discussions, and wikis capture design decisions, trade-offs, and rationale not present in commit history.
- Better RAG answers: Retrieval-augmented systems produce higher-quality answers when they can cite issue/PR threads and changelogs.
- Faster triage and automation: Structured metadata (labels, linked PRs, status) enables triage automation and prioritization rules.
- Release notes & changelogs: PRs + linked issues make automated changelog generation more reliable.
- Historical provenance: Linking issues↔PRs↔commits explains when/why behavior changed.
Cost–Benefit Summary
- Low-cost, high-value first tier: ingest structured metadata (titles, labels, authors, timestamps, links) with minimal LLM use.
- Moderate-cost, high-value second tier: selective LLM summarization and fact extraction for high-signal threads (e.g., labelled "bug", high-activity PRs).
- High-cost, low-value third tier: full-thread LLM parsing of everything (expensive, noisy, and likely to produce non-deterministic extractions).
Recommended Approach (staged, cost-sensitive)
- Metadata-first: index all issues/PRs metadata (title, labels, state, linked numbers, authors, timestamps) and surface it in search. No LLM required.
- Heuristics filtering: filter threads by labels, recency, comment count, or link to commits. Only these go to stage 3.
- Two-stage LLM processing (summarize → extract): use a small/cheaper model to produce short thread summaries, then run structured extraction on summaries for selected threads only.
- Store provenance and confidence scores; keep raw text for audits and reprocessing.
- Human-in-the-loop verification on a sample to tune prompts and measure precision.
Pilot proposal (recommended)
- Scope: last 12 months of issues + merged PRs + their reviews.
- Ingestion: metadata-only ingestion for entire scope; select top 200 high-activity or high-priority threads for two-stage processing.
- Metrics: extraction precision@100, coverage of priority issues, time-to-triage improvement, cost per processed thread.
- Acceptance criteria: >80% precision for extracted summaries on sample; measurable improvement in common RAG queries (benchmarked).
Implementation notes & tech suggestions
- Storage: object store for raw text (S3); vector DB for embeddings (PGVector, Pinecone, Milvus).
- Embeddings: batch by thread; reuse embeddings for later reprocessing.
- Models: start with a cheap summarizer; use a stronger model selectively for extraction if needed.
- Orchestration: small ETL jobs and webhook-based incremental updates.
- Governance: redact PII and secrets; store provenance links; respect licensing.
Risks & mitigations
- Noise & hallucination: mitigate with human sampling, confidence scoring, and storing raw text plus provenance.
- Cost overruns: avoid full-history LLM parsing; use heuristics & sampling; monitor costs.
- Legal/privacy: implement PII detection & redaction.
Next steps
- Run a lightweight inventory: counts for issues, PRs, discussions, wiki pages (last 12 months + all-time).
- Run the pilot as scoped above and report on cost and precision.
Generated by Co-Pilot:
Proposal: Ingest repository "non-commit" data (issues, PRs, PR reviews, discussions, wikis, project boards) into our knowledge and search infrastructure to improve developer productivity, triage, provenance, and RAG-based QA.
Rationale / Benefits
Cost–Benefit Summary
Recommended Approach (staged, cost-sensitive)
Pilot proposal (recommended)
Implementation notes & tech suggestions
Risks & mitigations
Next steps