Skip to content

RFC: Ingest repository issues/PRs/discussions (metadata-first, cost-sensitive pipeline) #234

Description

@adityamukho

Generated by Co-Pilot:

Proposal: Ingest repository "non-commit" data (issues, PRs, PR reviews, discussions, wikis, project boards) into our knowledge and search infrastructure to improve developer productivity, triage, provenance, and RAG-based QA.

Rationale / Benefits

  • Developer context & rationale: Issues, PR discussions, and wikis capture design decisions, trade-offs, and rationale not present in commit history.
  • Better RAG answers: Retrieval-augmented systems produce higher-quality answers when they can cite issue/PR threads and changelogs.
  • Faster triage and automation: Structured metadata (labels, linked PRs, status) enables triage automation and prioritization rules.
  • Release notes & changelogs: PRs + linked issues make automated changelog generation more reliable.
  • Historical provenance: Linking issues↔PRs↔commits explains when/why behavior changed.

Cost–Benefit Summary

  • Low-cost, high-value first tier: ingest structured metadata (titles, labels, authors, timestamps, links) with minimal LLM use.
  • Moderate-cost, high-value second tier: selective LLM summarization and fact extraction for high-signal threads (e.g., labelled "bug", high-activity PRs).
  • High-cost, low-value third tier: full-thread LLM parsing of everything (expensive, noisy, and likely to produce non-deterministic extractions).

Recommended Approach (staged, cost-sensitive)

  1. Metadata-first: index all issues/PRs metadata (title, labels, state, linked numbers, authors, timestamps) and surface it in search. No LLM required.
  2. Heuristics filtering: filter threads by labels, recency, comment count, or link to commits. Only these go to stage 3.
  3. Two-stage LLM processing (summarize → extract): use a small/cheaper model to produce short thread summaries, then run structured extraction on summaries for selected threads only.
  4. Store provenance and confidence scores; keep raw text for audits and reprocessing.
  5. Human-in-the-loop verification on a sample to tune prompts and measure precision.

Pilot proposal (recommended)

  • Scope: last 12 months of issues + merged PRs + their reviews.
  • Ingestion: metadata-only ingestion for entire scope; select top 200 high-activity or high-priority threads for two-stage processing.
  • Metrics: extraction precision@100, coverage of priority issues, time-to-triage improvement, cost per processed thread.
  • Acceptance criteria: >80% precision for extracted summaries on sample; measurable improvement in common RAG queries (benchmarked).

Implementation notes & tech suggestions

  • Storage: object store for raw text (S3); vector DB for embeddings (PGVector, Pinecone, Milvus).
  • Embeddings: batch by thread; reuse embeddings for later reprocessing.
  • Models: start with a cheap summarizer; use a stronger model selectively for extraction if needed.
  • Orchestration: small ETL jobs and webhook-based incremental updates.
  • Governance: redact PII and secrets; store provenance links; respect licensing.

Risks & mitigations

  • Noise & hallucination: mitigate with human sampling, confidence scoring, and storing raw text plus provenance.
  • Cost overruns: avoid full-history LLM parsing; use heuristics & sampling; monitor costs.
  • Legal/privacy: implement PII detection & redaction.

Next steps

  1. Run a lightweight inventory: counts for issues, PRs, discussions, wiki pages (last 12 months + all-time).
  2. Run the pilot as scoped above and report on cost and precision.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentationenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions