A reproducible AI agent workflow for research-style queries across X-style discourse and research-paper datasets.
This project is designed and packaged like an AI backend / agent systems project, not just a chat demo. It focuses on:
- multi-step planning, execution, analysis, and synthesis
- tool orchestration across heterogeneous local datasets
- deterministic offline demos without live model credentials
- graceful degradation when semantic retrieval backends are unavailable
- local benchmark and regression coverage for repeatable iteration
π GitHub Repository: https://github.com/Kylinny/agentic-research-workflow
Many agent demos look impressive but are hard to reproduce locally because they depend on API keys, unstable model behavior, or brittle retrieval pipelines. This repository turns that problem into an engineering exercise: build an agent workflow that still runs, benchmarks, and demos cleanly even when parts of the runtime are unavailable.
Open the full sample here: examples/sample_output.md
# 1. Create a Python 3.10+ environment
python3.10 -m venv .venv310
source .venv310/bin/activate
pip install -r requirements.txt
# 2. Run the offline workflow
python main.py --offline --query "How does X discourse on biotechnology compare to academic research?"
# 3. Run the offline regression suite
python -m unittest discover -s tests
# 4. Run a small offline benchmark
python evaluation/run_benchmark.py --offline --max-queries 2The workflow supports:
- Planning: break complex research questions into executable tasks
- Execution: run retrieval and analysis tools with dependency-aware orchestration
- Analysis: score intermediate results for quality and completeness
- Synthesis: generate a final structured answer from tool outputs
- Offline reproducibility: run demos and benchmarks without a live xAI key
- Retrieval fallback: degrade from semantic retrieval to keyword-only mode when needed
- Good for demos:
--offlinemode makes the workflow easy to show without external setup - Good for iteration: regression tests and benchmark scripts catch breakage quickly
- Good for storytelling: the architecture highlights agent control flow, retrieval design, and runtime resilience
- Iterative loop: plan β decompose β select tools β analyze β refine β summarize
- Context-aware decision making with multi-step memory
- Adaptive replanning when encountering ambiguous or insufficient results
- Tool chaining for complex multi-hop reasoning
- X (Twitter) Data: Real-time style posts with replies, threads, timestamps, and engagement metrics
- Research Papers: Structured documents with abstracts, methodologies, results, and citations
- Realistic noise, sarcasm detection challenges, and multilingual content
- Semantic search using sentence transformers
- Keyword-based retrieval for precise matching
- FAISS-powered vector indexing for scalability
- Cross-document analysis and citation tracking
- Multiple model variants (grok-beta, grok-2-latest, etc.)
- Optimized prompting for planning and reasoning
- Error handling and retry logic
- Token-aware context management
- 20-40 complex research queries
- Metrics: completion rate, step efficiency, answer quality
- Comparative analysis across Grok model variants
- Automated benchmarking
# Clone the repository
git clone https://github.com/Kylinny/agentic-research-workflow.git
cd agentic-research-workflow
# Set up your API key
cp .env.example .env
# Edit .env and add your X.AI API key
# Build and run
docker-compose up --build# Clone the repository
git clone https://github.com/Kylinny/agentic-research-workflow.git
cd agentic-research-workflow
# Create virtual environment
python3 -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
# Set up environment variables
cp .env.example .env
# Edit .env and add your X.AI API key
# Generate datasets
python scripts/generate_datasets.py
# Run the agent
python main.py --query "What are the latest trends in AI safety research?"# Interactive mode
python main.py --interactive
# Single query
python main.py --query "Analyze the sentiment around recent AI regulation discussions"
# Batch evaluation
python evaluation/run_benchmark.py
# Compare models
python evaluation/compare_models.py# Run the full workflow without an xAI API key
python main.py --offline --query "Compare public discourse and research trends in AI safety"
# Run a small offline benchmark
python evaluation/run_benchmark.py --offline --max-queries 3Offline mode uses a deterministic heuristic client so you can demo planning, tool execution, and synthesis locally before wiring in a live model.
python -m unittest discover -s tests# X Data Analysis
python main.py --query "What are users saying about climate change? Identify key influencers and sentiment trends over time."
# Research Paper Analysis
python main.py --query "Compare methodologies used in transformer architecture papers from 2017-2023. What are the key innovations?"
# Cross-domain
python main.py --query "How does public discourse on X about quantum computing align with academic research?"flowchart TD
A["User Query"] --> B["ResearchAgent Core"]
B --> C["Planner"]
B --> D["Executor"]
B --> E["Analyzer"]
B --> F["Context Manager"]
D --> G["Tool Layer"]
G --> H["X Search"]
G --> I["Paper Search"]
G --> J["Hybrid Retrieval"]
G --> K["Sentiment Analysis"]
G --> L["Citation Tracker"]
B --> M["Live Grok Client or Offline Client"]
J --> N["Semantic Retrieval if enabled"]
J --> O["Keyword-only fallback"]
Xai/
βββ agent/
β βββ __init__.py
β βββ core.py # Main agent loop
β βββ planner.py # Task decomposition & planning
β βββ executor.py # Tool execution
β βββ analyzer.py # Result analysis & refinement
β βββ context_manager.py # Context & memory management
βββ grok/
β βββ __init__.py
β βββ client.py # Grok API client
β βββ prompts.py # Optimized prompts
β βββ models.py # Model configurations
βββ tools/
β βββ __init__.py
β βββ x_search.py # X data retrieval
β βββ paper_search.py # Research paper search
β βββ sentiment.py # Sentiment analysis
β βββ citations.py # Citation tracking
β βββ hybrid_retrieval.py # Hybrid search engine
βββ data/
β βββ x_posts.json # Generated X data
β βββ research_papers.json # Generated papers
β βββ embeddings/ # Vector indices
βββ evaluation/
β βββ queries.json # Test queries
β βββ run_benchmark.py # Evaluation script
β βββ compare_models.py # Model comparison
β βββ metrics.py # Evaluation metrics
βββ scripts/
β βββ generate_datasets.py # Dataset generation
βββ Dockerfile
βββ docker-compose.yml
βββ main.py
βββ requirements.txt
βββ README.md
The system tracks:
- Completion Rate: % of queries successfully answered
- Step Efficiency: Average steps per query
- Answer Quality: Relevance, coherence, citation accuracy
- Replanning Rate: Frequency of adaptive replanning
- Context Utilization: Effective use of conversation history
See TROUBLESHOOTING.md for common issues and solutions.
- Technical Documentation - Architecture and design decisions
- Troubleshooting Guide - Common issues and solutions
MIT License - See LICENSE file for details.