Skip to content

About

An AI-powered application that generates multiple creative variations of an uploaded image. It uses a lightweight VAE model to modify images in latent space, with a LangChain agent and MCP-like tool server managing the generation process. A Streamlit interface allows users to upload images, preview variations, and download the generated results.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

AI Image Variator

Generate 4–8 creative image variations using a lightweight VAE and an MCP-like tool server, orchestrated by a LangChain agent with a Streamlit UI.

Overview

  • Upload an image in the Streamlit app
  • LangChain agent calls a local MCP-like JSON-RPC tool server
  • Server encodes the image to a VAE latent, perturbs it (noise + random interpolation), decodes variations
  • Variations are saved to mcp_image_variator/output/ and displayed in the UI

Architecture

ai-image-variator/
├── mcp_image_variator/
│   ├── server.py           # MCP-like JSON-RPC server (stdin/stdout) exposing generate_variations
│   ├── vae_model.py        # Tiny Conv-VAE model with optional torchvision vqvae fallback
│   ├── utils.py            # Image I/O, tensor utils
│   ├── output/             # Generated variations written here
│   └── __init__.py
├── mcp_llm/
│   ├── server.py           # MCP-like JSON-RPC server exposing chat (OpenAI)
│   └── __init__.py
├── agent/
│   ├── agent.py            # LangChain-based workflow to call the MCP tool and summarize results
│   ├── tools.py            # MCP clients: image variations + LLM chat
│   └── __init__.py
├── ui/
│   ├── app.py              # Streamlit app for upload, triggering, grid display, and zip download
│   └── __init__.py
├── run_server.sh           # Convenience script to run server
├── run_ui.sh               # Convenience script to run UI
├── requirements.txt
└── README.md

Notes on "MCP"

This repo implements a minimal JSON-RPC over stdin/stdout interface similar in spirit to MCP tool servers. The agent package includes a tiny client that spawns the server, sends a JSON-RPC request for generate_variations, and parses the response. This keeps the demo self-contained and runnable without extra MCP dependencies.

There are two MCP-like servers:

  • mcp_image_variator/server.py → method: generate_variations(image_path, num_variations)
  • mcp_llm/server.py → method: chat(messages, model, temperature, max_tokens) (uses OpenAI via OPENAI_API_KEY)

Requirements

  • Python 3.9+
  • Windows supported
  • See requirements.txt

Installation

python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt

If you prefer Conda:

conda create -n ai-image-variator python=3.10 -y
conda activate ai-image-variator
pip install -r requirements.txt

Running

  • Start the UI:
streamlit run ui/app.py
  • Optionally, run the server standalone (the agent will spawn it on demand):
python mcp_image_variator/server.py

Or via helper scripts (on Windows with Git Bash or WSL):

bash run_ui.sh
bash run_server.sh

Usage

  1. Open the Streamlit app
  2. Upload a JPG/PNG
  3. Click "Generate Variations"
  4. Preview the 2x2 or 3x3 grid of outputs
  5. Use "Download All" to save a ZIP

Generated images are written to mcp_image_variator/output/.

Optional: LLM summary via OpenAI

The UI sidebar has a toggle: "Use LLM summary (OpenAI)". When enabled, the agent calls the LLM MCP server to generate a richer textual summary of the variations.

Setup:

  1. Create ./.streamlit/secrets.toml and set:
OPENAI_API_KEY = "sk-..."
  1. Install deps (already included):
pip install -r requirements.txt
  1. Start the UI and enable the toggle. If the key is missing, the app falls back to the heuristic summary.

Model

A compact convolutional VAE is included for demo purposes. It will attempt to use torchvision VQ-VAE if available; otherwise, it falls back to a tiny VAE implemented locally. Latent-space variation is produced by:

  • Adding Gaussian noise to the latent vector
  • Randomly interpolating with another latent sample

You can swap in a small pretrained VAE by editing vae_model.py.

Troubleshooting

  • If Torch/Torchvision install fails on Windows, consult the official install matrix for compatible versions with your Python and CUDA/CPU.
  • If images do not appear, check console logs for errors and ensure the mcp_image_variator/output/ directory is writeable.
  • Gray outputs from the VAE? The server includes a fallback that applies strong augmentations (color/contrast/brightness/sharpness/blur/rotation/shift) to guarantee visible differences.
  • LLM not responding? Verify OPENAI_API_KEY in .streamlit/secrets.toml, restart Streamlit, and check network access.

License

MIT

About

An AI-powered application that generates multiple creative variations of an uploaded image. It uses a lightweight VAE model to modify images in latent space, with a LangChain agent and MCP-like tool server managing the generation process. A Streamlit interface allows users to upload images, preview variations, and download the generated results.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages