Generate 4–8 creative image variations using a lightweight VAE and an MCP-like tool server, orchestrated by a LangChain agent with a Streamlit UI.
- Upload an image in the Streamlit app
- LangChain agent calls a local MCP-like JSON-RPC tool server
- Server encodes the image to a VAE latent, perturbs it (noise + random interpolation), decodes variations
- Variations are saved to
mcp_image_variator/output/and displayed in the UI
ai-image-variator/
├── mcp_image_variator/
│ ├── server.py # MCP-like JSON-RPC server (stdin/stdout) exposing generate_variations
│ ├── vae_model.py # Tiny Conv-VAE model with optional torchvision vqvae fallback
│ ├── utils.py # Image I/O, tensor utils
│ ├── output/ # Generated variations written here
│ └── __init__.py
├── mcp_llm/
│ ├── server.py # MCP-like JSON-RPC server exposing chat (OpenAI)
│ └── __init__.py
├── agent/
│ ├── agent.py # LangChain-based workflow to call the MCP tool and summarize results
│ ├── tools.py # MCP clients: image variations + LLM chat
│ └── __init__.py
├── ui/
│ ├── app.py # Streamlit app for upload, triggering, grid display, and zip download
│ └── __init__.py
├── run_server.sh # Convenience script to run server
├── run_ui.sh # Convenience script to run UI
├── requirements.txt
└── README.md
This repo implements a minimal JSON-RPC over stdin/stdout interface similar in spirit to MCP tool servers. The agent package includes a tiny client that spawns the server, sends a JSON-RPC request for generate_variations, and parses the response. This keeps the demo self-contained and runnable without extra MCP dependencies.
There are two MCP-like servers:
mcp_image_variator/server.py→ method:generate_variations(image_path, num_variations)mcp_llm/server.py→ method:chat(messages, model, temperature, max_tokens)(uses OpenAI viaOPENAI_API_KEY)
- Python 3.9+
- Windows supported
- See
requirements.txt
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txtIf you prefer Conda:
conda create -n ai-image-variator python=3.10 -y
conda activate ai-image-variator
pip install -r requirements.txt- Start the UI:
streamlit run ui/app.py- Optionally, run the server standalone (the agent will spawn it on demand):
python mcp_image_variator/server.pyOr via helper scripts (on Windows with Git Bash or WSL):
bash run_ui.sh
bash run_server.sh- Open the Streamlit app
- Upload a JPG/PNG
- Click "Generate Variations"
- Preview the 2x2 or 3x3 grid of outputs
- Use "Download All" to save a ZIP
Generated images are written to mcp_image_variator/output/.
The UI sidebar has a toggle: "Use LLM summary (OpenAI)". When enabled, the agent calls the LLM MCP server to generate a richer textual summary of the variations.
Setup:
- Create
./.streamlit/secrets.tomland set:
OPENAI_API_KEY = "sk-..."
- Install deps (already included):
pip install -r requirements.txt
- Start the UI and enable the toggle. If the key is missing, the app falls back to the heuristic summary.
A compact convolutional VAE is included for demo purposes. It will attempt to use torchvision VQ-VAE if available; otherwise, it falls back to a tiny VAE implemented locally. Latent-space variation is produced by:
- Adding Gaussian noise to the latent vector
- Randomly interpolating with another latent sample
You can swap in a small pretrained VAE by editing vae_model.py.
- If Torch/Torchvision install fails on Windows, consult the official install matrix for compatible versions with your Python and CUDA/CPU.
- If images do not appear, check console logs for errors and ensure the
mcp_image_variator/output/directory is writeable. - Gray outputs from the VAE? The server includes a fallback that applies strong augmentations (color/contrast/brightness/sharpness/blur/rotation/shift) to guarantee visible differences.
- LLM not responding? Verify
OPENAI_API_KEYin.streamlit/secrets.toml, restart Streamlit, and check network access.
MIT