Welcome to the developer documentation for lollms_client! This guide is intended for developers who want to understand the project's architecture, contribute new features, fix bugs, or create new bindings.
Project Links:
- GitHub Repository: https://github.com/ParisNeo/simplified_lollms (Note: The client is part of this larger ecosystem, this doc focuses on the
lollms_clientlibrary itself) - PyPI Package: https://pypi.org/project/lollms-client/
- License: Apache 2.0
- Introduction
- Getting Started for Developers
- Project Structure
- Core Concepts & Architecture
- Adding New Bindings (A Practical Guide)
- Running Examples & Tests
- Coding Standards & Conventions
- Contribution Guidelines
- Reporting Issues
- Roadmap & Future Ideas
- Community & Contact
lollms_client is a Python client library designed to provide a unified and easy-to-use interface for interacting with various AI model backends and services. It aims to abstract the complexities of different APIs and local model execution, allowing developers to seamlessly switch between different AI modalities (LLM, TTS, TTI, STT, TTM, TTV), and enable powerful function calling via the Model Context Protocol (MCP).
The "LoLLMS" ecosystem (Lord of Large Language and Multimodal Systems) aims to be a versatile platform for AI interaction, and this client is a key component for Python-based applications and integrations.
- Modular Binding System: Easily add support for new AI backends or libraries.
- Multi-Modality: Supports Large Language Models (LLM), Text-to-Speech (TTS), Text-to-Image (TTI), Speech-to-Text (STT), Text-to-Music/Sound (TTM), and Text-to-Video (TTV).
- Function Calling (MCP): Integrated Model Context Protocol support, including a
local_mcpbinding for executing local Python tools with default utilities. - Unified API: Provides a consistent
LollmsClientinterface for common operations across different bindings. - Local & Remote Backends: Supports both local model execution and remote services.
- Helper Utilities: Includes tools for discussion management and various AI-driven tasks directly on the
LollmsClient. - Dependency Management: Uses
pipmasterwithin bindings to attempt to ensure necessary Python packages are installed.
(Content remains largely the same: Prerequisites, Cloning, Virtual Env, Editable Install, Dev Dependencies)
Here's an overview of the lollms_client directory structure:
π lollms_client/
ββ π ai_documentation/ # Markdown files generated by AI documentation tools
ββ π dist/ # Build artifacts
ββ π examples/ # Example scripts demonstrating client usage
β ββ π article_summary/
β ββ π deep_analyze/
β ββ π function_calling_with_local_custom_mcp.py # Shows custom MCP tools
β ββ π local_mcp.py # Shows default MCP tools
β ββ π generate_and_speak/
β ββ ... (other examples)
ββ π lollms_client/ # The core library source code
β ββ π llm_bindings/ # LLM-specific binding implementations
β β ββ ...
β ββ π tts_bindings/ # Text-to-Speech binding implementations
β β ββ ...
β ββ π tti_bindings/ # Text-to-Image binding implementations
β β ββ ...
β ββ π stt_bindings/ # Speech-to-Text binding implementations
β β ββ ...
β ββ π ttm_bindings/ # Text-to-Music/Sound binding implementations
β β ββ ...
β ββ π ttv_bindings/ # Text-to-Video binding implementations
β β ββ ...
β ββ π tools_bindings/ # Model Context Protocol binding implementations
β β ββ π local_mcp/ # Example: Local MCP tool executor
β β ββ π default_tools/ # Packaged tools for local_mcp
β β β ββ π file_writer/
β β β ββ π generate_image_from_prompt/
β β β ββ π internet_search/
β β β ββ π python_interpreter/
β β ββ π __init__.py
β ββ π __init__.py # Makes lollms_client a package, exports key classes
β ββ π lollms_config.py # Configuration classes
β ββ π lollms_core.py # LollmsClient class, main orchestrator (includes generate_with_mcp, summarize, etc.)
β ββ π lollms_discussion.py # LollmsDiscussion and LollmsMessage classes
β ββ π lollms_llm_binding.py # ABC for LLM bindings and LollmsLLMBindingManager
β ββ π lollms_tts_binding.py # ABC for TTS bindings and LollmsTTSBindingManager
β ββ π lollms_tti_binding.py # ABC for TTI bindings and LollmsTTIBindingManager
β ββ π lollms_stt_binding.py # ABC for STT bindings and LollmsSTTBindingManager
β ββ π lollms_ttm_binding.py # ABC for TTM bindings and LollmsTTMBindingManager
β ββ π lollms_ttv_binding.py # ABC for TTV bindings and LollmsTTVBindingManager
β ββ π lollms_tools_binding.py # ABC for MCP bindings and LollmsTOOLBindingManager
β ββ π lollms_types.py # Enums (MSG_TYPE, ELF_COMPLETION_FORMAT, etc.)
β ββ π lollms_utilities.py # Helper functions
ββ π lollms_client.egg-info/ # Packaging metadata
ββ π CHANGELOG.md
ββ π DOC_DEV.md # This developer documentation file
ββ π DOC_USE.md # User-focused documentation
ββ π pyproject.toml
ββ π README.md
ββ π requirements.txt
Located in lollms_client/lollms_core.py, the LollmsClient class is the main entry point. Its responsibilities include:
- Initializing and managing different types of bindings (LLM, TTS, TTI, STT, TTM, TTV, MCP).
- Providing a unified API for common operations like text generation (
generate_text), speech synthesis (tts.generate_audio), and function calling (generate_with_mcp). - Orchestrating interactions involving MCP tools by leveraging an active MCP binding.
- Offering direct methods for high-level tasks like
sequential_summarize,deep_analyze,generate_code,yes_no, etc. - Storing default generation parameters.
(General description of Bindings, ABCs, and Binding Managers remains the same)
- Zero Port Drift Singletons: Every local daemon binding (
whisperon 9633,xttson 9634,piper_ttson 9635,barkon 9636,diffusersTTI on 9632,diffusersTTM on 9637) adheres to the "first instance wins" pattern. The first process spawns the daemon with a cross-processFileLock(timeout=120)and loopback probe; subsequent processes attach to the same running port with zero VRAM overhead. - Continuous Dynamic Micro-Batching: Request queues gather concurrent jobs across an adaptive window (e.g. 20ms) and dispatch them in a single tensor core forward pass.
- HMAC Token Security: Daemons write a random token (
*.tokenwith0o600permissions) and validate all requests in constant time usinghmac.compare_digest. - Text-to-Music & Full Song Synthesis: The
diffusersTTM binding provides high-throughput generation forMiniMaxAI/MiniMax-Music3,stabilityai/stable-audio-open-1.0, andcvssp/audioldm2-music.
- Cross-Process Server Daemon:
VLLMBindingmanages an inter-process daemon. The first process callingload_model()acquires an inter-processFileLock(~/.lollms/bindings_models/vllm_models/global_vllm_manager.lock), spawns an OpenAI-compatible daemon (vllm.entrypoints.openai.api_server), and registers its metadata. Subsequent processes attach directly to the running endpoint with zero VRAM duplication. - Recipe Presets in
description.yaml: Exposes pre-validated deployment recipes (deepseek_r1_mtp,qwen2_5_coder_fp8_kv,qwen2_5_vl_multimodal,llama_3_3_70b_throughput,glm_5_flash_long_ctx,low_vram_consumer). Front-ends can querydesc.get("presets")to populate instant configuration dropdowns. - Clean Command Construction: Any parameter evaluating to
None, empty string, orFalse(for boolean flags) is omitted from the daemon launch arguments.
- Execution Layer Support: LLM bindings accept
videos: Optional[List[str]] = Noneingenerate_text()and structured{"type": "video_url", ...}items ingenerate_from_messages(). - Normalization:
normalize_video_input()converts local files (.mp4,.webm,.mov,.mkv), raw base64 strings, and HTTP/HTTPS URLs into standardized OpenAI/vLLMvideo_urlcontent blocks. - Model Profiles: Set
video_enabled=TrueonLollmsModelProfileto indicate native video comprehension capability.LollmsClient.has_video_capability()enables programmatic capability inspection.
-
Topological Effort Projection:
LollmsLLMBinding.translate_reasoning_effort(effort, supported_efforts)maps continuous compute intensities ($0.0$ to$1.0$ ) to the model's declaredsupported_reasoning_efforts(e.g. mapping"max"to"high"on OpenAI, or"medium"to"high"on GLM-5.3-Flash). -
Streaming Stream Handler (
_StreamThinkingHandler): Routes thinking chunks toMSG_TYPE.MSG_TYPE_THOUGHT_CHUNKwhile preserving visible<think>...</think>XML delimiters in output text. This ensuresLollmsTextProcessor.remove_thinking_blocks()reliably identifies and strips thoughts during multi-turn discussion reprompting. -
GLM-5.3-Flash Sequential Embedding (
glm_image_embedding): Formats Chat Completions content blocks in strict text-first order ([{"type": "text", ...}, {"type": "image_url", ...}]) and suppresses incompatible vLLMenable_thinkingkwargs.
- Directory:
lollms_client/tools_bindings/ - ABC:
LollmsToolBinding(inlollms_tools_binding.py) - Manager:
LollmsTOOLBindingManager(inlollms_tools_binding.py) - Implemented Bindings (Examples):
local_mcp: Discovers and executes local Python tools. Each tool is defined by a<tool_name>.pyfile (with anexecutefunction) and a<tool_name>.mcp.jsonfile (describing the tool's name, description, input/output schemas).- Default Tools for
local_mcp:file_writer: Writes or appends text to files.internet_search: Performs web searches using DuckDuckGo.python_interpreter: Executes Python code snippets in a restricted environment.generate_image_from_prompt: Generates an image by calling theLollmsClient's active TTI binding.
- Default Tools for
The LollmsClient class itself now includes several high-level methods for common AI tasks, previously housed in TasksLibrary. These methods often combine multiple calls to the LLM binding or other utilities. Examples:
sequential_summarize,deep_analyze: For processing and understanding long texts.generate_code,generate_codes: For structured code generation.yes_no,multichoice_question,multichoice_ranking: For specific Q&A formats.extract_code_blocks,extract_thinking_blocks,remove_thinking_blocks: Text processing utilities.
These methods leverage the active LLM binding for their core AI operations.
(Content remains the same)
(Content remains the same)
(Content remains the same, though lollms_functions.py and lollms_tasks.py are gone)
(General steps for adding modality bindings remain the same: Create dir, __init__.py, BindingName, implement ABC, handle deps, add test block)
Contributing a new MCP binding (e.g., to interact with a remote MCP tool server) follows a similar pattern:
- Create Binding Directory:
lollms_client/tools_bindings/<your_mcp_binding_name>/ - Create
__init__.py: Inside the new directory. - Define
BindingName: At the top of your__init__.py.# lollms_client/tools_bindings/my_remote_mcp/__init__.py BindingName = "MyRemoteMCPBinding" # Must match your class name
- Implement the
LollmsToolBindingClass:You must implementfrom lollms_client.lollms_tools_binding import LollmsToolBinding from typing import List, Dict, Any # Import any SDKs or libraries needed to talk to your remote MCP server BindingName = "MyRemoteMCPBinding" class MyRemoteMCPBinding(LollmsToolBinding): def __init__(self, tool_server_url: str, api_key: str = None, **kwargs): super().__init__(binding_name="my_remote_mcp") self.tool_server_url = tool_server_url self.api_key = api_key # Initialize your client for the remote server # self.remote_client = MyRemoteSDK(base_url=tool_server_url, api_key=api_key) def discover_tools(self, **kwargs) -> List[Dict[str, Any]]: # Implement logic to fetch tool definitions from self.tool_server_url # try: # response = self.remote_client.get_tools_list() # return response.json().get("tools", []) # except Exception as e: # # Log error # return [] return [{"name": "remote_tool_example", "description": "Example tool from remote server", "input_schema":{}}] # Placeholder def execute_tool(self, tool_name: str, params: Dict[str, Any], **kwargs) -> Dict[str, Any]]: # Implement logic to call the specific tool on self.tool_server_url # try: # response = self.remote_client.execute(tool_name, params) # return response.json().get("result", {"error": "No result field"}) # except Exception as e: # # Log error # return {"error": f"Failed to execute remote tool {tool_name}: {e}"} return {"output": f"Executed {tool_name} remotely with {params}. Placeholder.", "status_code": 200} # Placeholder
discover_toolsandexecute_tool. - Handle Dependencies: Use
pipmaster.ensure_packages()if your binding requires specific Python packages for communication. - Add a Test Block: Include an
if __name__ == "__main__":block for standalone testing.
If you want to add a new tool to be used by the existing local_mcp binding:
- Choose/Create a Tools Folder: This can be any folder. You'll pass its path to
LollmsClientviatools_binding_config={"tools_folder_path": "your/tools/dir"}.- The
local_mcpbinding also has adefault_toolssubdirectory packaged with it. You can add tools there if modifying the library directly, but using a custom external folder is cleaner for user-defined tools.
- The
- Create Tool Subdirectory: Inside your chosen tools folder, create a subdirectory for your new tool, e.g.,
my_custom_tool/. - Create
<tool_name>.mcp.json: In the tool's subdirectory (e.g.,my_custom_tool/my_custom_tool.mcp.json), define the tool's metadata.{ "name": "my_custom_tool", "description": "A brief description of what my custom tool does.", "input_schema": { "type": "object", "properties": { "param1": {"type": "string", "description": "Description of param1"}, "param2": {"type": "integer", "default": 10} }, "required": ["param1"] }, "output_schema": { /* Define expected output structure */ } } - Create
<tool_name>.py: In the tool's subdirectory (e.g.,my_custom_tool/my_custom_tool.py), implement the tool's logic.Thefrom typing import Dict, Any def execute(params: Dict[str, Any], lollms_client_instance: Any = None) -> Dict[str, Any]: # params will contain {'param1': 'value', 'param2': value_or_default} # lollms_client_instance is the LollmsClient that invoked this, # useful if your tool needs to call TTI, TTS, or even another LLM call. param1_value = params.get("param1") param2_value = params.get("param2", 10) # Use default if not provided # Your tool's logic here result_data = f"Tool processed {param1_value} and {param2_value}" # Return a dictionary (which will be nested under "output" by local_mcp) return {"processed_data": result_data, "status_message": "Custom tool executed successfully."}
executefunction must acceptparams: Dict[str, Any]and can optionally acceptlollms_client_instance: Any. It should return a dictionary.
The local_mcp binding will automatically discover and make this tool available to the LLM when generate_with_mcp is called.
As of v1.21.0, both LollmsDiscussion and LollmsPersonality implement Round-End State Checkpointing and Turn Resumption. This allows applications to survive client disconnections, accidental tab closures, or user pauses during multi-round reasoning tasks.
In agentic mode, an agent turn often spans multiple rounds (tool executions, file writes, sub-agent delegations). Rather than waiting until the entire turn finishes to persist changes, the system saves an atomic checkpoint at the conclusion of every single round.
[Round 1 Start] ββ> [Tool Execution] ββ> [Round 1 End: CHECKPOINT SAVED]
β
[Round 2 Start] ββ> [Artifact Stream] ββ> [Round 2 End: CHECKPOINT SAVED]
β
(Session cut / Stopped)
β
[RESUME TURN TRIGGER]
β
[Round 3 Start: Virtual History Restored] ββ> Continues to completion
Every round checkpoint records:
round_count(int): The current zero-indexed or 1-indexed round number.turn_status(str):"in_progress","cancelled", or"completed".virtual_history(list[dict]): The full sequence of assistant thoughts, tool calls, and<tool_result>messages accumulated in this turn.tool_calls(list[dict]): Tools dispatched and their execution status.workspace_changes(list[dict]): Files created or modified during the turn.timestamp(float/str): UTC time of the checkpoint.
from lollms_client.lollms_personality import LollmsPersonality
personality = LollmsPersonality(...)
# 1. Check if an interrupted turn exists
if personality.has_resumable_turn():
print("Incomplete turn detected. Resuming...")
result = personality.chat(
prompt="organize this folder", # Original prompt is loaded from checkpoint automatically
lollms_client=client,
resume_turn=True, # Re-hydrates virtual history and continues execution
)
else:
result = personality.chat(prompt="organize this folder", lollms_client=client)from lollms_client import LollmsDiscussion
discussion = LollmsDiscussion(...)
# 1. Inspect if the branch tip has an incomplete or paused turn
if discussion.has_resumable_turn():
# 2. Resume execution from the last round checkpoint
result = discussion.resume_turn(
streaming_callback=my_streaming_callback,
)When displaying live streaming code or artifact generation in a UI:
- Header Discipline: The collapsible header/subtitle MUST display only the file name and high-level structural units (
Section: <name>,def <function_name>(),class <ClassName>). Raw content lines (table rows, assignment expressions, code statements) must NEVER be placed in the header. - Body Discipline: The verbatim code content must stream strictly inside the collapsible box.
Dynamic Mode provides real-time adaptive control over the reasoning loop across both LollmsDiscussion and LollmsPersonality.
[Round 1 (Initial)] ββ> Starts at effort="none" (fast execution, zero latency)
β
[Model encounters complex refactor / bug]
β
ββ> Emits: <effort level="high"/> on a new line
β
[Stream Parser] ββ> Intercepts tag, scrubs from user content, sets next_reasoning_effort="high"
β
[Round 2] ββ> Binding translates "high" -> backend parameters:
β’ Ollama: think=True, options["think"]="high"
β’ OpenAI / vLLM: reasoning_effort="high", chat_template_kwargs={"enable_thinking": True}
β
[Task finished or simplified]
β
ββ> Emits: <effort level="none"/> ββ> Concludes with <done/>
In _mixin_chat.py (_StreamState.feed()) and lollms_agent_state.py (_AgentStreamState.feed()):
effort_match = re.search(
r'(?m)^\s*(?!`)(?!.*\|)<effort\b([^>]*)(?:/>|>.*?</effort>)',
self._pending_buffer,
re.IGNORECASE | re.DOTALL
)
if effort_match:
tag_full = effort_match.group(0)
attrs_part = effort_match.group(1)
lvl_match = re.search(r'(?:level|value)=["\']([^"\']+)["\']', attrs_part, re.IGNORECASE)
self.next_reasoning_effort = lvl_match.group(1).lower().strip() if lvl_match else "medium"The tag is immediately removed from the pending buffer so it never enters the visible message content or database history.
The active LLM binding projects the requested level using LollmsLLMBinding.translate_reasoning_effort():
- Translates topological effort scores (
$0.0 \to 1.0$ ) to the model's declaredsupported_reasoning_efforts. - For backends where thinking must be completely silenced (
effort="none"), bindings inject:extra_body = { "thinking": False, "chat_template_kwargs": {"enable_thinking": False, "thinking": False} }
During agent execution in LollmsPersonality.chat():
if active_temperature is None:
has_code_intent = bool(ss.artifact_trigger or ss.tool_trigger or "<artifact" in raw_llm_output_buffer or "<tool" in raw_llm_output_buffer)
gen_kwargs["temperature"] = 0.15 if has_code_intent else 0.70When repetition or preambles loop consecutively, the temperature dynamically escalates:
active_temperature = min(0.95, base_temperature + (0.15 * consecutive_stall_count))Specialist workers spawned via SubAgentSpawner.spawn() or tool_spinoff_agent inherit effort parameters:
agent_spawner.spawn(
instruction="Optimize the inner database indexing routine.",
effort="high",
dynamic_effort=True # Allows child worker to adjust its own effort
)(Content updated to reflect MCP examples and removal of TasksLibrary examples)
- Examples: The
examples/directory contains various scripts.examples/function_calling_with_local_custom_mcp.py: Demonstratesgenerate_with_mcpusing custom local tools.examples/local_mcp.py: Demonstratesgenerate_with_mcpusing the default tools packaged with thelocal_mcpbinding.- Other examples show text generation, multimodal features, etc.
- Binding Self-Tests: Most binding
__init__.pyfiles have anif __name__ == "__main__":block. Run these directly to test that specific binding. - Adding Formal Tests: Contributions of unit/integration tests are welcome.
(Content remains the same)
(Content remains the same)
(Content remains the same)
(Content updated to remove TasksLibrary ideas and potentially add MCP-related ones)
- More Bindings: LLM, TTS, TTI, STT, TTM, TTV, and MCP bindings.
- Enhanced High-Level Ops: Add more sophisticated pre-built operations directly to
LollmsClient. - Asynchronous Operations: Explore
async/await. - Improved Error Handling.
- Standardized Configuration.
- Comprehensive Testing Framework.
- Plugin System for Bindings.
- Documentation Generation.
- More Default Tools for
local_mcp: Consider adding more generally useful local tools. - Support for Remote MCP Tool Servers: Beyond
local_mcp, bindings for standardized remote MCP endpoints.
(Content remains the same)
Thank you for your interest in contributing to lollms_client!