Skip to content

Latest commit

Β 

History

History
480 lines (394 loc) Β· 27.3 KB

File metadata and controls

480 lines (394 loc) Β· 27.3 KB

LoLLMS Client - Developer Documentation

Welcome to the developer documentation for lollms_client! This guide is intended for developers who want to understand the project's architecture, contribute new features, fix bugs, or create new bindings.

Project Links:

Table of Contents

  1. Introduction
  2. Getting Started for Developers
  3. Project Structure
  4. Core Concepts & Architecture
  5. Adding New Bindings (A Practical Guide)
  6. Running Examples & Tests
  7. Coding Standards & Conventions
  8. Contribution Guidelines
  9. Reporting Issues
  10. Roadmap & Future Ideas
  11. Community & Contact

1. Introduction

Project Goal

lollms_client is a Python client library designed to provide a unified and easy-to-use interface for interacting with various AI model backends and services. It aims to abstract the complexities of different APIs and local model execution, allowing developers to seamlessly switch between different AI modalities (LLM, TTS, TTI, STT, TTM, TTV), and enable powerful function calling via the Model Context Protocol (MCP).

The "LoLLMS" ecosystem (Lord of Large Language and Multimodal Systems) aims to be a versatile platform for AI interaction, and this client is a key component for Python-based applications and integrations.

Key Features

  • Modular Binding System: Easily add support for new AI backends or libraries.
  • Multi-Modality: Supports Large Language Models (LLM), Text-to-Speech (TTS), Text-to-Image (TTI), Speech-to-Text (STT), Text-to-Music/Sound (TTM), and Text-to-Video (TTV).
  • Function Calling (MCP): Integrated Model Context Protocol support, including a local_mcp binding for executing local Python tools with default utilities.
  • Unified API: Provides a consistent LollmsClient interface for common operations across different bindings.
  • Local & Remote Backends: Supports both local model execution and remote services.
  • Helper Utilities: Includes tools for discussion management and various AI-driven tasks directly on the LollmsClient.
  • Dependency Management: Uses pipmaster within bindings to attempt to ensure necessary Python packages are installed.

2. Getting Started for Developers

(Content remains largely the same: Prerequisites, Cloning, Virtual Env, Editable Install, Dev Dependencies)

3. Project Structure

Here's an overview of the lollms_client directory structure:

πŸ“ lollms_client/
β”œβ”€ πŸ“ ai_documentation/         # Markdown files generated by AI documentation tools
β”œβ”€ πŸ“ dist/                     # Build artifacts
β”œβ”€ πŸ“ examples/                 # Example scripts demonstrating client usage
β”‚  β”œβ”€ πŸ“ article_summary/
β”‚  β”œβ”€ πŸ“ deep_analyze/
β”‚  β”œβ”€ πŸ“ function_calling_with_local_custom_mcp.py # Shows custom MCP tools
β”‚  β”œβ”€ πŸ“ local_mcp.py                              # Shows default MCP tools
β”‚  β”œβ”€ πŸ“ generate_and_speak/
β”‚  └─ ... (other examples)
β”œβ”€ πŸ“ lollms_client/            # The core library source code
β”‚  β”œβ”€ πŸ“ llm_bindings/           # LLM-specific binding implementations
β”‚  β”‚  └─ ...
β”‚  β”œβ”€ πŸ“ tts_bindings/           # Text-to-Speech binding implementations
β”‚  β”‚  └─ ...
β”‚  β”œβ”€ πŸ“ tti_bindings/           # Text-to-Image binding implementations
β”‚  β”‚  └─ ...
β”‚  β”œβ”€ πŸ“ stt_bindings/           # Speech-to-Text binding implementations
β”‚  β”‚  └─ ...
β”‚  β”œβ”€ πŸ“ ttm_bindings/           # Text-to-Music/Sound binding implementations
β”‚  β”‚  └─ ...
β”‚  β”œβ”€ πŸ“ ttv_bindings/           # Text-to-Video binding implementations
β”‚  β”‚  └─ ...
β”‚  β”œβ”€ πŸ“ tools_bindings/           # Model Context Protocol binding implementations
β”‚  β”‚  └─ πŸ“ local_mcp/          # Example: Local MCP tool executor
β”‚  β”‚     β”œβ”€ πŸ“ default_tools/    # Packaged tools for local_mcp
β”‚  β”‚     β”‚  β”œβ”€ πŸ“ file_writer/
β”‚  β”‚     β”‚  β”œβ”€ πŸ“ generate_image_from_prompt/
β”‚  β”‚     β”‚  β”œβ”€ πŸ“ internet_search/
β”‚  β”‚     β”‚  └─ πŸ“ python_interpreter/
β”‚  β”‚     └─ πŸ“„ __init__.py
β”‚  β”œβ”€ πŸ“„ __init__.py             # Makes lollms_client a package, exports key classes
β”‚  β”œβ”€ πŸ“„ lollms_config.py        # Configuration classes
β”‚  β”œβ”€ πŸ“„ lollms_core.py          # LollmsClient class, main orchestrator (includes generate_with_mcp, summarize, etc.)
β”‚  β”œβ”€ πŸ“„ lollms_discussion.py    # LollmsDiscussion and LollmsMessage classes
β”‚  β”œβ”€ πŸ“„ lollms_llm_binding.py   # ABC for LLM bindings and LollmsLLMBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_tts_binding.py   # ABC for TTS bindings and LollmsTTSBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_tti_binding.py   # ABC for TTI bindings and LollmsTTIBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_stt_binding.py   # ABC for STT bindings and LollmsSTTBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_ttm_binding.py   # ABC for TTM bindings and LollmsTTMBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_ttv_binding.py   # ABC for TTV bindings and LollmsTTVBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_tools_binding.py   # ABC for MCP bindings and LollmsTOOLBindingManager
β”‚  β”œβ”€ πŸ“„ lollms_types.py         # Enums (MSG_TYPE, ELF_COMPLETION_FORMAT, etc.)
β”‚  └─ πŸ“„ lollms_utilities.py     # Helper functions
β”œβ”€ πŸ“ lollms_client.egg-info/   # Packaging metadata
β”œβ”€ πŸ“„ CHANGELOG.md
β”œβ”€ πŸ“„ DOC_DEV.md                # This developer documentation file
β”œβ”€ πŸ“„ DOC_USE.md                # User-focused documentation
β”œβ”€ πŸ“„ pyproject.toml
β”œβ”€ πŸ“„ README.md
└─ πŸ“„ requirements.txt

4. Core Concepts & Architecture

LollmsClient (The Orchestrator)

Located in lollms_client/lollms_core.py, the LollmsClient class is the main entry point. Its responsibilities include:

  • Initializing and managing different types of bindings (LLM, TTS, TTI, STT, TTM, TTV, MCP).
  • Providing a unified API for common operations like text generation (generate_text), speech synthesis (tts.generate_audio), and function calling (generate_with_mcp).
  • Orchestrating interactions involving MCP tools by leveraging an active MCP binding.
  • Offering direct methods for high-level tasks like sequential_summarize, deep_analyze, generate_code, yes_no, etc.
  • Storing default generation parameters.

Bindings (The Backends)

(General description of Bindings, ABCs, and Binding Managers remains the same)

Supported Modalities and Bindings

Shared Model Server Daemon Architecture (Project Phenix)

  • Zero Port Drift Singletons: Every local daemon binding (whisper on 9633, xtts on 9634, piper_tts on 9635, bark on 9636, diffusers TTI on 9632, diffusers TTM on 9637) adheres to the "first instance wins" pattern. The first process spawns the daemon with a cross-process FileLock(timeout=120) and loopback probe; subsequent processes attach to the same running port with zero VRAM overhead.
  • Continuous Dynamic Micro-Batching: Request queues gather concurrent jobs across an adaptive window (e.g. 20ms) and dispatch them in a single tensor core forward pass.
  • HMAC Token Security: Daemons write a random token (*.token with 0o600 permissions) and validate all requests in constant time using hmac.compare_digest.
  • Text-to-Music & Full Song Synthesis: The diffusers TTM binding provides high-throughput generation for MiniMaxAI/MiniMax-Music3, stabilityai/stable-audio-open-1.0, and cvssp/audioldm2-music.

vLLM Server Mutualization & Recipe Presets

  • Cross-Process Server Daemon: VLLMBinding manages an inter-process daemon. The first process calling load_model() acquires an inter-process FileLock (~/.lollms/bindings_models/vllm_models/global_vllm_manager.lock), spawns an OpenAI-compatible daemon (vllm.entrypoints.openai.api_server), and registers its metadata. Subsequent processes attach directly to the running endpoint with zero VRAM duplication.
  • Recipe Presets in description.yaml: Exposes pre-validated deployment recipes (deepseek_r1_mtp, qwen2_5_coder_fp8_kv, qwen2_5_vl_multimodal, llama_3_3_70b_throughput, glm_5_flash_long_ctx, low_vram_consumer). Front-ends can query desc.get("presets") to populate instant configuration dropdowns.
  • Clean Command Construction: Any parameter evaluating to None, empty string, or False (for boolean flags) is omitted from the daemon launch arguments.

Multimodal Video Input & Comprehension

  • Execution Layer Support: LLM bindings accept videos: Optional[List[str]] = None in generate_text() and structured {"type": "video_url", ...} items in generate_from_messages().
  • Normalization: normalize_video_input() converts local files (.mp4, .webm, .mov, .mkv), raw base64 strings, and HTTP/HTTPS URLs into standardized OpenAI/vLLM video_url content blocks.
  • Model Profiles: Set video_enabled=True on LollmsModelProfile to indicate native video comprehension capability. LollmsClient.has_video_capability() enables programmatic capability inspection.

Reasoning Effort Projection & Thought Stream Handling

  • Topological Effort Projection: LollmsLLMBinding.translate_reasoning_effort(effort, supported_efforts) maps continuous compute intensities ($0.0$ to $1.0$) to the model's declared supported_reasoning_efforts (e.g. mapping "max" to "high" on OpenAI, or "medium" to "high" on GLM-5.3-Flash).
  • Streaming Stream Handler (_StreamThinkingHandler): Routes thinking chunks to MSG_TYPE.MSG_TYPE_THOUGHT_CHUNK while preserving visible <think>...</think> XML delimiters in output text. This ensures LollmsTextProcessor.remove_thinking_blocks() reliably identifies and strips thoughts during multi-turn discussion reprompting.
  • GLM-5.3-Flash Sequential Embedding (glm_image_embedding): Formats Chat Completions content blocks in strict text-first order ([{"type": "text", ...}, {"type": "image_url", ...}]) and suppresses incompatible vLLM enable_thinking kwargs.

MCP (Model Context Protocol) Bindings

  • Directory: lollms_client/tools_bindings/
  • ABC: LollmsToolBinding (in lollms_tools_binding.py)
  • Manager: LollmsTOOLBindingManager (in lollms_tools_binding.py)
  • Implemented Bindings (Examples):
    • local_mcp: Discovers and executes local Python tools. Each tool is defined by a <tool_name>.py file (with an execute function) and a <tool_name>.mcp.json file (describing the tool's name, description, input/output schemas).
      • Default Tools for local_mcp:
        • file_writer: Writes or appends text to files.
        • internet_search: Performs web searches using DuckDuckGo.
        • python_interpreter: Executes Python code snippets in a restricted environment.
        • generate_image_from_prompt: Generates an image by calling the LollmsClient's active TTI binding.

High-Level Operations (in LollmsClient)

The LollmsClient class itself now includes several high-level methods for common AI tasks, previously housed in TasksLibrary. These methods often combine multiple calls to the LLM binding or other utilities. Examples:

  • sequential_summarize, deep_analyze: For processing and understanding long texts.
  • generate_code, generate_codes: For structured code generation.
  • yes_no, multichoice_question, multichoice_ranking: For specific Q&A formats.
  • extract_code_blocks, extract_thinking_blocks, remove_thinking_blocks: Text processing utilities.

These methods leverage the active LLM binding for their core AI operations.

LollmsDiscussion & LollmsMessage

(Content remains the same)

Configuration (lollms_config.py)

(Content remains the same)

Utilities & Types

(Content remains the same, though lollms_functions.py and lollms_tasks.py are gone)

5. Adding New Bindings (A Practical Guide)

General Steps for Modality Bindings (LLM, TTS, TTI, STT, TTM, TTV)

(General steps for adding modality bindings remain the same: Create dir, __init__.py, BindingName, implement ABC, handle deps, add test block)

Adding New MCP Bindings

Contributing a new MCP binding (e.g., to interact with a remote MCP tool server) follows a similar pattern:

  1. Create Binding Directory: lollms_client/tools_bindings/<your_mcp_binding_name>/
  2. Create __init__.py: Inside the new directory.
  3. Define BindingName: At the top of your __init__.py.
    # lollms_client/tools_bindings/my_remote_mcp/__init__.py
    BindingName = "MyRemoteMCPBinding" # Must match your class name
  4. Implement the LollmsToolBinding Class:
    from lollms_client.lollms_tools_binding import LollmsToolBinding
    from typing import List, Dict, Any
    # Import any SDKs or libraries needed to talk to your remote MCP server
    
    BindingName = "MyRemoteMCPBinding"
    
    class MyRemoteMCPBinding(LollmsToolBinding):
        def __init__(self, tool_server_url: str, api_key: str = None, **kwargs):
            super().__init__(binding_name="my_remote_mcp")
            self.tool_server_url = tool_server_url
            self.api_key = api_key
            # Initialize your client for the remote server
            # self.remote_client = MyRemoteSDK(base_url=tool_server_url, api_key=api_key)
    
        def discover_tools(self, **kwargs) -> List[Dict[str, Any]]:
            # Implement logic to fetch tool definitions from self.tool_server_url
            # try:
            #     response = self.remote_client.get_tools_list()
            #     return response.json().get("tools", [])
            # except Exception as e:
            #     # Log error
            #     return []
            return [{"name": "remote_tool_example", "description": "Example tool from remote server", "input_schema":{}}] # Placeholder
    
        def execute_tool(self, tool_name: str, params: Dict[str, Any], **kwargs) -> Dict[str, Any]]:
            # Implement logic to call the specific tool on self.tool_server_url
            # try:
            #     response = self.remote_client.execute(tool_name, params)
            #     return response.json().get("result", {"error": "No result field"})
            # except Exception as e:
            #     # Log error
            #     return {"error": f"Failed to execute remote tool {tool_name}: {e}"}
            return {"output": f"Executed {tool_name} remotely with {params}. Placeholder.", "status_code": 200} # Placeholder
    You must implement discover_tools and execute_tool.
  5. Handle Dependencies: Use pipmaster.ensure_packages() if your binding requires specific Python packages for communication.
  6. Add a Test Block: Include an if __name__ == "__main__": block for standalone testing.

Adding New Local MCP Tools (for local_mcp binding)

If you want to add a new tool to be used by the existing local_mcp binding:

  1. Choose/Create a Tools Folder: This can be any folder. You'll pass its path to LollmsClient via tools_binding_config={"tools_folder_path": "your/tools/dir"}.
    • The local_mcp binding also has a default_tools subdirectory packaged with it. You can add tools there if modifying the library directly, but using a custom external folder is cleaner for user-defined tools.
  2. Create Tool Subdirectory: Inside your chosen tools folder, create a subdirectory for your new tool, e.g., my_custom_tool/.
  3. Create <tool_name>.mcp.json: In the tool's subdirectory (e.g., my_custom_tool/my_custom_tool.mcp.json), define the tool's metadata.
    {
        "name": "my_custom_tool",
        "description": "A brief description of what my custom tool does.",
        "input_schema": {
            "type": "object",
            "properties": {
                "param1": {"type": "string", "description": "Description of param1"},
                "param2": {"type": "integer", "default": 10}
            },
            "required": ["param1"]
        },
        "output_schema": { /* Define expected output structure */ }
    }
  4. Create <tool_name>.py: In the tool's subdirectory (e.g., my_custom_tool/my_custom_tool.py), implement the tool's logic.
    from typing import Dict, Any
    
    def execute(params: Dict[str, Any], lollms_client_instance: Any = None) -> Dict[str, Any]:
        # params will contain {'param1': 'value', 'param2': value_or_default}
        # lollms_client_instance is the LollmsClient that invoked this,
        # useful if your tool needs to call TTI, TTS, or even another LLM call.
        
        param1_value = params.get("param1")
        param2_value = params.get("param2", 10) # Use default if not provided
    
        # Your tool's logic here
        result_data = f"Tool processed {param1_value} and {param2_value}"
        
        # Return a dictionary (which will be nested under "output" by local_mcp)
        return {"processed_data": result_data, "status_message": "Custom tool executed successfully."}
    The execute function must accept params: Dict[str, Any] and can optionally accept lollms_client_instance: Any. It should return a dictionary.

The local_mcp binding will automatically discover and make this tool available to the LLM when generate_with_mcp is called.

πŸ”„ Turn Checkpointing & Resumption Architecture

As of v1.21.0, both LollmsDiscussion and LollmsPersonality implement Round-End State Checkpointing and Turn Resumption. This allows applications to survive client disconnections, accidental tab closures, or user pauses during multi-round reasoning tasks.

1. Conceptual Model

In agentic mode, an agent turn often spans multiple rounds (tool executions, file writes, sub-agent delegations). Rather than waiting until the entire turn finishes to persist changes, the system saves an atomic checkpoint at the conclusion of every single round.

[Round 1 Start] ──> [Tool Execution] ──> [Round 1 End: CHECKPOINT SAVED]
                                                      β”‚
[Round 2 Start] ──> [Artifact Stream] ──> [Round 2 End: CHECKPOINT SAVED]
                                                      β”‚
                                          (Session cut / Stopped)
                                                      β”‚
                                            [RESUME TURN TRIGGER]
                                                      β”‚
[Round 3 Start: Virtual History Restored] ──> Continues to completion

2. Checkpoint Data Structure

Every round checkpoint records:

  • round_count (int): The current zero-indexed or 1-indexed round number.
  • turn_status (str): "in_progress", "cancelled", or "completed".
  • virtual_history (list[dict]): The full sequence of assistant thoughts, tool calls, and <tool_result> messages accumulated in this turn.
  • tool_calls (list[dict]): Tools dispatched and their execution status.
  • workspace_changes (list[dict]): Files created or modified during the turn.
  • timestamp (float/str): UTC time of the checkpoint.

3. Upgrading Third-Party Applications

Mode A: Using LollmsPersonality (Headless or Agent Apps)

from lollms_client.lollms_personality import LollmsPersonality

personality = LollmsPersonality(...)

# 1. Check if an interrupted turn exists
if personality.has_resumable_turn():
    print("Incomplete turn detected. Resuming...")
    result = personality.chat(
        prompt="organize this folder", # Original prompt is loaded from checkpoint automatically
        lollms_client=client,
        resume_turn=True, # Re-hydrates virtual history and continues execution
    )
else:
    result = personality.chat(prompt="organize this folder", lollms_client=client)

Mode B: Using LollmsDiscussion (Chat, WebUI, and Branching Apps)

from lollms_client import LollmsDiscussion

discussion = LollmsDiscussion(...)

# 1. Inspect if the branch tip has an incomplete or paused turn
if discussion.has_resumable_turn():
    # 2. Resume execution from the last round checkpoint
    result = discussion.resume_turn(
        streaming_callback=my_streaming_callback,
    )

🎨 Collapsible Artifact UI Protocol

When displaying live streaming code or artifact generation in a UI:

  1. Header Discipline: The collapsible header/subtitle MUST display only the file name and high-level structural units (Section: <name>, def <function_name>(), class <ClassName>). Raw content lines (table rows, assignment expressions, code statements) must NEVER be placed in the header.
  2. Body Discipline: The verbatim code content must stream strictly inside the collapsible box.

⚑ Dynamic Mode Architecture & Autonomous Reasoning Escalation

Dynamic Mode provides real-time adaptive control over the reasoning loop across both LollmsDiscussion and LollmsPersonality.

[Round 1 (Initial)] ──> Starts at effort="none" (fast execution, zero latency)
                                     β”‚
                 [Model encounters complex refactor / bug]
                                     β”‚
           ──> Emits: <effort level="high"/> on a new line
                                     β”‚
[Stream Parser] ──> Intercepts tag, scrubs from user content, sets next_reasoning_effort="high"
                                     β”‚
[Round 2]       ──> Binding translates "high" -> backend parameters:
                    β€’ Ollama: think=True, options["think"]="high"
                    β€’ OpenAI / vLLM: reasoning_effort="high", chat_template_kwargs={"enable_thinking": True}
                                     β”‚
                 [Task finished or simplified]
                                     β”‚
           ──> Emits: <effort level="none"/> ──> Concludes with <done/>

1. Interception and State Machine

In _mixin_chat.py (_StreamState.feed()) and lollms_agent_state.py (_AgentStreamState.feed()):

effort_match = re.search(
    r'(?m)^\s*(?!`)(?!.*\|)<effort\b([^>]*)(?:/>|>.*?</effort>)',
    self._pending_buffer,
    re.IGNORECASE | re.DOTALL
)
if effort_match:
    tag_full = effort_match.group(0)
    attrs_part = effort_match.group(1)
    lvl_match = re.search(r'(?:level|value)=["\']([^"\']+)["\']', attrs_part, re.IGNORECASE)
    self.next_reasoning_effort = lvl_match.group(1).lower().strip() if lvl_match else "medium"

The tag is immediately removed from the pending buffer so it never enters the visible message content or database history.

2. Multi-Backend Projection

The active LLM binding projects the requested level using LollmsLLMBinding.translate_reasoning_effort():

  • Translates topological effort scores ($0.0 \to 1.0$) to the model's declared supported_reasoning_efforts.
  • For backends where thinking must be completely silenced (effort="none"), bindings inject:
    extra_body = {
        "thinking": False,
        "chat_template_kwargs": {"enable_thinking": False, "thinking": False}
    }

3. Task-Adapted Sampling Temperature

During agent execution in LollmsPersonality.chat():

if active_temperature is None:
    has_code_intent = bool(ss.artifact_trigger or ss.tool_trigger or "<artifact" in raw_llm_output_buffer or "<tool" in raw_llm_output_buffer)
    gen_kwargs["temperature"] = 0.15 if has_code_intent else 0.70

When repetition or preambles loop consecutively, the temperature dynamically escalates:

active_temperature = min(0.95, base_temperature + (0.15 * consecutive_stall_count))

4. Sub-Agent and Spinoff Delegation

Specialist workers spawned via SubAgentSpawner.spawn() or tool_spinoff_agent inherit effort parameters:

agent_spawner.spawn(
    instruction="Optimize the inner database indexing routine.",
    effort="high",
    dynamic_effort=True  # Allows child worker to adjust its own effort
)

6. Running Examples & Tests

(Content updated to reflect MCP examples and removal of TasksLibrary examples)

  • Examples: The examples/ directory contains various scripts.
    • examples/function_calling_with_local_custom_mcp.py: Demonstrates generate_with_mcp using custom local tools.
    • examples/local_mcp.py: Demonstrates generate_with_mcp using the default tools packaged with the local_mcp binding.
    • Other examples show text generation, multimodal features, etc.
  • Binding Self-Tests: Most binding __init__.py files have an if __name__ == "__main__": block. Run these directly to test that specific binding.
  • Adding Formal Tests: Contributions of unit/integration tests are welcome.

7. Coding Standards & Conventions

(Content remains the same)

8. Contribution Guidelines

(Content remains the same)

9. Reporting Issues

(Content remains the same)

10. Roadmap & Future Ideas

(Content updated to remove TasksLibrary ideas and potentially add MCP-related ones)

  • More Bindings: LLM, TTS, TTI, STT, TTM, TTV, and MCP bindings.
  • Enhanced High-Level Ops: Add more sophisticated pre-built operations directly to LollmsClient.
  • Asynchronous Operations: Explore async/await.
  • Improved Error Handling.
  • Standardized Configuration.
  • Comprehensive Testing Framework.
  • Plugin System for Bindings.
  • Documentation Generation.
  • More Default Tools for local_mcp: Consider adding more generally useful local tools.
  • Support for Remote MCP Tool Servers: Beyond local_mcp, bindings for standardized remote MCP endpoints.

11. Community & Contact

(Content remains the same)

Thank you for your interest in contributing to lollms_client!