Skip to content

Latest commit

 

History

History
328 lines (269 loc) · 69.3 KB

File metadata and controls

328 lines (269 loc) · 69.3 KB

AGENTS.md

This document is the LeapFlow engineering collaboration contract. It is not only a style guide: it defines the design, runtime, UX, safety, and verification rules that every implementation change must follow.

Design Philosophy

  1. Signal-Driven Intelligence — All agent intelligence derives from observing real-world signals, not from hardcoded rules. If a behavior cannot be learned from signals, it is not in scope.

  2. Context Pipeline as Core — Signal → Filter (SNR) → Compress (intent-preserving) → Store (multi-layer) → Retrieve (goal-dependent) → Decide. Every feature and every external signal source, including IM collaboration events, must map to this pipeline before it can drive action.

  3. Everything Is a Plugin — Capability is composed, not built in. Tools, LLM backends, platform adapters, signal sources, and vision processors all arrive behind a runtime_checkable Protocol and are discovered, injected, and disposed by the same machinery. A capability that can only exist by editing core is a design failure; the answer is a new Protocol, not a special case.

  4. Progressive Trust — Never auto-execute on first encounter. Autonomy is earned through repeated observed success and lost the same way: DRAFT → CANDIDATE → VERIFIED → PRODUCTION on consecutive successes, demotion on consecutive failures, permanent freeze on an internal defect. Trust is per plugin, persisted, and the only legitimate source of an approval exemption.

  5. Occam's Razor — The simplest correct solution wins. Reject complexity that doesn't directly serve user value. Every abstraction must pay for itself.

  6. LLM-Native Design — Design for LLM reasoning first. Protocols over classes. Declarative over imperative. Context over configuration.

  7. User-Centric Reliability — User experience is part of correctness. Every change must keep common paths easy, predictable, recoverable, and must not degrade adjacent workflows.

  8. Environment–Harness Co-evolution — An environmental change is an observable hypothesis, never authority to mutate the framework. Only a typed, task-relevant delta that survives evidence validation and live capability resolution may produce a governed capability requirement; policy, approval, trust, and lifecycle control every subsequent plugin addition, change, disablement, or removal. The complete causal record — including rejected and no-op branches — must remain available so evolution is explainable, reproducible, and reversible.

Code Quality Requirements

  • SOLID principles are non-negotiable; implementations must be cohesive, well-factored, and easy to reason about
  • Occam's Razor: maximize elegance and efficiency, reject unnecessary complexity
  • Design for generalization and universality; prefer reusable domain concepts over one-off special cases
  • Easy to extend, avoid hardcoding and hard rules
  • Industrial-grade robustness: every external call has timeout, retry, and fallback
  • Structured error propagation: failures must flow as typed envelopes (FailureEnvelope) with source, category, recoverability, and side-effect state — never as ad-hoc dicts, bare strings, or unclassified exceptions
  • User experience is a first-class quality bar: optimize for clarity, ease of use, fast feedback, and graceful recovery
  • All comments and docstrings in English
  • Type annotations on all public APIs
  • No bare except — always specify exception types

Architecture Principles

  • System Boundary Awareness: LeapFlow is a multi-entry, multi-module runtime. Changes must account for the affected path across CLI/TUI, leapd, engine, plugins, skills/tools, LLM, storage, memory, gateway, hub, and platform adapters.

  • TUI as the Primary User Entry: The interactive TUI is the default product surface. Preserve streaming feedback, command queue behavior, approval prompts, status bar accuracy, long-input robustness, history, and session continuity.

  • Concurrent TUI Instances Are a Supported Scenario (MANDATORY): several TUIs in different workspaces, sharing one leapd and one profile, is a normal way to use LeapFlow — not an edge case. Each instance must remain fully usable and must see only its own session, conversation, context usage, and turn state. A change to session routing, status(), stream metadata, the client lease, or anything the status bar renders is not verified until it has been exercised with two instances in two workspaces at the same time. One instance degrading another is a release blocker, not a limitation to document.

  • TUI Command Clarity: Global task-control commands stay short and unambiguous (/cancel, /skip, /pause, /resume, /queue, /drop); teach-mode controls must use the /teach ... namespace and should not keep bare compatibility aliases during early iteration.

  • TUI Prompt Ownership: Input prompt and placeholder rendering must have a single owner. Avoid duplicate prompt sources; placeholder text stays visually subordinate, offset after the prompt, and disappears as soon as the user types.

  • leapd Runtime Consistency: Daemon-backed behavior must preserve lifecycle correctness: start, stop, restart, status, RPC streaming, cancellation, pending approvals, runtime config reload, multi-client state, and version consistency.

  • Progressive Context Disclosure (PCD): Keep one unified execution loop, but never default every turn to full disclosure. Each LLM call must use the smallest sufficient PromptAssemblyPlan for tools, memory, history, reasoning, streaming, and risk; upgrade progressively only when observable signals require it.

    Prefix cache stability is part of PCD correctness. The prompt prefix — system prompt static body, tool catalog text, and frozen append-only summary segments — must be byte-stable across consecutive turns within a committed task. Volatile per-turn content (memory retrieval, distilled knowledge, semantic focus, session summary) must be injected as a separate _volatile_context: True system message so the cache optimizer excludes it from the stable prefix. A change that moves per-turn content into the stable prefix, or that reorders stable-prefix messages after cache markers are applied, silently degrades cache hit rate and is a correctness defect, not a performance trade-off.

  • System Prompt Template Anchors Are Structural Contracts (MANDATORY): AnthropicCacheStrategy splits the system prompt into a cached static head and an uncached dynamic tail using deterministic anchors (_STATIC_TERMINAL_ANCHOR and _KNOWN_STATIC_HEADERS in prompt_cache.py). An edit to UNIFIED_SYSTEM_TEMPLATE that moves, renames, or removes a known static header, or relocates the terminal anchor line, must update the anchor set in prompt_cache.py in the same change. Adding a new always-present section to the static body requires adding its header to _KNOWN_STATIC_HEADERS.

  • Prefix Commitment Is Monotonic; Enforcement Is Revocable: CommitmentStatus transitions only from UNCOMMITTED to COMMITTED within a task. The enforcement snapshot (frozen disclosure level, tool set, system prompt hash) is broken by structural disruptions (posture change, tool error, slash command, transform retry) and re-established when the prefix stabilizes. Code that introduces a new structural disruption must call _maybe_break_commitment with the appropriate flag; code that changes commitment conditions (difficulty threshold, posture gates, amortization model) must update PrefixCommitmentConfig and verify test_prefix_commitment_enforcement.py.

  • Compression Must Preserve Append-Only Frozen Segments: SummarizeStage marks its output with _compressed_summary: True; the append-only mode freezes earlier summaries into the head. DropStage preserves contiguous frozen segments after the system head. A new compression stage must not mutate, reorder, or remove messages carrying _compressed_summary: True — these segments are part of the byte-stable prefix.

  • Cache Hit Rate Is a Dual-Caliber Observable: TurnUsageTracker reports per-turn, token-weighted cumulative, and steady-state (cold-start-excluded) cache hit rates. Provider adapters must propagate cached_tokens (OpenAI/DeepSeek) or both cache_read_input_tokens and cache_creation_input_tokens (Anthropic) in the usage dict. A new provider that omits these keys renders the cache metric blind.

  • Session Resume Cache Policy Must Be End-to-End: session_resume_cache_policy: "cache_priority" persists the system prompt and tool schema at turn end and freezes them on resume so the first turn hits the provider prefix cache. Adding a new field to system prompt assembly that is not captured by _last_system_prompt will silently diverge the resumed prefix, causing a miss on the most expensive turn.

  • Task Environment Is a First-Class Signal: Environment adaptation identifies task environments through declared descriptors and typed structural or affordance deltas. Host capabilities and application affordances are separate concepts; an opaque "changed" hash or a host fingerprint alone cannot establish task compatibility, loss, or a need to evolve.

  • Resolution Before Acquisition: A classified environment signal carries redacted provenance and becomes a capability requirement only when it is relevant to the active task. Resolve that requirement against the live catalog before opening a proposal: a satisfiable requirement is recorded as a no-op, while an unmet task-critical requirement follows the typed capability-unavailable/recovery path. No environment change may create a duplicate capability merely because it was observed.

  • Experiment Control Plane Is Not the Subject: Environment sources, fault injection, and harness adapters are discoverable plugins that drive only real production seams. They must not register tools directly, hand-write observation or proposal records, impersonate approval, or introduce a parallel mutation path. Synthetic signals are explicitly enabled only in an experiment profile; production defaults remain unchanged until an operator opts in through the normal configuration surface.

  • Evolution Causality Is Durable and Queryable: For each observed environmental delta, preserve a redacted causal chain linking source and before/after descriptors, evidence and requirement identifiers, live-resolution result, policy and approval decisions, affected plugin/version/fiber, validation and trust outcomes, and any rollback or retirement. A rejected gate, duplicate resolution, failed validation, or no-op is a first-class result, not missing telemetry; it proves why the framework did not change.

  • Gateway as Signal Boundary: External IM/platform integrations are not just messaging features; they extend LeapFlow's Observe/Orient boundary into collaboration environments. Inbound platform events must enter as structured signals (BackendEvent → normalized domain event/message), pass SNR filtering and privacy/safety gates, then feed memory, decision, and action paths according to their classification.

  • Transport-Lifecycle Separation: Short-lived actions (ExecutionBackend/CliBackend) and long-lived observations (BackendEventSource) are separate responsibilities. Do not implement streaming subscribers, webhooks, polling loops, or CLI NDJSON consumers inside one-shot action execution code.

  • Platform-Neutral Gateway Core: Gateway core owns protocols, lifecycle, routing, session isolation, approval, audit, and memory integration. Platform adapters own authentication, send semantics, event-source configuration, and schema normalization. Core modules must not import platform SDKs directly. Per-vendor code — including credential validators — lives in a platform sub-package (adapters/, normalizers/, action_packs/, validators/<platform>.py), never in a core module; core keeps only the neutral registry and contracts.

  • Platform vs App Business Boundary: Platform layers may define stable contracts and governance primitives (ActionSpec, ActionFailure, ActionAuthSpec, CapabilityHealthLedger, approval/feasibility gates, audit, and metadata propagation). Third-party app or vendor specifics — SDK/CLI wire formats, scope names, auth commands, console URLs, error JSON shapes, resource naming, and recovery playbooks — must live in that app's action pack, adapter, backend, or normalizer, never in gateway core.

  • Dependency Inversion: Core logic depends on Protocol abstractions, never on concrete implementations

  • Protocol over ABC: Use typing.Protocol with runtime_checkable for all extension points

  • Event-Driven Communication: Modules interact through typed events on EventBus, not direct imports

  • Session Engine is the Only Reporting Source (MANDATORY): conversation state lives on the per-session engines built by SessionRegistry. ctx.engine is only the template they are cloned from and never accumulates turns, context, or history. Any code reporting runtime state — stream chunk metadata, status(), session analysis, dashboards, status bar values — must resolve the engine through the single entry point (RuntimeLeapService._active_engine() / SessionCoordinator.resolve_session_engine()), never getattr(ctx, "engine"). Reading the template silently yields zeros, which is invisible in review and has surfaced repeatedly as an empty LeapBoard and a status bar frozen at 0/<limit>. When a value is produced by a specific engine (e.g. a stream event), pass that engine explicitly instead of re-resolving, so concurrent sessions cannot be cross-reported.

  • Client-Visible Runtime State Must Be Pushed, Not Inferred: a daemon-mode TUI is a separate process; it seeds model, context length, and usage at startup and can only learn about later changes from metadata the daemon returns. Any runtime value the status bar renders must travel on ordinary status/stream metadata (and on the mutation payload for command RPCs) — never rely on a one-off change notification, which change-detection can legitimately skip.

  • Session Identity Belongs to the Client That Created It (MANDATORY): a session_id is the client's own identity, not shared daemon state. The daemon must never report a session a caller did not name, and a client must never adopt a session id it did not ask for. Concretely: every RPC that returns session-scoped state (status(), stream chunk metadata, history, analysis) takes the caller's session_id and resolves that session; when the caller names none, the reply carries no session identity and no per-session figures rather than a substitute. A client accepts a reported session id only when it matches its own or it has none yet (first assignment, or an explicit --resume). Violating this is not cosmetic: a second TUI adopted the first's session, sent it with its own workspace, and was rejected on every turn — with advice ("start a fresh session") that could not work, because a fresh client re-adopted the same id on its first status poll.

  • Cross-Session Fallbacks Must Be Named For What They Do: a resolver that answers "whichever session was most recently active" ignores workspace and client identity, so it is valid only for genuinely aggregate views (a dashboard summarizing all activity). Such helpers must say so in their name (e.g. most_recent_any_client()), and no code path that describes one caller may use them. A friendly name like "the current session" invites exactly the misuse that leaked one client's identity to another.

  • Workspace Binding Is Part of Session Identity: a session is bound to the workspace of its first request, and reuse from another workspace is refused. Because correct clients cannot trigger this, a mismatch means either an explicit --resume into another workspace (the only legitimate cause, and what the message must name) or a defect — never something to tell the user to work around.

  • Terminal Output Must Wrap at the Console Layer: LeapConsole owns wrapping (soft_wrap=False), because prompt_toolkit's renderer clips at the window edge rather than reflowing. Never enable soft_wrap on the shared console or hand it pre-formatted long lines; long answers silently lose their tail. A standalone Console for fixed-width art (e.g. the banner) may opt out, but must set an explicit width.

  • Immutable Domain Types: Use @dataclass(frozen=True) or NamedTuple for domain objects — but never for an exception type. CPython stores the traceback on the instance and every Python-level re-raise assigns __traceback__ through __setattr__ (contextlib.__exit__, asyncio, and pytest all do), so a frozen exception raises FrozenInstanceError there and replaces the real failure with that noise: a driver answering "Unknown tool" was reported for months as "cannot assign to field 'traceback'".

  • Config-Driven Behavior: Thresholds, intervals, feature flags, model budgets, platform capabilities, hub backends, gateway manifests, and paths must be configurable through Settings/env/config layers.

  • Graceful Degradation: Every optional component (LLM, Hub) can be absent without crash

  • Single Source of Truth: DuckDB for persistence, EventBus for communication, Settings for configuration

  • Inbound Signal Classification: Platform events must be classified before they activate the agent. Message/callback events may enter Decide; signal/lifecycle events should be stored or routed without triggering LLM by default; ignored events must be explicit (e.g. self-message, duplicate, blocked scope).

  • Single Recovery Decision Point: All agent loop errors (LLM, tool, system, security) enter one RecoveryCoordinator. No parallel decision paths, no scattered if/break logic. The pipeline is always: FailureEnvelope → RecoveryDecision → StrategyOutcome feedback.

  • A Local Defect Is Never a Provider Failure (MANDATORY): an except around a provider call must wrap only that call. Post-response bookkeeping — usage recording, calibration, capability learning — belongs outside it, in a helper that contains its own failures, because telemetry must never fail a turn. Exceptions that mean "LeapFlow has a bug" (AttributeError, TypeError, NameError, KeyError, IndexError, ImportError, AssertionError, NotImplementedError) are classified by type into the non-recoverable internal_defect category before the provider taxonomy is consulted. The provider classifier matches on message text, so one mistyped attribute whose name contained "context" was read as a context overflow and driven through three compressions, a provider failover, and a credential rotation before halting — every turn, for every user, with the suite green.

  • Every Terminal Decision Must Be Actionable: a halt is the last thing a stopped turn can say, so it always carries an InteractionRequest naming the failure and the next step. The raw reason is written for the audit log; surfacing it alone produces internal jargon as the user's entire answer ("No applicable recovery strategy found").

  • Recovery Must Leave Evidence: every recovery entry point logs the exception with exc_info, and the audit sink is constructed with the profile layout's audit path. An in-memory sink and an unlogged branch made a reproducible outage undiagnosable after the fact — the incident had to be reconstructed from token counts in an unrelated log line.

  • Side-Effect Gating: Recovery is gated by SideEffectState at two levels. Within a tool batch, a failed side-effecting call stops the remaining calls in that batch, decided by the declared execution_policy rather than any tool-name list. Within RecoveryCoordinator, any state other than NONE blocks replaying actions (retry, transform-and-retry, failover) and yields a checkpointed halt carrying an InteractionRequest; only user-mediated or checkpoint-based resumption is permitted. UNKNOWN blocks like COMMITTED and PARTIAL do: it is the classifier's fallback and the state assigned to external_side_effect (outbound sends, external API calls), so exempting it would leave the highest-risk case ungated.

  • Uncertain Effects Are Reported, Not Retried Blindly: a failed call whose effect may already have landed (external_side_effect, mutating_once) must carry that verdict in its result so the next turn verifies before repeating it. An error is not proof that nothing happened. Idempotent mutations are exempt — re-applying them converges, so flagging them would only stall safe retries.

  • Budget-Constrained Recovery: Turn-level deadlines, per-category limits, and a global recovery budget prevent infinite retry loops. Every recovery action has an explicit cost; exhaustion triggers a clean halt or user escalation.

  • Recovery Strategy as Protocol: Recovery strategies implement a RecoveryStrategy Protocol (can_apply + decide), registered by priority, composable, and extensible without modifying the coordinator.

  • Tool Output Must Survive Rendering Intact (MANDATORY): text produced by tool execution — especially error tracebacks, JSON payloads, and structured diagnostics — must reach the user without silent corruption. Angle-bracketed identifiers (<string>, <module>, <stdin>) in Python tracebacks, and any content that resembles HTML tags, must be escaped or code-fenced before passing through Markdown renderers. A silently stripped traceback is worse than no traceback — it misdirects investigation. The _sanitize_final_response pipeline owns this guarantee for the TUI path.

Engine Module Architecture Rules

The engine module (leapflow/engine/) is the agent's core execution surface — prompt assembly, tool dispatch, recovery, context management, and session orchestration all live here. A structural mistake in engine/ degrades every turn for every user, so the rules below encode the decomposition invariants established during the refactoring that broke a 3,000+ line monolith into composable sub-packages.

  • Sub-package structure is the module's type system: engine/ is organized into five sub-packages — recovery/ (error classification, recovery coordination, strategies), context/ (compression, control, disclosure, focus), task_planning/ (task graph, planner, scheduler), tools/ (execution policy, concurrency, guardrails, action executor), session/ (session controller, session factory). New functionality must be placed in the sub-package whose boundary it fits. If no existing sub-package fits, propose and justify a new one with clear responsibilities before implementation — do not add domain-specific logic to the engine root.

  • Files in the engine root are orchestrators or thin helpers: the root holds the core orchestrator (engine.py), delegate components (prompt_assembler.py, calibration.py, skill_dispatcher.py, session_persistence.py, learning_bridge.py, tool_dispatch_engine.py), and small focused helpers (message sanitization, cost calculation, budget, audit). A file that grows a sub-domain of its own belongs in a sub-package.

  • 800-line soft limit per file (MANDATORY justification above): any file exceeding 800 lines must document in its module docstring why it cannot be decomposed further. New files must not exceed this limit. engine.py is the sole historical exception and is actively being reduced through delegation.

  • 100-line soft limit per function/method: functions exceeding 100 lines indicate mixed concerns and must be decomposed. The previous _run_agent_loop (512 lines) and _unified_tool_loop_stream (780 lines) were near-identical streaming/non-streaming forks that diverged silently over months — the kind of duplication-then-explosion these limits exist to prevent.

  • AgentEngine composes, it does not accumulate (MANDATORY): AgentEngine orchestrates via six delegate components, each owning a distinct concern:

    • PromptAssembler — prompt construction, context assembly, PCD plan execution
    • CalibrationManager — budget calibration, prefix commitment, threshold tuning
    • SkillDispatcher — skill matching, teach-mode commands, evolution actions
    • SessionPersistence — session load/save, message persistence, memory prefetch
    • LearningBridge — episode persistence, capability observation, coevolution recording
    • ToolDispatchEngine — tool execution, concurrency, guardrails, catalog management

    New engine behavior must be added to an existing component or a new component — never directly to AgentEngine. The only methods that belong on AgentEngine itself are the core agent loop, public entry points (run, run_stream), and thin delegation wrappers for externally-referenced public methods. A method that does not need self._run_agent_loop's iteration state does not belong in the loop body.

  • Back-reference pattern for delegate components (MANDATORY): all delegate components hold a back-reference to the engine (self._engine: AgentEngine) and access engine state via self._engine._xxx, never via captured references passed at construction time. AgentEngine attributes are mutated at runtime by set_* injector methods (e.g., set_evolution_store, set_event_bus); a component holding a captured reference silently uses stale state after re-injection — a bug class that is invisible in tests because tests rarely call set_* after construction.

  • OutputSink is the single streaming contract: the agent loop has ONE code path for both streaming and non-streaming execution, unified through the OutputSink Protocol. BufferSink collects output for run(). StreamSink pushes StreamEvents via asyncio.Queue for run_stream(). Adding a new output mode (SSE, WebSocket, RPC relay) means implementing a new OutputSink, not forking the agent loop. Duplicating the loop previously caused 1,000+ lines of near-identical code that diverged in subtle, untested ways.

  • Recovery sub-package is self-contained: all recovery types (FailureEnvelope, RecoveryDecision, RecoveryAction, RecoveryBudget) and coordination (RecoveryCoordinator) live in engine/recovery/. Recovery strategies implement the RecoveryStrategy Protocol and are registered in engine/recovery/strategies/. Adding a new recovery strategy means adding a file to strategies/ and registering it in strategies/__init__.py — no other engine file should require modification.

  • Context sub-package owns all context shaping: context management — compression, budgeting, disclosure, and focus — lives in engine/context/. The compression pipeline is stage-based (CompressionStage Protocol); adding a compression stage means implementing the Protocol and registering it, not modifying ContextCompressor. Context disclosure reads tool-declared x_leapflow metadata first; substring inference is a deprecated fallback that logs a warning.

  • No cross-sub-package imports below the root: sub-packages (recovery/, context/, tools/, task_planning/, session/) must not import from each other at module level. Cross-cutting coordination flows through the engine root's delegate components or through typed events on EventBus. A direct import between sub-packages creates a coupling that defeats the decomposition. Function-local imports of pure stateless helpers (e.g., exit_code_from used by context/ for evidence rendering) are permitted when the alternative — injecting a one-line function through the engine root — adds indirection without value. Such exceptions must remain function-local (never top-level) and must not create circular dependency chains.

Plugin and Extension Rules

The plugin subsystem is not a feature area — it is how the product is composed. Every capability the agent has, and every capability it can acquire at runtime, enters through it, so a mistake here changes what the agent is able to do rather than how well it does it.

  • leapflow.plugins owns extension mechanics; leapflow.tools owns tool behaviour (MANDATORY): contracts (protocol.py), discovery/DI/assembly (registry.py), fiber lifecycle (scoped_registry.py), isolation (sandbox/), and distribution (marketplace/) live in the plugin package and nowhere else. The dependency direction is one-way and executable: plugin core must never import a tool module, and tool_plugins/ is the single layer allowed to wrap one (tests/test_architecture_contracts.py). The relocated leapflow.tools.{plugins,protocol,plugin_registry,scoped_registry,marketplace,sandbox} paths must stay physically absent — a compatibility shim would split the registry's single source of truth and let two divergent registries coexist.
  • ToolMetadata is the single source of truth for a tool: one declaration produces the provider schema, the handler mapping, and the PCD/capability metadata. x_leapflow is mandatory and must carry category and risk_level; a mutating tool declares mutates_state plus its approval/idempotency metadata; capability tags are declared, not inferred. Never hand-write a second schema, a parallel handler table, or a capability list beside it — the disclosure layer reads declared metadata first, and substring inference is a deprecated fallback that logs a warning.
  • Tool names are one global namespace, arbitrated first-wins and never silently: the incumbent keeps the name and the challenger is recorded as a CapabilityConflict surfaced through plugin_list. Rejection is deliberately non-fatal so one colliding plugin cannot break assembly for every other plugin. Never overwrite a live handler or emit a duplicate schema to claim a name another plugin already owns.
  • Dependencies arrive late, through bind_runtime, and absence degrades rather than crashes: plugins declare names in dependencies and receive matching services injected in provider→consumer topological order, independent of discovery order. A plugin module must not import a runtime service at module level, and must not perform I/O, network calls, or state mutation at import time — every module is required to be importable standalone. A handler whose dependency was never bound returns a structured refusal; it does not raise.
  • Lifecycle is a fiber, and cleanup is a scope (MANDATORY): every plugin instance lives under a PluginFiber (PENDING → LOADING → ACTIVE → UNLOADING → DISPOSED, with a LOADING → FAILED → LOADING retry path) whose EffectScope disposes registered effects LIFO, children before parents, idempotently and exception-safely. Anything a plugin registers process-globally — an interceptor on registry.tool_pipeline, an EventBus subscription, a background task — must be registered as an effect on that scope in the same change, or disable and reload leak it. Illegal transitions raise IllegalStateTransition instead of being silently corrected.
  • Hot-reload is safe because handlers are snapshotted per turn, not because reload is atomic: each turn copies dict(registry.tool_handlers) at its start, so an in-flight turn finishes against the handlers it began with while notify_mutation()'s version bump invalidates the catalog cache for turns that start later. Reload re-imports through importlib.util.spec_from_file_location using the source path recorded as __leapflow_plugin_path__ — never by mutating global sys.path, and never in a way that requires the plugin to be importable from the process's import path. Hot-reload's catalog version bump also invalidates _full_tools_tokens (the cached full-catalog token estimate used by prefix commitment amortization); a registration path that bypasses notify_mutation() leaves the commitment controller operating on a stale cost estimate.
  • Plugin governance is cold-path (MANDATORY): fiber state, trust ledgers, usage statistics, health producers, advisors, proposal queues, and marketplace work must add no per-turn cost to the hot path. Trust is flushed to DuckDB only on level transitions (plus a final atexit flush), and usage samples stay in bounded deques. A governance feature that measurably slows an ordinary turn is a defect in the feature, not a cost to accept.
  • Plugin mutation is uniformly HIGH risk and never permanently granted: any action whose metadata.platform == "plugin_management" is forced to RiskLevel.HIGH with allow_permanent=False in security/risk.py — defense-in-depth that holds even when caller metadata is wrong. Install, reload, rollback, enable, disable, and remove each build an ActionDescriptor and go through ApprovalOrchestrator per invocation. The single exemption is plugin_reload at PRODUCTION trust, which is earned evidence rather than a configured bypass. With no gate installed (in-process CLI binds none), every mutation is denied: code that can rewrite the agent's own composition must never be installable through an unguarded path.
  • Self-evolution is a governed pipeline, not a code-writing shortcut: capability gap → proposal → generate → validate (syntax → structure → import/Protocol conformance) → compatibility assessment → approval → write → sandbox smoke → register at DRAFT → behavior tests → probation → trust accrual → verify, with quarantine and rollback as the failure path. An INCOMPATIBLE verdict is rejected before any file write; a failure at any later stage rolls back the fiber, the sys.modules entry, and the written file. Each next action comes from AdaptiveEvolutionPolicy reading structured requirement, risk, trust, and status — never from natural-language intent — and the autonomy level is configuration, so raising it is a deliberate operator decision rather than a code path.
  • Co-evolution makes every capability transition visible, including retirement: an environment-driven install, reload, disable, rollback, or remove records its causal requirement/evidence, pre- and post-mutation catalog state, artifact identity or digest, compatibility and behavior verdicts, approval, and resulting fiber/trust state through the existing audit, proposal, version, and lifecycle stores. An unselected or superseded proposal must be resolved, expired, or retained with an explicit reason; no proposal, plugin artifact, or registered effect may become an untracked permanent residue.
  • Untrusted code is isolated before it is trusted: requires_sandbox defaults to True; sandboxed plugins run in a subprocess over JSON-RPC with a bounded invoke timeout and receive no host-side runtime dependencies. Marketplace artifacts are verified by SHA-256 checksum and, when trusted pubkeys are configured, by Ed25519 signature over the canonical name|version|entry_point|checksum_sha256 payload. Validation re-runs on the install path even for marketplace code that was already checked.
  • Plugins are process-global; sessions are not: the registry is a daemon-wide singleton, so install, reload, disable, and remove change the capability set for every connected client at its next turn, and trust accrues from all of them. Any change to plugin state must be assessed against the concurrent-TUI contract — per-turn snapshots are the only isolation, and there is deliberately no per-workspace plugin set.
  • Self-capability answers come from the live registry, never from documentation: when LeapFlow reports what it supports — plugins, self-evolution, hot reload, version management — the evidence is plugin_list's live capability_report or an equivalent runtime registry read. If runtime introspection fails, state that the running state could not be verified; never infer a capability from README, design docs, or memory.
  • self_describe is the canonical tool for agent self-cognition: all facets read live runtime state (registry, daemon, engine, build_info) through bind_runtime injected services, never from documentation or static config. A new runtime observable (e.g., a new daemon metric, a new engine state) that the agent should be aware of must be wired into the appropriate self_describe facet in the same change — an observable that exists only in daemon status but not in any tool is invisible to the agent.
  • The plugin contract is published, so it changes with the code: docs/plugins/third_party_plugin_development.md (interfaces, deployment, security model) and docs/plugins/plugin_lifecycle_management.md (lifecycle, governance matrix, enforcement status) are third-party-facing specifications whose tables state what the code does today. A change to a Protocol, a lifecycle transition, an approval rule, a config key, or an injectable dependency name updates them in the same change — and never promotes a roadmap entry to ENFORCED ahead of the wiring.

Extension Ladder

LeapFlow exposes four levels of extension, ordered from lowest barrier to deepest integration. Pick the lowest level that satisfies the requirement — it will ship faster, carry less maintenance cost, and stay compatible across upgrades.

Level 1: Skill (SKILL.md) — Lowest Barrier

  • Pure Markdown file with YAML frontmatter; no code changes required.
  • Add a SKILL.md to the skills directory and the agent discovers it at startup.
  • Declares tool dependencies, trigger phrases, category, and platform constraints.
  • The LLM reads the skill document and autonomously calls existing tools to execute the workflow.
  • Compatible with Hermes skill format (metadata.hermes namespace).
  • Best for: custom workflows, domain knowledge, operational playbooks, guided procedures.

Level 2: MCP Server — Low Barrier

  • External process communicating via Model Context Protocol (JSON-RPC over stdio/SSE).
  • Brings external service capabilities into the agent as discoverable tools.
  • Language-agnostic — any runtime that speaks MCP can serve tools.
  • Best for: external API integrations, third-party service connectors, language-specific tooling.

Level 3: Plugin (Python Module) — Medium Barrier

  • Python module implementing the ToolPlugin Protocol (runtime_checkable).
  • Full access to LeapFlow's runtime: EventBus, memory, storage, settings via bind_runtime.
  • Subject to Progressive Trust lifecycle: DRAFT → CANDIDATE → VERIFIED → PRODUCTION.
  • Sandbox isolation via subprocess JSON-RPC until trust is earned.
  • Best for: deep framework integration, new LLM providers, custom storage backends, platform adapters.

Level 4: Core Tool — High Barrier

  • Direct modification to LeapFlow's core tool system (leapflow/tools/).
  • Requires understanding of internal architecture, review process, and compliance with all rules in this document.
  • Best for: fundamental capabilities that all plugins and skills may depend on.

Summary

Level Mechanism Barrier Use Case Example
1 Skill (SKILL.md) Lowest Workflows, playbooks, domain knowledge Deployment checklist, code-review guide
2 MCP Server Low External services, cross-language tools GitHub API connector, database explorer
3 Plugin (Python) Medium Runtime integration, providers, adapters LLM provider, gateway adapter
4 Core Tool High Foundational agent capabilities File I/O, code search

Community contributions should start at Level 1 (Skill). It requires no code changes, has the fastest feedback loop, and can be shared as a single Markdown file. Escalate to a higher level only when the skill layer cannot express the needed capability.

Path Tree, Configuration, and Secrets Rules

  • Path tree is a product contract: every LeapFlow-managed path must be declared by PathLayout, ProfileLayout, CacheLayout, or a child layout object. Runtime code must consume layout APIs, never assemble managed paths with ad-hoc string joins.
  • Profile is the runtime boundary: profiles/<profile>/ owns profile metadata, config, DBs, memory, skills, gateway state, approval state, audit logs, runtime files, cache roots, and profile-scoped secrets. Cross-profile access requires an explicit layout object.
  • Workspace is context, not ownership: workspace-local files are limited to .leapflow/config.yaml and .leapflow/workspace.yaml; profile data and caches stay under the active profile and are addressed by workspace/session ids.
  • Config is layered, not scattered: durable settings live in config/user.yaml, profiles/<profile>/config/*.yaml, and optional workspace config. LEAPFLOW_* values are process overrides only; env files are not a supported configuration source.
  • leap config is the user-facing control plane: every durable, user-writable setting must be discoverable and mutable through ConfigService, leap config, and TUI /config; do not add one-off setup commands, hidden YAML-only knobs, or new persistent env-first flows.
  • Config catalog is the discovery contract: writable config fields must expose key, effective value, type, scopes, hot-reload semantics, category, value hint, and description. leap config keys stays compact and script-friendly; leap config list and /config list are the human-readable catalog; leap config show <key> and /config show <key> are the single-field detail views.
  • TUI config parity is mandatory: /config mirrors leap config, supports active-session reload when possible, and must remain self-discoverable through slash completion for subcommands, keys, and simple values. Any new config subcommand must update CLI parser, TUI payload/rendering, completion, README, and tests together.
  • Secrets are refs, never durable plaintext: long-lived LLM, VLM, aux-provider, Gateway, and Hub credentials must be stored in the vault and referenced as secret://profile/... or secret://global/.... Config may contain refs, never tokens.
  • Cache declares scope and sensitivity: cache paths must route through CacheLayout/CacheManager with profile, workspace, or session scope. Session visual/video/VLM/signal artifacts are sensitive, non-syncable, and TTL/quota managed by index.
  • Safety follows path semantics: daemon sockets, pid/lock files, runtime state, DuckDB files, vault files, approval grants, audit logs, and memory stores must flow through layout descriptors, path sensitivity, risk, approval, and redaction gates.
  • No legacy aliases: do not reintroduce global .env as persistent config, flat cache roots, profile-root gateway config, inline credential files, .credential_key, or run/ runtime paths.

Sensitive Capability and Approval Rules

Every capability that changes the world outside the current turn — shell execution, sensitive file read/write, config mutation, outbound sends, platform actions, network egress, desktop control, plugin self-modification — reaches the user through one approval chain. Wiring a new one is a fixed sequence, not a design exercise: each rule below is the residue of a defect that shipped with a green suite.

  • One orchestrator is the only entry point (MANDATORY): build an ActionDescriptor — its kind/effect/resource/metadata is the contract — and call ApprovalOrchestrator.evaluate(), which supplies risk classification, policy, existing grants, and the audit record for free. Never hand-roll a confirmation prompt, a per-tool approval allow-list, or a second gate implementation. Never call the orchestrator's legacy check(command): it is the single-argument shell adapter, and a gate exposing only check() must be treated as unusable and denied rather than assumed permissive. config_set copied file_write's four-argument gate call and failed on every invocation with check() takes 2 positional arguments but 4 were given, so no model-driven config change ever reached approval — invisibly, because the test fakes implemented the same wrong signature.
  • Register the gate in both installation sites: gates are process-global and injected twice — in-process through cli/context.py, daemon-side through ApprovalCoordinator.install_gate(). A new sensitive capability must be wired in both, or it behaves differently depending on whether leapd is running. Where a mode legitimately cannot supply a gate (in-process CLI binds no plugin_approval_gate), the tool documents it and fails closed instead of proceeding unguarded.
  • Fail closed on absence and on exception (MANDATORY): no gate installed, no per-turn route, or a gate that raises all mean deny, with a message the model can act on — a broken gate must never become an open door. Keep the except narrow, and never raise from inside one except branch expecting a sibling except to catch it; that is how a TypeError fallback escaped its handler and surfaced raw instead of failing closed.
  • Feasibility precedes consent: an action that cannot succeed must never reach a human. The order is fixed — dedup → payload validation → capability feasibility (CapabilityHealthLedger, blocks_approval) → resource provenance → approval → execute. Missing scopes, degraded capabilities, and admin-required failures return a deterministic repair instruction and hard-stop the turn, with security/permission_failures.py as the single authority both engine and TUI consult. Prompting for consent to a call that will be refused for lack of permission teaches users to click through prompts.
  • Gate the action that actually executes, at every hop: classify once and the transport can still go somewhere else. web_fetch's egress gate only saw the first hop while the transports followed redirects themselves, so any server answering 302 could bounce an approved public URL to loopback or a cloud metadata endpoint. Transports are single-hop and report Location; the caller re-classifies and re-gates each hop, and reachability asks is_global rather than not is_private so unenumerated non-routable ranges fail closed.
  • The decision vocabulary is one enum across the process boundary: ApprovalDecision is the entire vocabulary, so adding or renaming a value means updating the enum, the orchestrator's _choices and scope mapping, ApprovalCoordinator._normalize_decision, the TUI modal, and the RPC together. The daemon normalizer kept a hardcoded allowed-set that predated ALLOW_ALL_SESSION, silently rewrote that choice to deny, and never armed the session bypass it was meant to grant — explicit user consent became a refusal with no error logged anywhere.
  • Scope is the grant's contract, and a bypass is resolved in one predicate: grants persist only at SESSION/PROFILE scope via ApprovalScope, and HIGH/CRITICAL risk sets allow_permanent=False so "always allow" is never offered for actions that change the agent's own composition or reach internal addresses. A bypass — config approval_bypass or a session-level "allow all" — is answered by a single predicate every gate consults, so it cannot mean "approved" at one gate and "still ask" at the next. Hardline blocks sit above all of it and are never bypassable by any grant, scope, or bypass flag.
  • Prompt and audit text is redacted at the descriptor: ApprovalRequest.detail is both rendered to the user and persisted to the JSONL audit log, so secrets are removed when the descriptor is built, not when it is displayed. URL query, fragment, and userinfo are stripped; config values are excluded entirely and only the key is shown. Approval details were carrying API keys and signed tokens out of query strings straight into the audit log.
  • Daemon approvals are turn-routed and terminally resolved: prompts travel on the per-turn approval_route ContextVar (queue, request_id), so a prompt never surfaces in a client that did not cause it. Every pending future needs a terminal path — deny_for_request when the turn ends, deny_for_queue when the stream closes, prune_stale on TTL — because an unresolved future blocks its tool call forever. The ContextVar must be set and reset within the same contextvars.Context; setting it in one per-chunk task and resetting it in another raised "was created in a different Context" and broke streaming for every gated action.
  • A denial is terminal and reaches the model verbatim: ApprovalResult.denial_message states that the user did not consent and that the outcome must not be retried, rephrased, or pursued through another tool. An adapter wrapping a gate must capture that message and return it to the caller; substituting a generic tool error lets the agent reroute around the human's refusal.
  • Approval invariants are tested against the real orchestrator: tests/test_approval_layer.py drives the production ApprovalOrchestrator — per-scope grant reuse, hardline deny without prompting, allow_permanent suppression, decision round-trips through the RPC vocabulary, denial-message pass-through. A hand-written fake is acceptable only as the human surface; a fake that reimplements the orchestrator's own interface will keep agreeing with the caller's mistake, which is precisely how a dead gate stayed green.

Implementation Guidelines

  • Define the Protocol first — the contract is the design
  • Implement against the Protocol, never against another implementation
  • Consider affected user journeys before changing shared flows; do not introduce regressions, broken links, or worse experiences in adjacent paths
  • Keep common paths transparent: long-running work must stream progress, surface recoverable errors clearly, and avoid silent stalls.
  • For context assembly, prefer manifest-driven progressive disclosure over shortcuts or intent-handler sprawl: expose compact capability indexes, selected schemas, and targeted memory only when the current plan needs them.
  • For gateway or IM work, define the signal contract first: event source, normalizer/classifier, trigger policy, session routing, memory/audit path, and outbound action path. Default inbound activation to least privilege (mention_only or equivalent), filter self-generated messages before LLM invocation, and keep cross-chat or proactive sends behind Progressive Trust and ApprovalGate.
  • Avoid rule-based natural-language fitting by default. Do not add keyword/action-verb/alias enumerations, intent-handler taxonomies, or brittle routing rules when LLM-native capability disclosure, manifests, schemas, protocols, or configuration-driven contracts can solve the problem. If a rule-based method is truly unavoidable for a stable protocol boundary, offline fallback, or safety hard gate, explain the necessity, scope, alternatives, and rollback path to a human and obtain explicit second confirmation before implementation.
  • Specifically within engine/: reference resolution, error classification, context disclosure, and focus tracking must not use keyword/substring matching to classify user intent or error type. Reference resolution uses SessionFocusState entity registry lookups; error classification uses data-table or registry-pattern matching (never if-elif chains on message text); context disclosure reads tool-declared x_leapflow metadata first and treats substring inference as a deprecated fallback that logs a warning.
  • Preserve security and audit paths: dangerous actions, file writes, outbound messages, credentials, and path access must flow through the existing policy, approval, redaction, and audit mechanisms.
  • Preserve gateway safety boundaries: inbound credentials stay in CredentialVault; outbound send/write/execute actions go through ApprovalGate; bot self-messages and duplicate events are filtered before routing; platform-specific metadata must remain in metadata escape hatches instead of polluting core message types.
  • Keep App Connector governance thin: platform core should consume normalized contracts and failures, while app-specific auth scopes, CLI/SDK error parsing, vendor recovery steps, and command templates remain in action packs, adapters, or backend-specific helpers. If a new platform requires changing gateway core business rules, first refactor toward a protocol hook or app-side classifier.
  • To add a built-in plugin, add its module path to _BUILTIN_PLUGIN_MODULES in plugins/tool_plugins/__init__.py and expose a module-level plugin instance; that wrapper module is the only place allowed to import the tool implementation it exposes.
  • Tool handlers are async, accept the parameters their schema declares, and return a structured dict. Expected failures come back as {"ok": False, "error": ...}, which records an ordinary trust-affecting failure; raising instead reserves the signal for a genuine internal defect. Permanent trust freezing is a hard_failure recorded through LifecycleGovernor, not something an escaping exception should trigger by accident.
  • Maintain backward-compatible migrations for persistent state, configuration, skills, trajectories, sessions, and profile data.
  • Write unit tests before or alongside the implementation
  • Integrate via EventBus events, not direct function calls between modules
  • Every module must be importable standalone without side effects
  • No placeholder stubs — implement fully or do not add the code
  • ANSI output must check sys.stdout.isatty() before emitting escape codes
  • For error recovery, route all failures through the RecoveryCoordinator — classify into a FailureEnvelope, receive a RecoveryDecision with an explainable reason and strategy_key, then feed the outcome back. Never handle errors with ad-hoc if/break in the loop body.
  • Recovery strategies are standalone Protocol implementations with can_apply() + decide(). Add new strategies by registration, never by modifying the coordinator's decision logic.
  • When automatic recovery exhausts its budget or encounters non-recoverable failures, emit a structured InteractionRequest (typed action, severity, suggested actions, timeout behavior, resumption key) — not raw text appended to conversation. A terminal decision that carries one must surface it: render its title, description, and suggested actions for the user, and pass the structured payload to the client so it can prompt and resume by resumption_key. Dropping it back to decision.reason tells the user a turn stopped without saying what to do.

Review Requirements

  • Deep review for large changes: When a change substantially affects architecture, runtime behavior, user flows, persistence, safety, or multiple modules, perform an additional deep review before considering the work complete.
  • Human confirmation for TUI changes: Any TUI layout or interaction-logic change requires a second human confirmation before it is considered ready.
  • Human confirmation for slash-command paths (MANDATORY, no exceptions): Any change to what a slash command (/...) does requires a second human confirmation before it is considered ready — never ship it on a single pass. This covers the whole surface: registry, router, in-process REPL, daemon REPL, command_execute (including its RPC signature and parameter plumbing), completion, and rendering.
    • Functional changes count even when the command surface is unchanged. The name, arguments, and help text staying identical does NOT waive confirmation. Altering what the command observes, targets, arms, schedules, sends, opens, or persists — or which session/workspace/profile it resolves against — is a functional change and must be confirmed.
    • Equally in scope: dispatch and routing, prompts and confirmations, emitted output, browser/dashboard launches, background work the command triggers (watches, schedules, re-entries), and error/recovery messaging.
    • Passing tests and a clean lint run are NOT a substitute for confirmation. Slash commands are the primary user-facing control plane; correctness of the visible behavior is only established by a human check.
    • State the pending confirmation explicitly in the handoff, and name the behavior a human should exercise to verify it.
  • Human confirmation for approval-path changes: any change to what reaches the approval chain — a newly gated capability, a new ActionKind, an ApprovalDecision/scope/bypass change, or gate registration — requires exercising the real prompt by hand in both in-process and daemon mode before it is considered ready. Gates are process-global and injected twice, so a green suite proves at most that one of the two wirings works; every approval defect recorded in this document passed its tests.
  • Deep review for plugin composition changes: a change to a Protocol signature, discovery, fiber lifecycle, trust thresholds, sandbox policy, or marketplace verification alters what the agent can load and execute at runtime. Exercise it against a real registry (register → publish → reload → dispose) rather than a fake, and re-check the concurrent-client and cold-path implications before considering it complete.
  • Design goal check: Verify that the implementation actually achieves the intended design goal and is not just a local patch.
  • Optimality check: Evaluate whether the solution is the simplest robust design, avoids unnecessary abstractions, and fits the existing architecture.
  • Regression impact check: Inspect affected modules and user journeys for logic bugs, degraded UX, broken compatibility, slower feedback, weaker diagnostics, or worse failure recovery.
  • SOLID and extensibility check: Look for responsibility leaks, tight coupling, hardcoded paths/thresholds/rules, magic strings, and choices that reduce generalization or future extension.
  • Fix what the review finds: If the review identifies correctness, design, UX, SOLID, hardcoding, or extensibility issues, fix and simplify them directly rather than only reporting them.
  • Anti-hardcoding audit for engine/: All regex patterns, keyword lists, and if-elif classification chains in engine/ must be audited against the rule-avoidance principle (Implementation Guidelines). Permitted only when the pattern falls into one of the three explicit exemptions (stable protocol boundary, offline fallback, safety hard gate). Non-exempt hard rules must be refactored to Protocol/registry/config-driven implementations. When reviewing engine changes, verify that new code does not introduce keyword enumerations, substring routing, or magic-number thresholds without a configuration escape hatch.

Testing Philosophy

The suite has two layers with different jobs, and keeping the boundary sharp is what stops each from doing the other's work badly.

Mock layer (tests/*.py, marked unit/component) — broad and fast. It owns pure algorithms, state-machine branches, error-classification tables, rare and malformed inputs, and single-module invariants. Branch combinations can only be enumerated here, and only here is the feedback measured in milliseconds.

Real layer (tests/journeys/, marked e2e) — six coarse journeys driving a real leapd subprocess over RPC, with the LLM boundary served by a local cassette proxy. It owns what a mock structurally cannot observe: cross-module wiring, process boundaries, session identity, workspace binding, real persistence, and pushed runtime metadata. Every incident recorded in this document shipped with a green mock suite.

Three rules keep the split honest:

  1. Anything provable with one mock-layer assertion must not enter the real layer.
  2. One journey is one test case. Express variation as ordered phases inside a single session; never parameterize a journey.
  3. The real layer has a hard case budget (tests/regression/test_suite_budget.py). When it is reached, merge a journey — do not raise the ceiling. The budget is what lets the real layer run on every push, and a suite that can be skipped will be skipped.

The two layers are joined by recorded traffic. tests/_fixtures/cassettes/ holds the deterministic inputs the offline lanes replay (rebuilt with make seed-cassettes); tests/_fixtures/recordings/ holds real provider traffic captured by make record-traffic. make sync-fixtures distils both into tests/_fixtures/llm_responses/response_shapes.json, which the mock layer asserts against — so a provider dropping a field turns the build red instead of passing forever against a body written from memory. Recording never writes to the replay store: a multi-turn agent conversation cannot be replayed from a recording, because each turn's prompt embeds the exact round-by-round history of the turns before it.

Each journey also declares two cost ceilings, both enforced at the proxy and reported by finish(). max_llm_calls is the convergence guard: a turn that stops converging is cut off instead of running to the engine's iteration cap. max_llm_tokens is the cost guard, and it catches what call count cannot — prompt growth raises the bill without adding a single round. Raising a ceiling is not the fix for hitting one.

Which tests run. The offline lanes never select: the mock layer runs in full, and the real journeys run in full on every push. Selection would save at most a few seconds, because the always-on tiers set the floor, and a suite that can be skipped will be skipped. The live lane is the exception — there a journey costs real tokens, so it selects: each journey declares SUBJECT_PATHS (the sources it exercises) and LIVE_SIGNAL (whether a real provider adds signal), and tools/impact.py --live-journeys picks from those. Declaring them inside the journey keeps that knowledge next to the assertions it describes. tools/impact.py can also scope the mock layer (make test-impact) for local work on a large change; it is not wired into CI.

  • Unit tests must be hermetic: no network, no LLM calls
  • Journeys must not mock anything: a journey that reaches for unittest.mock has become an expensive unit test
  • Provider bodies are recorded, not written: use a cassette or a derived fixture; a hand-authored body keeps passing after the provider changes
  • Faked construction needs real-instance cover: object.__new__ plus private-attribute assignment is acceptable only for ordering contracts, and only when the same file also builds the class properly and drives the production path
  • py_compile all modified files: syntax errors caught before test run
  • Import chain verification: every new module must be importable standalone
  • Existing tests must not regress: all tests must pass after every change
  • User-facing flows must not regress: preserve or improve usability, feedback clarity, and failure recovery for impacted paths
  • Verification sequence: compile → import → mock layer → real journeys
  • Behavior contracts over snapshots: assert invariants, not frozen values
  • Mock at boundaries only: mock external I/O (network, disk), never internal logic
  • A test may not fabricate the wiring it claims to cover: building an object with object.__new__ and assigning the private attributes the code reads cannot detect a wrong attribute name — the test simply agrees with the typo. Calibration tests did exactly that and stayed green while every real turn raised AttributeError. Any test whose stated purpose is wiring must construct the real object and drive the production path.
  • Multi-client behavior needs multi-client tests: session routing, status(), stream metadata, and client-lease changes require two sessions in two workspaces asserting that neither sees the other's identity, usage, or turn state. Single-session tests cannot observe cross-client leakage, which is why a leak shipped with a green suite.
  • Co-evolution experiments need longitudinal counterfactual evidence: exercise an unchanged baseline, an irrelevant delta that is correctly rejected, a task-relevant delta already satisfied by the live catalog, and an unmet requirement that traverses the real governed lifecycle. Use isolated profiles and real integration seams; deterministic reruns are not independent environmental units. Freeze the experiment configuration and corpus, preserve an immutable evidence bundle with a digest and known threats, and report the evidence level, unit count, no-op/rejection outcomes, mutations, and final steady state.
  • Change-scoped validation: Run the most specific relevant tests first, then broaden only as needed: CLI/TUI changes require CLI/TUI tests; leapd changes require daemon RPC/lifecycle tests; storage or memory changes require persistence tests; gateway, IM, event-source, or approval changes require connector lifecycle, event normalization, routing, idempotency, self-message filtering, security/approval, and failure-recovery tests; plugin contract, registry, lifecycle, sandbox, marketplace, or trust changes require the plugin reload, scoped-registry, fiber/effect-scope, sandbox, marketplace-signing, trust-learning, and architecture-contract tests; skills, learning, perception, and copilot changes require their lifecycle or pipeline tests.
  • Recovery strategy isolation: Each RecoveryStrategy must be testable in isolation — verify can_apply predicates, decide outputs, and side-effect-state gating independently of the coordinator and other strategies.
  • Budget boundary tests: Verify that recovery budgets exhaust correctly (per-category, per-turn, deadline), that exhaustion produces a deterministic halt decision, and that cost accounting is exact.

What to Avoid

  • Over-engineering: if you need 3+ files for a simple feature, rethink
  • Premature optimization: measure first, optimize only bottlenecks
  • God objects: no class should exceed 500 lines; approaching this limit requires checking whether policy, state, rendering, protocol, storage, or adapter responsibilities should be split.
  • Magic strings: use constants or enums
  • Blocking the event loop: all IO must be async or run_in_executor
  • Hardcoded paths, URLs, thresholds without config escape hatch
  • Chinese comments in source code (English only)
  • Speculative infrastructure: no hooks or extension points without a concrete consumer
  • Mixing long-lived event observation into short-lived action execution; CliBackend is for bounded commands, while streaming CLI consumers, webhooks, polling, and WebSocket subscriptions belong behind BackendEventSource-style contracts.
  • Activating IM agents on all inbound messages by default, skipping self-message filtering, or allowing cross-chat/proactive sends before Progressive Trust and approval policies explicitly permit them.
  • Putting third-party app business code into platform core: do not add vendor scopes, lark-cli/SDK JSON parsing, auth command construction, console-specific recovery instructions, or resource-specific branching to gateway-wide modules such as action registries, capability ledgers, approval gates, or engine recovery paths.
  • Shortcut-style natural-language fitting and large intent-handler taxonomies; use stable runtime gates plus capability manifests instead. Rule-based keyword/action-verb/alias matching is prohibited by default and requires explicit human second confirmation before implementation when unavoidable.
  • Scattered if/break/continue recovery decisions inside the agent loop body; all error handling enters the RecoveryCoordinator as a FailureEnvelope
  • Magic retry counts or unbounded retry loops without budget constraints and deadline enforcement
  • Feeding unstructured error text back to the LLM without classification, recoverability assessment, or side-effect awareness
  • Multiple parallel error-handling paths for the same failure domain (LLM errors in one handler, tool errors in another, security errors in a third); use a unified classification and coordination pipeline
  • Widening a provider call's try block to cover bookkeeping, or classifying a Python-level defect by matching its message text against provider conditions
  • Answering "the caller's session" with "the most recently active session", or letting a client adopt a session id the daemon happened to report
  • Hand-rolled confirmation prompts, per-tool approval allow-lists, or any second gate beside ApprovalOrchestrator; a sensitive capability declares an ActionDescriptor and evaluates it
  • Treating a missing, unbound, or raising approval gate as permission to proceed
  • Putting a secret, token, or config value into ApprovalRequest.detail — it is rendered to the user and persisted to the audit log
  • Extending ApprovalDecision without updating the daemon normalizer, TUI modal, and RPC in the same change
  • Importing a tool implementation from plugin core, or re-creating a leapflow.tools shim for a relocated plugin-subsystem module
  • Module-level I/O, network calls, or runtime-service imports in a plugin module; dependencies arrive through bind_runtime
  • Registering a process-global interceptor, subscription, or background task without a matching cleanup effect on the plugin's EffectScope
  • Reloading a plugin by injecting into sys.path instead of a file-backed import spec, or overwriting a live handler to claim a tool name another plugin owns
  • Adding per-turn cost for plugin governance (trust, stats, health, advisor, proposals) — governance is cold-path
  • Treating an environment delta, a model-authored evolution hypothesis, or a successful experiment run as authorization to mutate the framework; each must still pass evidence validation, live resolution, policy, approval, sandboxing, lifecycle, and trust gates
  • Using an experiment harness to register plugins directly, forge observations/proposals, bypass approval, or write synthetic evidence into production state without an explicit experiment-profile boundary
  • Answering a question about LeapFlow's own capabilities from documentation or memory instead of a live registry read
  • Bare except: clauses — always specify the exception type
  • # TODO: implement stubs — implement or don't commit

Naming Conventions

  • Files: snake_case.py — noun for types (signal_event.py), verb for actions (compress_context.py)
  • Classes: PascalCase — Protocols suffixed with purpose (SignalSource, SkillStore)
  • Functions/Methods: snake_case — verb-first (filter_noise, retrieve_context)
  • Constants: UPPER_SNAKE_CASE — grouped in module-level or dedicated constants.py
  • Private: single underscore prefix (_internal_method) — never double underscore
  • Modules/Packages: short, singular nouns (perception, causal, memory)
  • Events: past-tense domain verbs (SignalReceived, SkillVerified, ContextCompressed)
  • Config keys: dot.separated.lowercase in env/settings (copilot.idle_threshold_ms)