This document is the LeapFlow engineering collaboration contract. It is not only a style guide: it defines the design, runtime, UX, safety, and verification rules that every implementation change must follow.
-
Signal-Driven Intelligence — All agent intelligence derives from observing real-world signals, not from hardcoded rules. If a behavior cannot be learned from signals, it is not in scope.
-
Context Pipeline as Core — Signal → Filter (SNR) → Compress (intent-preserving) → Store (multi-layer) → Retrieve (goal-dependent) → Decide. Every feature and every external signal source, including IM collaboration events, must map to this pipeline before it can drive action.
-
Everything Is a Plugin — Capability is composed, not built in. Tools, LLM backends, platform adapters, signal sources, and vision processors all arrive behind a
runtime_checkableProtocol and are discovered, injected, and disposed by the same machinery. A capability that can only exist by editing core is a design failure; the answer is a new Protocol, not a special case. -
Progressive Trust — Never auto-execute on first encounter. Autonomy is earned through repeated observed success and lost the same way: DRAFT → CANDIDATE → VERIFIED → PRODUCTION on consecutive successes, demotion on consecutive failures, permanent freeze on an internal defect. Trust is per plugin, persisted, and the only legitimate source of an approval exemption.
-
Occam's Razor — The simplest correct solution wins. Reject complexity that doesn't directly serve user value. Every abstraction must pay for itself.
-
LLM-Native Design — Design for LLM reasoning first. Protocols over classes. Declarative over imperative. Context over configuration.
-
User-Centric Reliability — User experience is part of correctness. Every change must keep common paths easy, predictable, recoverable, and must not degrade adjacent workflows.
-
Environment–Harness Co-evolution — An environmental change is an observable hypothesis, never authority to mutate the framework. Only a typed, task-relevant delta that survives evidence validation and live capability resolution may produce a governed capability requirement; policy, approval, trust, and lifecycle control every subsequent plugin addition, change, disablement, or removal. The complete causal record — including rejected and no-op branches — must remain available so evolution is explainable, reproducible, and reversible.
- SOLID principles are non-negotiable; implementations must be cohesive, well-factored, and easy to reason about
- Occam's Razor: maximize elegance and efficiency, reject unnecessary complexity
- Design for generalization and universality; prefer reusable domain concepts over one-off special cases
- Easy to extend, avoid hardcoding and hard rules
- Industrial-grade robustness: every external call has timeout, retry, and fallback
- Structured error propagation: failures must flow as typed envelopes (
FailureEnvelope) with source, category, recoverability, and side-effect state — never as ad-hoc dicts, bare strings, or unclassified exceptions - User experience is a first-class quality bar: optimize for clarity, ease of use, fast feedback, and graceful recovery
- All comments and docstrings in English
- Type annotations on all public APIs
- No bare except — always specify exception types
-
System Boundary Awareness: LeapFlow is a multi-entry, multi-module runtime. Changes must account for the affected path across CLI/TUI, leapd, engine, plugins, skills/tools, LLM, storage, memory, gateway, hub, and platform adapters.
-
TUI as the Primary User Entry: The interactive TUI is the default product surface. Preserve streaming feedback, command queue behavior, approval prompts, status bar accuracy, long-input robustness, history, and session continuity.
-
Concurrent TUI Instances Are a Supported Scenario (MANDATORY): several TUIs in different workspaces, sharing one leapd and one profile, is a normal way to use LeapFlow — not an edge case. Each instance must remain fully usable and must see only its own session, conversation, context usage, and turn state. A change to session routing,
status(), stream metadata, the client lease, or anything the status bar renders is not verified until it has been exercised with two instances in two workspaces at the same time. One instance degrading another is a release blocker, not a limitation to document. -
TUI Command Clarity: Global task-control commands stay short and unambiguous (
/cancel,/skip,/pause,/resume,/queue,/drop); teach-mode controls must use the/teach ...namespace and should not keep bare compatibility aliases during early iteration. -
TUI Prompt Ownership: Input prompt and placeholder rendering must have a single owner. Avoid duplicate prompt sources; placeholder text stays visually subordinate, offset after the prompt, and disappears as soon as the user types.
-
leapd Runtime Consistency: Daemon-backed behavior must preserve lifecycle correctness: start, stop, restart, status, RPC streaming, cancellation, pending approvals, runtime config reload, multi-client state, and version consistency.
-
Progressive Context Disclosure (PCD): Keep one unified execution loop, but never default every turn to full disclosure. Each LLM call must use the smallest sufficient PromptAssemblyPlan for tools, memory, history, reasoning, streaming, and risk; upgrade progressively only when observable signals require it.
Prefix cache stability is part of PCD correctness. The prompt prefix — system prompt static body, tool catalog text, and frozen append-only summary segments — must be byte-stable across consecutive turns within a committed task. Volatile per-turn content (memory retrieval, distilled knowledge, semantic focus, session summary) must be injected as a separate
_volatile_context: Truesystem message so the cache optimizer excludes it from the stable prefix. A change that moves per-turn content into the stable prefix, or that reorders stable-prefix messages after cache markers are applied, silently degrades cache hit rate and is a correctness defect, not a performance trade-off. -
System Prompt Template Anchors Are Structural Contracts (MANDATORY):
AnthropicCacheStrategysplits the system prompt into a cached static head and an uncached dynamic tail using deterministic anchors (_STATIC_TERMINAL_ANCHORand_KNOWN_STATIC_HEADERSinprompt_cache.py). An edit toUNIFIED_SYSTEM_TEMPLATEthat moves, renames, or removes a known static header, or relocates the terminal anchor line, must update the anchor set inprompt_cache.pyin the same change. Adding a new always-present section to the static body requires adding its header to_KNOWN_STATIC_HEADERS. -
Prefix Commitment Is Monotonic; Enforcement Is Revocable:
CommitmentStatustransitions only fromUNCOMMITTEDtoCOMMITTEDwithin a task. The enforcement snapshot (frozen disclosure level, tool set, system prompt hash) is broken by structural disruptions (posture change, tool error, slash command, transform retry) and re-established when the prefix stabilizes. Code that introduces a new structural disruption must call_maybe_break_commitmentwith the appropriate flag; code that changes commitment conditions (difficulty threshold, posture gates, amortization model) must updatePrefixCommitmentConfigand verifytest_prefix_commitment_enforcement.py. -
Compression Must Preserve Append-Only Frozen Segments:
SummarizeStagemarks its output with_compressed_summary: True; the append-only mode freezes earlier summaries into the head.DropStagepreserves contiguous frozen segments after the system head. A new compression stage must not mutate, reorder, or remove messages carrying_compressed_summary: True— these segments are part of the byte-stable prefix. -
Cache Hit Rate Is a Dual-Caliber Observable:
TurnUsageTrackerreports per-turn, token-weighted cumulative, and steady-state (cold-start-excluded) cache hit rates. Provider adapters must propagatecached_tokens(OpenAI/DeepSeek) or bothcache_read_input_tokensandcache_creation_input_tokens(Anthropic) in the usage dict. A new provider that omits these keys renders the cache metric blind. -
Session Resume Cache Policy Must Be End-to-End:
session_resume_cache_policy: "cache_priority"persists the system prompt and tool schema at turn end and freezes them on resume so the first turn hits the provider prefix cache. Adding a new field to system prompt assembly that is not captured by_last_system_promptwill silently diverge the resumed prefix, causing a miss on the most expensive turn. -
Task Environment Is a First-Class Signal: Environment adaptation identifies task environments through declared descriptors and typed structural or affordance deltas. Host capabilities and application affordances are separate concepts; an opaque "changed" hash or a host fingerprint alone cannot establish task compatibility, loss, or a need to evolve.
-
Resolution Before Acquisition: A classified environment signal carries redacted provenance and becomes a capability requirement only when it is relevant to the active task. Resolve that requirement against the live catalog before opening a proposal: a satisfiable requirement is recorded as a no-op, while an unmet task-critical requirement follows the typed capability-unavailable/recovery path. No environment change may create a duplicate capability merely because it was observed.
-
Experiment Control Plane Is Not the Subject: Environment sources, fault injection, and harness adapters are discoverable plugins that drive only real production seams. They must not register tools directly, hand-write observation or proposal records, impersonate approval, or introduce a parallel mutation path. Synthetic signals are explicitly enabled only in an experiment profile; production defaults remain unchanged until an operator opts in through the normal configuration surface.
-
Evolution Causality Is Durable and Queryable: For each observed environmental delta, preserve a redacted causal chain linking source and before/after descriptors, evidence and requirement identifiers, live-resolution result, policy and approval decisions, affected plugin/version/fiber, validation and trust outcomes, and any rollback or retirement. A rejected gate, duplicate resolution, failed validation, or no-op is a first-class result, not missing telemetry; it proves why the framework did not change.
-
Gateway as Signal Boundary: External IM/platform integrations are not just messaging features; they extend LeapFlow's Observe/Orient boundary into collaboration environments. Inbound platform events must enter as structured signals (
BackendEvent→ normalized domain event/message), pass SNR filtering and privacy/safety gates, then feed memory, decision, and action paths according to their classification. -
Transport-Lifecycle Separation: Short-lived actions (
ExecutionBackend/CliBackend) and long-lived observations (BackendEventSource) are separate responsibilities. Do not implement streaming subscribers, webhooks, polling loops, or CLI NDJSON consumers inside one-shot action execution code. -
Platform-Neutral Gateway Core: Gateway core owns protocols, lifecycle, routing, session isolation, approval, audit, and memory integration. Platform adapters own authentication, send semantics, event-source configuration, and schema normalization. Core modules must not import platform SDKs directly. Per-vendor code — including credential validators — lives in a platform sub-package (
adapters/,normalizers/,action_packs/,validators/<platform>.py), never in a core module; core keeps only the neutral registry and contracts. -
Platform vs App Business Boundary: Platform layers may define stable contracts and governance primitives (
ActionSpec,ActionFailure,ActionAuthSpec,CapabilityHealthLedger, approval/feasibility gates, audit, and metadata propagation). Third-party app or vendor specifics — SDK/CLI wire formats, scope names, auth commands, console URLs, error JSON shapes, resource naming, and recovery playbooks — must live in that app's action pack, adapter, backend, or normalizer, never in gateway core. -
Dependency Inversion: Core logic depends on Protocol abstractions, never on concrete implementations
-
Protocol over ABC: Use
typing.Protocolwithruntime_checkablefor all extension points -
Event-Driven Communication: Modules interact through typed events on EventBus, not direct imports
-
Session Engine is the Only Reporting Source (MANDATORY): conversation state lives on the per-session engines built by
SessionRegistry.ctx.engineis only the template they are cloned from and never accumulates turns, context, or history. Any code reporting runtime state — stream chunk metadata,status(), session analysis, dashboards, status bar values — must resolve the engine through the single entry point (RuntimeLeapService._active_engine()/SessionCoordinator.resolve_session_engine()), nevergetattr(ctx, "engine"). Reading the template silently yields zeros, which is invisible in review and has surfaced repeatedly as an empty LeapBoard and a status bar frozen at0/<limit>. When a value is produced by a specific engine (e.g. a stream event), pass that engine explicitly instead of re-resolving, so concurrent sessions cannot be cross-reported. -
Client-Visible Runtime State Must Be Pushed, Not Inferred: a daemon-mode TUI is a separate process; it seeds model, context length, and usage at startup and can only learn about later changes from metadata the daemon returns. Any runtime value the status bar renders must travel on ordinary status/stream metadata (and on the mutation payload for command RPCs) — never rely on a one-off change notification, which change-detection can legitimately skip.
-
Session Identity Belongs to the Client That Created It (MANDATORY): a
session_idis the client's own identity, not shared daemon state. The daemon must never report a session a caller did not name, and a client must never adopt a session id it did not ask for. Concretely: every RPC that returns session-scoped state (status(), stream chunk metadata, history, analysis) takes the caller'ssession_idand resolves that session; when the caller names none, the reply carries no session identity and no per-session figures rather than a substitute. A client accepts a reported session id only when it matches its own or it has none yet (first assignment, or an explicit--resume). Violating this is not cosmetic: a second TUI adopted the first's session, sent it with its own workspace, and was rejected on every turn — with advice ("start a fresh session") that could not work, because a fresh client re-adopted the same id on its first status poll. -
Cross-Session Fallbacks Must Be Named For What They Do: a resolver that answers "whichever session was most recently active" ignores workspace and client identity, so it is valid only for genuinely aggregate views (a dashboard summarizing all activity). Such helpers must say so in their name (e.g.
most_recent_any_client()), and no code path that describes one caller may use them. A friendly name like "the current session" invites exactly the misuse that leaked one client's identity to another. -
Workspace Binding Is Part of Session Identity: a session is bound to the workspace of its first request, and reuse from another workspace is refused. Because correct clients cannot trigger this, a mismatch means either an explicit
--resumeinto another workspace (the only legitimate cause, and what the message must name) or a defect — never something to tell the user to work around. -
Terminal Output Must Wrap at the Console Layer:
LeapConsoleowns wrapping (soft_wrap=False), because prompt_toolkit's renderer clips at the window edge rather than reflowing. Never enablesoft_wrapon the shared console or hand it pre-formatted long lines; long answers silently lose their tail. A standaloneConsolefor fixed-width art (e.g. the banner) may opt out, but must set an explicitwidth. -
Immutable Domain Types: Use
@dataclass(frozen=True)orNamedTuplefor domain objects — but never for an exception type. CPython stores the traceback on the instance and every Python-level re-raise assigns__traceback__through__setattr__(contextlib.__exit__, asyncio, and pytest all do), so a frozen exception raisesFrozenInstanceErrorthere and replaces the real failure with that noise: a driver answering "Unknown tool" was reported for months as "cannot assign to field 'traceback'". -
Config-Driven Behavior: Thresholds, intervals, feature flags, model budgets, platform capabilities, hub backends, gateway manifests, and paths must be configurable through Settings/env/config layers.
-
Graceful Degradation: Every optional component (LLM, Hub) can be absent without crash
-
Single Source of Truth: DuckDB for persistence, EventBus for communication, Settings for configuration
-
Inbound Signal Classification: Platform events must be classified before they activate the agent. Message/callback events may enter Decide; signal/lifecycle events should be stored or routed without triggering LLM by default; ignored events must be explicit (e.g. self-message, duplicate, blocked scope).
-
Single Recovery Decision Point: All agent loop errors (LLM, tool, system, security) enter one
RecoveryCoordinator. No parallel decision paths, no scattered if/break logic. The pipeline is always:FailureEnvelope→RecoveryDecision→StrategyOutcomefeedback. -
A Local Defect Is Never a Provider Failure (MANDATORY): an
exceptaround a provider call must wrap only that call. Post-response bookkeeping — usage recording, calibration, capability learning — belongs outside it, in a helper that contains its own failures, because telemetry must never fail a turn. Exceptions that mean "LeapFlow has a bug" (AttributeError,TypeError,NameError,KeyError,IndexError,ImportError,AssertionError,NotImplementedError) are classified by type into the non-recoverableinternal_defectcategory before the provider taxonomy is consulted. The provider classifier matches on message text, so one mistyped attribute whose name contained "context" was read as a context overflow and driven through three compressions, a provider failover, and a credential rotation before halting — every turn, for every user, with the suite green. -
Every Terminal Decision Must Be Actionable: a halt is the last thing a stopped turn can say, so it always carries an
InteractionRequestnaming the failure and the next step. The rawreasonis written for the audit log; surfacing it alone produces internal jargon as the user's entire answer ("No applicable recovery strategy found"). -
Recovery Must Leave Evidence: every recovery entry point logs the exception with
exc_info, and the audit sink is constructed with the profile layout's audit path. An in-memory sink and an unlogged branch made a reproducible outage undiagnosable after the fact — the incident had to be reconstructed from token counts in an unrelated log line. -
Side-Effect Gating: Recovery is gated by
SideEffectStateat two levels. Within a tool batch, a failed side-effecting call stops the remaining calls in that batch, decided by the declaredexecution_policyrather than any tool-name list. WithinRecoveryCoordinator, any state other thanNONEblocks replaying actions (retry, transform-and-retry, failover) and yields a checkpointed halt carrying anInteractionRequest; only user-mediated or checkpoint-based resumption is permitted.UNKNOWNblocks likeCOMMITTEDandPARTIALdo: it is the classifier's fallback and the state assigned toexternal_side_effect(outbound sends, external API calls), so exempting it would leave the highest-risk case ungated. -
Uncertain Effects Are Reported, Not Retried Blindly: a failed call whose effect may already have landed (
external_side_effect,mutating_once) must carry that verdict in its result so the next turn verifies before repeating it. An error is not proof that nothing happened. Idempotent mutations are exempt — re-applying them converges, so flagging them would only stall safe retries. -
Budget-Constrained Recovery: Turn-level deadlines, per-category limits, and a global recovery budget prevent infinite retry loops. Every recovery action has an explicit cost; exhaustion triggers a clean halt or user escalation.
-
Recovery Strategy as Protocol: Recovery strategies implement a
RecoveryStrategyProtocol (can_apply+decide), registered by priority, composable, and extensible without modifying the coordinator. -
Tool Output Must Survive Rendering Intact (MANDATORY): text produced by tool execution — especially error tracebacks, JSON payloads, and structured diagnostics — must reach the user without silent corruption. Angle-bracketed identifiers (
<string>,<module>,<stdin>) in Python tracebacks, and any content that resembles HTML tags, must be escaped or code-fenced before passing through Markdown renderers. A silently stripped traceback is worse than no traceback — it misdirects investigation. The_sanitize_final_responsepipeline owns this guarantee for the TUI path.
The engine module (leapflow/engine/) is the agent's core execution surface — prompt assembly, tool dispatch, recovery, context management, and session orchestration all live here. A structural mistake in engine/ degrades every turn for every user, so the rules below encode the decomposition invariants established during the refactoring that broke a 3,000+ line monolith into composable sub-packages.
-
Sub-package structure is the module's type system:
engine/is organized into five sub-packages —recovery/(error classification, recovery coordination, strategies),context/(compression, control, disclosure, focus),task_planning/(task graph, planner, scheduler),tools/(execution policy, concurrency, guardrails, action executor),session/(session controller, session factory). New functionality must be placed in the sub-package whose boundary it fits. If no existing sub-package fits, propose and justify a new one with clear responsibilities before implementation — do not add domain-specific logic to the engine root. -
Files in the engine root are orchestrators or thin helpers: the root holds the core orchestrator (
engine.py), delegate components (prompt_assembler.py,calibration.py,skill_dispatcher.py,session_persistence.py,learning_bridge.py,tool_dispatch_engine.py), and small focused helpers (message sanitization, cost calculation, budget, audit). A file that grows a sub-domain of its own belongs in a sub-package. -
800-line soft limit per file (MANDATORY justification above): any file exceeding 800 lines must document in its module docstring why it cannot be decomposed further. New files must not exceed this limit.
engine.pyis the sole historical exception and is actively being reduced through delegation. -
100-line soft limit per function/method: functions exceeding 100 lines indicate mixed concerns and must be decomposed. The previous
_run_agent_loop(512 lines) and_unified_tool_loop_stream(780 lines) were near-identical streaming/non-streaming forks that diverged silently over months — the kind of duplication-then-explosion these limits exist to prevent. -
AgentEngine composes, it does not accumulate (MANDATORY):
AgentEngineorchestrates via six delegate components, each owning a distinct concern:PromptAssembler— prompt construction, context assembly, PCD plan executionCalibrationManager— budget calibration, prefix commitment, threshold tuningSkillDispatcher— skill matching, teach-mode commands, evolution actionsSessionPersistence— session load/save, message persistence, memory prefetchLearningBridge— episode persistence, capability observation, coevolution recordingToolDispatchEngine— tool execution, concurrency, guardrails, catalog management
New engine behavior must be added to an existing component or a new component — never directly to
AgentEngine. The only methods that belong onAgentEngineitself are the core agent loop, public entry points (run,run_stream), and thin delegation wrappers for externally-referenced public methods. A method that does not needself._run_agent_loop's iteration state does not belong in the loop body. -
Back-reference pattern for delegate components (MANDATORY): all delegate components hold a back-reference to the engine (
self._engine: AgentEngine) and access engine state viaself._engine._xxx, never via captured references passed at construction time.AgentEngineattributes are mutated at runtime byset_*injector methods (e.g.,set_evolution_store,set_event_bus); a component holding a captured reference silently uses stale state after re-injection — a bug class that is invisible in tests because tests rarely callset_*after construction. -
OutputSink is the single streaming contract: the agent loop has ONE code path for both streaming and non-streaming execution, unified through the
OutputSinkProtocol.BufferSinkcollects output forrun().StreamSinkpushesStreamEvents viaasyncio.Queueforrun_stream(). Adding a new output mode (SSE, WebSocket, RPC relay) means implementing a newOutputSink, not forking the agent loop. Duplicating the loop previously caused 1,000+ lines of near-identical code that diverged in subtle, untested ways. -
Recovery sub-package is self-contained: all recovery types (
FailureEnvelope,RecoveryDecision,RecoveryAction,RecoveryBudget) and coordination (RecoveryCoordinator) live inengine/recovery/. Recovery strategies implement theRecoveryStrategyProtocol and are registered inengine/recovery/strategies/. Adding a new recovery strategy means adding a file tostrategies/and registering it instrategies/__init__.py— no other engine file should require modification. -
Context sub-package owns all context shaping: context management — compression, budgeting, disclosure, and focus — lives in
engine/context/. The compression pipeline is stage-based (CompressionStageProtocol); adding a compression stage means implementing the Protocol and registering it, not modifyingContextCompressor. Context disclosure reads tool-declaredx_leapflowmetadata first; substring inference is a deprecated fallback that logs a warning. -
No cross-sub-package imports below the root: sub-packages (
recovery/,context/,tools/,task_planning/,session/) must not import from each other at module level. Cross-cutting coordination flows through the engine root's delegate components or through typed events on EventBus. A direct import between sub-packages creates a coupling that defeats the decomposition. Function-local imports of pure stateless helpers (e.g.,exit_code_fromused by context/ for evidence rendering) are permitted when the alternative — injecting a one-line function through the engine root — adds indirection without value. Such exceptions must remain function-local (never top-level) and must not create circular dependency chains.
The plugin subsystem is not a feature area — it is how the product is composed. Every capability the agent has, and every capability it can acquire at runtime, enters through it, so a mistake here changes what the agent is able to do rather than how well it does it.
leapflow.pluginsowns extension mechanics;leapflow.toolsowns tool behaviour (MANDATORY): contracts (protocol.py), discovery/DI/assembly (registry.py), fiber lifecycle (scoped_registry.py), isolation (sandbox/), and distribution (marketplace/) live in the plugin package and nowhere else. The dependency direction is one-way and executable: plugin core must never import a tool module, andtool_plugins/is the single layer allowed to wrap one (tests/test_architecture_contracts.py). The relocatedleapflow.tools.{plugins,protocol,plugin_registry,scoped_registry,marketplace,sandbox}paths must stay physically absent — a compatibility shim would split the registry's single source of truth and let two divergent registries coexist.ToolMetadatais the single source of truth for a tool: one declaration produces the provider schema, the handler mapping, and the PCD/capability metadata.x_leapflowis mandatory and must carrycategoryandrisk_level; a mutating tool declaresmutates_stateplus its approval/idempotency metadata; capability tags are declared, not inferred. Never hand-write a second schema, a parallel handler table, or a capability list beside it — the disclosure layer reads declared metadata first, and substring inference is a deprecated fallback that logs a warning.- Tool names are one global namespace, arbitrated first-wins and never silently: the incumbent keeps the name and the challenger is recorded as a
CapabilityConflictsurfaced throughplugin_list. Rejection is deliberately non-fatal so one colliding plugin cannot break assembly for every other plugin. Never overwrite a live handler or emit a duplicate schema to claim a name another plugin already owns. - Dependencies arrive late, through
bind_runtime, and absence degrades rather than crashes: plugins declare names independenciesand receive matching services injected in provider→consumer topological order, independent of discovery order. A plugin module must not import a runtime service at module level, and must not perform I/O, network calls, or state mutation at import time — every module is required to be importable standalone. A handler whose dependency was never bound returns a structured refusal; it does not raise. - Lifecycle is a fiber, and cleanup is a scope (MANDATORY): every plugin instance lives under a
PluginFiber(PENDING → LOADING → ACTIVE → UNLOADING → DISPOSED, with aLOADING → FAILED → LOADINGretry path) whoseEffectScopedisposes registered effects LIFO, children before parents, idempotently and exception-safely. Anything a plugin registers process-globally — an interceptor onregistry.tool_pipeline, an EventBus subscription, a background task — must be registered as an effect on that scope in the same change, or disable and reload leak it. Illegal transitions raiseIllegalStateTransitioninstead of being silently corrected. - Hot-reload is safe because handlers are snapshotted per turn, not because reload is atomic: each turn copies
dict(registry.tool_handlers)at its start, so an in-flight turn finishes against the handlers it began with whilenotify_mutation()'s version bump invalidates the catalog cache for turns that start later. Reload re-imports throughimportlib.util.spec_from_file_locationusing the source path recorded as__leapflow_plugin_path__— never by mutating globalsys.path, and never in a way that requires the plugin to be importable from the process's import path. Hot-reload's catalog version bump also invalidates_full_tools_tokens(the cached full-catalog token estimate used by prefix commitment amortization); a registration path that bypassesnotify_mutation()leaves the commitment controller operating on a stale cost estimate. - Plugin governance is cold-path (MANDATORY): fiber state, trust ledgers, usage statistics, health producers, advisors, proposal queues, and marketplace work must add no per-turn cost to the hot path. Trust is flushed to DuckDB only on level transitions (plus a final
atexitflush), and usage samples stay in bounded deques. A governance feature that measurably slows an ordinary turn is a defect in the feature, not a cost to accept. - Plugin mutation is uniformly HIGH risk and never permanently granted: any action whose
metadata.platform == "plugin_management"is forced toRiskLevel.HIGHwithallow_permanent=Falseinsecurity/risk.py— defense-in-depth that holds even when caller metadata is wrong. Install, reload, rollback, enable, disable, and remove each build anActionDescriptorand go throughApprovalOrchestratorper invocation. The single exemption isplugin_reloadatPRODUCTIONtrust, which is earned evidence rather than a configured bypass. With no gate installed (in-process CLI binds none), every mutation is denied: code that can rewrite the agent's own composition must never be installable through an unguarded path. - Self-evolution is a governed pipeline, not a code-writing shortcut: capability gap → proposal → generate → validate (syntax → structure → import/Protocol conformance) → compatibility assessment → approval → write → sandbox smoke → register at DRAFT → behavior tests → probation → trust accrual → verify, with quarantine and rollback as the failure path. An
INCOMPATIBLEverdict is rejected before any file write; a failure at any later stage rolls back the fiber, thesys.modulesentry, and the written file. Each next action comes fromAdaptiveEvolutionPolicyreading structured requirement, risk, trust, and status — never from natural-language intent — and the autonomy level is configuration, so raising it is a deliberate operator decision rather than a code path. - Co-evolution makes every capability transition visible, including retirement: an environment-driven install, reload, disable, rollback, or remove records its causal requirement/evidence, pre- and post-mutation catalog state, artifact identity or digest, compatibility and behavior verdicts, approval, and resulting fiber/trust state through the existing audit, proposal, version, and lifecycle stores. An unselected or superseded proposal must be resolved, expired, or retained with an explicit reason; no proposal, plugin artifact, or registered effect may become an untracked permanent residue.
- Untrusted code is isolated before it is trusted:
requires_sandboxdefaults toTrue; sandboxed plugins run in a subprocess over JSON-RPC with a bounded invoke timeout and receive no host-side runtime dependencies. Marketplace artifacts are verified by SHA-256 checksum and, when trusted pubkeys are configured, by Ed25519 signature over the canonicalname|version|entry_point|checksum_sha256payload. Validation re-runs on the install path even for marketplace code that was already checked. - Plugins are process-global; sessions are not: the registry is a daemon-wide singleton, so install, reload, disable, and remove change the capability set for every connected client at its next turn, and trust accrues from all of them. Any change to plugin state must be assessed against the concurrent-TUI contract — per-turn snapshots are the only isolation, and there is deliberately no per-workspace plugin set.
- Self-capability answers come from the live registry, never from documentation: when LeapFlow reports what it supports — plugins, self-evolution, hot reload, version management — the evidence is
plugin_list's livecapability_reportor an equivalent runtime registry read. If runtime introspection fails, state that the running state could not be verified; never infer a capability from README, design docs, or memory. self_describeis the canonical tool for agent self-cognition: all facets read live runtime state (registry, daemon, engine, build_info) throughbind_runtimeinjected services, never from documentation or static config. A new runtime observable (e.g., a new daemon metric, a new engine state) that the agent should be aware of must be wired into the appropriateself_describefacet in the same change — an observable that exists only in daemon status but not in any tool is invisible to the agent.- The plugin contract is published, so it changes with the code:
docs/plugins/third_party_plugin_development.md(interfaces, deployment, security model) anddocs/plugins/plugin_lifecycle_management.md(lifecycle, governance matrix, enforcement status) are third-party-facing specifications whose tables state what the code does today. A change to a Protocol, a lifecycle transition, an approval rule, a config key, or an injectable dependency name updates them in the same change — and never promotes a roadmap entry to ENFORCED ahead of the wiring.
LeapFlow exposes four levels of extension, ordered from lowest barrier to deepest integration. Pick the lowest level that satisfies the requirement — it will ship faster, carry less maintenance cost, and stay compatible across upgrades.
- Pure Markdown file with YAML frontmatter; no code changes required.
- Add a
SKILL.mdto the skills directory and the agent discovers it at startup. - Declares tool dependencies, trigger phrases, category, and platform constraints.
- The LLM reads the skill document and autonomously calls existing tools to execute the workflow.
- Compatible with Hermes skill format (
metadata.hermesnamespace). - Best for: custom workflows, domain knowledge, operational playbooks, guided procedures.
- External process communicating via Model Context Protocol (JSON-RPC over stdio/SSE).
- Brings external service capabilities into the agent as discoverable tools.
- Language-agnostic — any runtime that speaks MCP can serve tools.
- Best for: external API integrations, third-party service connectors, language-specific tooling.
- Python module implementing the
ToolPluginProtocol (runtime_checkable). - Full access to LeapFlow's runtime: EventBus, memory, storage, settings via
bind_runtime. - Subject to Progressive Trust lifecycle: DRAFT → CANDIDATE → VERIFIED → PRODUCTION.
- Sandbox isolation via subprocess JSON-RPC until trust is earned.
- Best for: deep framework integration, new LLM providers, custom storage backends, platform adapters.
- Direct modification to LeapFlow's core tool system (
leapflow/tools/). - Requires understanding of internal architecture, review process, and compliance with all rules in this document.
- Best for: fundamental capabilities that all plugins and skills may depend on.
| Level | Mechanism | Barrier | Use Case | Example |
|---|---|---|---|---|
| 1 | Skill (SKILL.md) |
Lowest | Workflows, playbooks, domain knowledge | Deployment checklist, code-review guide |
| 2 | MCP Server | Low | External services, cross-language tools | GitHub API connector, database explorer |
| 3 | Plugin (Python) | Medium | Runtime integration, providers, adapters | LLM provider, gateway adapter |
| 4 | Core Tool | High | Foundational agent capabilities | File I/O, code search |
Community contributions should start at Level 1 (Skill). It requires no code changes, has the fastest feedback loop, and can be shared as a single Markdown file. Escalate to a higher level only when the skill layer cannot express the needed capability.
- Path tree is a product contract: every LeapFlow-managed path must be declared by
PathLayout,ProfileLayout,CacheLayout, or a child layout object. Runtime code must consume layout APIs, never assemble managed paths with ad-hoc string joins. - Profile is the runtime boundary:
profiles/<profile>/owns profile metadata, config, DBs, memory, skills, gateway state, approval state, audit logs, runtime files, cache roots, and profile-scoped secrets. Cross-profile access requires an explicit layout object. - Workspace is context, not ownership: workspace-local files are limited to
.leapflow/config.yamland.leapflow/workspace.yaml; profile data and caches stay under the active profile and are addressed by workspace/session ids. - Config is layered, not scattered: durable settings live in
config/user.yaml,profiles/<profile>/config/*.yaml, and optional workspace config.LEAPFLOW_*values are process overrides only; env files are not a supported configuration source. leap configis the user-facing control plane: every durable, user-writable setting must be discoverable and mutable throughConfigService,leap config, and TUI/config; do not add one-off setup commands, hidden YAML-only knobs, or new persistent env-first flows.- Config catalog is the discovery contract: writable config fields must expose key, effective value, type, scopes, hot-reload semantics, category, value hint, and description.
leap config keysstays compact and script-friendly;leap config listand/config listare the human-readable catalog;leap config show <key>and/config show <key>are the single-field detail views. - TUI config parity is mandatory:
/configmirrorsleap config, supports active-session reload when possible, and must remain self-discoverable through slash completion for subcommands, keys, and simple values. Any new config subcommand must update CLI parser, TUI payload/rendering, completion, README, and tests together. - Secrets are refs, never durable plaintext: long-lived LLM, VLM, aux-provider, Gateway, and Hub credentials must be stored in the vault and referenced as
secret://profile/...orsecret://global/.... Config may contain refs, never tokens. - Cache declares scope and sensitivity: cache paths must route through
CacheLayout/CacheManagerwithprofile,workspace, orsessionscope. Session visual/video/VLM/signal artifacts are sensitive, non-syncable, and TTL/quota managed by index. - Safety follows path semantics: daemon sockets, pid/lock files, runtime state, DuckDB files, vault files, approval grants, audit logs, and memory stores must flow through layout descriptors, path sensitivity, risk, approval, and redaction gates.
- No legacy aliases: do not reintroduce global
.envas persistent config, flat cache roots, profile-root gateway config, inline credential files,.credential_key, orrun/runtime paths.
Every capability that changes the world outside the current turn — shell execution, sensitive file read/write, config mutation, outbound sends, platform actions, network egress, desktop control, plugin self-modification — reaches the user through one approval chain. Wiring a new one is a fixed sequence, not a design exercise: each rule below is the residue of a defect that shipped with a green suite.
- One orchestrator is the only entry point (MANDATORY): build an
ActionDescriptor— itskind/effect/resource/metadatais the contract — and callApprovalOrchestrator.evaluate(), which supplies risk classification, policy, existing grants, and the audit record for free. Never hand-roll a confirmation prompt, a per-tool approval allow-list, or a second gate implementation. Never call the orchestrator's legacycheck(command): it is the single-argument shell adapter, and a gate exposing onlycheck()must be treated as unusable and denied rather than assumed permissive.config_setcopiedfile_write's four-argument gate call and failed on every invocation withcheck() takes 2 positional arguments but 4 were given, so no model-driven config change ever reached approval — invisibly, because the test fakes implemented the same wrong signature. - Register the gate in both installation sites: gates are process-global and injected twice — in-process through
cli/context.py, daemon-side throughApprovalCoordinator.install_gate(). A new sensitive capability must be wired in both, or it behaves differently depending on whetherleapdis running. Where a mode legitimately cannot supply a gate (in-process CLI binds noplugin_approval_gate), the tool documents it and fails closed instead of proceeding unguarded. - Fail closed on absence and on exception (MANDATORY): no gate installed, no per-turn route, or a gate that raises all mean deny, with a message the model can act on — a broken gate must never become an open door. Keep the
exceptnarrow, and never raise from inside oneexceptbranch expecting a siblingexceptto catch it; that is how aTypeErrorfallback escaped its handler and surfaced raw instead of failing closed. - Feasibility precedes consent: an action that cannot succeed must never reach a human. The order is fixed — dedup → payload validation → capability feasibility (
CapabilityHealthLedger,blocks_approval) → resource provenance → approval → execute. Missing scopes, degraded capabilities, and admin-required failures return a deterministic repair instruction and hard-stop the turn, withsecurity/permission_failures.pyas the single authority both engine and TUI consult. Prompting for consent to a call that will be refused for lack of permission teaches users to click through prompts. - Gate the action that actually executes, at every hop: classify once and the transport can still go somewhere else.
web_fetch's egress gate only saw the first hop while the transports followed redirects themselves, so any server answering302could bounce an approved public URL to loopback or a cloud metadata endpoint. Transports are single-hop and reportLocation; the caller re-classifies and re-gates each hop, and reachability asksis_globalrather thannot is_privateso unenumerated non-routable ranges fail closed. - The decision vocabulary is one enum across the process boundary:
ApprovalDecisionis the entire vocabulary, so adding or renaming a value means updating the enum, the orchestrator's_choicesand scope mapping,ApprovalCoordinator._normalize_decision, the TUI modal, and the RPC together. The daemon normalizer kept a hardcoded allowed-set that predatedALLOW_ALL_SESSION, silently rewrote that choice todeny, and never armed the session bypass it was meant to grant — explicit user consent became a refusal with no error logged anywhere. - Scope is the grant's contract, and a bypass is resolved in one predicate: grants persist only at
SESSION/PROFILEscope viaApprovalScope, andHIGH/CRITICALrisk setsallow_permanent=Falseso "always allow" is never offered for actions that change the agent's own composition or reach internal addresses. A bypass — configapproval_bypassor a session-level "allow all" — is answered by a single predicate every gate consults, so it cannot mean "approved" at one gate and "still ask" at the next. Hardline blocks sit above all of it and are never bypassable by any grant, scope, or bypass flag. - Prompt and audit text is redacted at the descriptor:
ApprovalRequest.detailis both rendered to the user and persisted to the JSONL audit log, so secrets are removed when the descriptor is built, not when it is displayed. URL query, fragment, and userinfo are stripped; config values are excluded entirely and only the key is shown. Approval details were carrying API keys and signed tokens out of query strings straight into the audit log. - Daemon approvals are turn-routed and terminally resolved: prompts travel on the per-turn
approval_routeContextVar(queue, request_id), so a prompt never surfaces in a client that did not cause it. Every pending future needs a terminal path —deny_for_requestwhen the turn ends,deny_for_queuewhen the stream closes,prune_staleon TTL — because an unresolved future blocks its tool call forever. The ContextVar must be set and reset within the samecontextvars.Context; setting it in one per-chunk task and resetting it in another raised "was created in a different Context" and broke streaming for every gated action. - A denial is terminal and reaches the model verbatim:
ApprovalResult.denial_messagestates that the user did not consent and that the outcome must not be retried, rephrased, or pursued through another tool. An adapter wrapping a gate must capture that message and return it to the caller; substituting a generic tool error lets the agent reroute around the human's refusal. - Approval invariants are tested against the real orchestrator:
tests/test_approval_layer.pydrives the productionApprovalOrchestrator— per-scope grant reuse, hardline deny without prompting,allow_permanentsuppression, decision round-trips through the RPC vocabulary, denial-message pass-through. A hand-written fake is acceptable only as the human surface; a fake that reimplements the orchestrator's own interface will keep agreeing with the caller's mistake, which is precisely how a dead gate stayed green.
- Define the Protocol first — the contract is the design
- Implement against the Protocol, never against another implementation
- Consider affected user journeys before changing shared flows; do not introduce regressions, broken links, or worse experiences in adjacent paths
- Keep common paths transparent: long-running work must stream progress, surface recoverable errors clearly, and avoid silent stalls.
- For context assembly, prefer manifest-driven progressive disclosure over shortcuts or intent-handler sprawl: expose compact capability indexes, selected schemas, and targeted memory only when the current plan needs them.
- For gateway or IM work, define the signal contract first: event source, normalizer/classifier, trigger policy, session routing, memory/audit path, and outbound action path. Default inbound activation to least privilege (
mention_onlyor equivalent), filter self-generated messages before LLM invocation, and keep cross-chat or proactive sends behind Progressive Trust and ApprovalGate. - Avoid rule-based natural-language fitting by default. Do not add keyword/action-verb/alias enumerations, intent-handler taxonomies, or brittle routing rules when LLM-native capability disclosure, manifests, schemas, protocols, or configuration-driven contracts can solve the problem. If a rule-based method is truly unavoidable for a stable protocol boundary, offline fallback, or safety hard gate, explain the necessity, scope, alternatives, and rollback path to a human and obtain explicit second confirmation before implementation.
- Specifically within
engine/: reference resolution, error classification, context disclosure, and focus tracking must not use keyword/substring matching to classify user intent or error type. Reference resolution usesSessionFocusStateentity registry lookups; error classification uses data-table or registry-pattern matching (never if-elif chains on message text); context disclosure reads tool-declaredx_leapflowmetadata first and treats substring inference as a deprecated fallback that logs a warning. - Preserve security and audit paths: dangerous actions, file writes, outbound messages, credentials, and path access must flow through the existing policy, approval, redaction, and audit mechanisms.
- Preserve gateway safety boundaries: inbound credentials stay in CredentialVault; outbound send/write/execute actions go through ApprovalGate; bot self-messages and duplicate events are filtered before routing; platform-specific metadata must remain in
metadataescape hatches instead of polluting core message types. - Keep App Connector governance thin: platform core should consume normalized contracts and failures, while app-specific auth scopes, CLI/SDK error parsing, vendor recovery steps, and command templates remain in action packs, adapters, or backend-specific helpers. If a new platform requires changing gateway core business rules, first refactor toward a protocol hook or app-side classifier.
- To add a built-in plugin, add its module path to
_BUILTIN_PLUGIN_MODULESinplugins/tool_plugins/__init__.pyand expose a module-levelplugininstance; that wrapper module is the only place allowed to import the tool implementation it exposes. - Tool handlers are
async, accept the parameters their schema declares, and return a structureddict. Expected failures come back as{"ok": False, "error": ...}, which records an ordinary trust-affecting failure; raising instead reserves the signal for a genuine internal defect. Permanent trust freezing is ahard_failurerecorded throughLifecycleGovernor, not something an escaping exception should trigger by accident. - Maintain backward-compatible migrations for persistent state, configuration, skills, trajectories, sessions, and profile data.
- Write unit tests before or alongside the implementation
- Integrate via EventBus events, not direct function calls between modules
- Every module must be importable standalone without side effects
- No placeholder stubs — implement fully or do not add the code
- ANSI output must check
sys.stdout.isatty()before emitting escape codes - For error recovery, route all failures through the
RecoveryCoordinator— classify into aFailureEnvelope, receive aRecoveryDecisionwith an explainablereasonandstrategy_key, then feed the outcome back. Never handle errors with ad-hoc if/break in the loop body. - Recovery strategies are standalone Protocol implementations with
can_apply()+decide(). Add new strategies by registration, never by modifying the coordinator's decision logic. - When automatic recovery exhausts its budget or encounters non-recoverable failures, emit a structured
InteractionRequest(typed action, severity, suggested actions, timeout behavior, resumption key) — not raw text appended to conversation. A terminal decision that carries one must surface it: render its title, description, and suggested actions for the user, and pass the structured payload to the client so it can prompt and resume byresumption_key. Dropping it back todecision.reasontells the user a turn stopped without saying what to do.
- Deep review for large changes: When a change substantially affects architecture, runtime behavior, user flows, persistence, safety, or multiple modules, perform an additional deep review before considering the work complete.
- Human confirmation for TUI changes: Any TUI layout or interaction-logic change requires a second human confirmation before it is considered ready.
- Human confirmation for slash-command paths (MANDATORY, no exceptions): Any change to what a slash command (
/...) does requires a second human confirmation before it is considered ready — never ship it on a single pass. This covers the whole surface: registry, router, in-process REPL, daemon REPL,command_execute(including its RPC signature and parameter plumbing), completion, and rendering.- Functional changes count even when the command surface is unchanged. The name, arguments, and help text staying identical does NOT waive confirmation. Altering what the command observes, targets, arms, schedules, sends, opens, or persists — or which session/workspace/profile it resolves against — is a functional change and must be confirmed.
- Equally in scope: dispatch and routing, prompts and confirmations, emitted output, browser/dashboard launches, background work the command triggers (watches, schedules, re-entries), and error/recovery messaging.
- Passing tests and a clean lint run are NOT a substitute for confirmation. Slash commands are the primary user-facing control plane; correctness of the visible behavior is only established by a human check.
- State the pending confirmation explicitly in the handoff, and name the behavior a human should exercise to verify it.
- Human confirmation for approval-path changes: any change to what reaches the approval chain — a newly gated capability, a new
ActionKind, anApprovalDecision/scope/bypass change, or gate registration — requires exercising the real prompt by hand in both in-process and daemon mode before it is considered ready. Gates are process-global and injected twice, so a green suite proves at most that one of the two wirings works; every approval defect recorded in this document passed its tests. - Deep review for plugin composition changes: a change to a Protocol signature, discovery, fiber lifecycle, trust thresholds, sandbox policy, or marketplace verification alters what the agent can load and execute at runtime. Exercise it against a real registry (register → publish → reload → dispose) rather than a fake, and re-check the concurrent-client and cold-path implications before considering it complete.
- Design goal check: Verify that the implementation actually achieves the intended design goal and is not just a local patch.
- Optimality check: Evaluate whether the solution is the simplest robust design, avoids unnecessary abstractions, and fits the existing architecture.
- Regression impact check: Inspect affected modules and user journeys for logic bugs, degraded UX, broken compatibility, slower feedback, weaker diagnostics, or worse failure recovery.
- SOLID and extensibility check: Look for responsibility leaks, tight coupling, hardcoded paths/thresholds/rules, magic strings, and choices that reduce generalization or future extension.
- Fix what the review finds: If the review identifies correctness, design, UX, SOLID, hardcoding, or extensibility issues, fix and simplify them directly rather than only reporting them.
- Anti-hardcoding audit for engine/: All regex patterns, keyword lists, and if-elif classification chains in
engine/must be audited against the rule-avoidance principle (Implementation Guidelines). Permitted only when the pattern falls into one of the three explicit exemptions (stable protocol boundary, offline fallback, safety hard gate). Non-exempt hard rules must be refactored to Protocol/registry/config-driven implementations. When reviewing engine changes, verify that new code does not introduce keyword enumerations, substring routing, or magic-number thresholds without a configuration escape hatch.
The suite has two layers with different jobs, and keeping the boundary sharp is what stops each from doing the other's work badly.
Mock layer (tests/*.py, marked unit/component) — broad and fast. It owns pure algorithms, state-machine branches, error-classification tables, rare and malformed inputs, and single-module invariants. Branch combinations can only be enumerated here, and only here is the feedback measured in milliseconds.
Real layer (tests/journeys/, marked e2e) — six coarse journeys driving a real leapd subprocess over RPC, with the LLM boundary served by a local cassette proxy. It owns what a mock structurally cannot observe: cross-module wiring, process boundaries, session identity, workspace binding, real persistence, and pushed runtime metadata. Every incident recorded in this document shipped with a green mock suite.
Three rules keep the split honest:
- Anything provable with one mock-layer assertion must not enter the real layer.
- One journey is one test case. Express variation as ordered phases inside a single session; never parameterize a journey.
- The real layer has a hard case budget (
tests/regression/test_suite_budget.py). When it is reached, merge a journey — do not raise the ceiling. The budget is what lets the real layer run on every push, and a suite that can be skipped will be skipped.
The two layers are joined by recorded traffic. tests/_fixtures/cassettes/ holds the deterministic inputs the offline lanes replay (rebuilt with make seed-cassettes); tests/_fixtures/recordings/ holds real provider traffic captured by make record-traffic. make sync-fixtures distils both into tests/_fixtures/llm_responses/response_shapes.json, which the mock layer asserts against — so a provider dropping a field turns the build red instead of passing forever against a body written from memory. Recording never writes to the replay store: a multi-turn agent conversation cannot be replayed from a recording, because each turn's prompt embeds the exact round-by-round history of the turns before it.
Each journey also declares two cost ceilings, both enforced at the proxy and reported by finish(). max_llm_calls is the convergence guard: a turn that stops converging is cut off instead of running to the engine's iteration cap. max_llm_tokens is the cost guard, and it catches what call count cannot — prompt growth raises the bill without adding a single round. Raising a ceiling is not the fix for hitting one.
Which tests run. The offline lanes never select: the mock layer runs in full, and the real journeys run in full on every push. Selection would save at most a few seconds, because the always-on tiers set the floor, and a suite that can be skipped will be skipped. The live lane is the exception — there a journey costs real tokens, so it selects: each journey declares SUBJECT_PATHS (the sources it exercises) and LIVE_SIGNAL (whether a real provider adds signal), and tools/impact.py --live-journeys picks from those. Declaring them inside the journey keeps that knowledge next to the assertions it describes. tools/impact.py can also scope the mock layer (make test-impact) for local work on a large change; it is not wired into CI.
- Unit tests must be hermetic: no network, no LLM calls
- Journeys must not mock anything: a journey that reaches for
unittest.mockhas become an expensive unit test - Provider bodies are recorded, not written: use a cassette or a derived fixture; a hand-authored body keeps passing after the provider changes
- Faked construction needs real-instance cover:
object.__new__plus private-attribute assignment is acceptable only for ordering contracts, and only when the same file also builds the class properly and drives the production path - py_compile all modified files: syntax errors caught before test run
- Import chain verification: every new module must be importable standalone
- Existing tests must not regress: all tests must pass after every change
- User-facing flows must not regress: preserve or improve usability, feedback clarity, and failure recovery for impacted paths
- Verification sequence: compile → import → mock layer → real journeys
- Behavior contracts over snapshots: assert invariants, not frozen values
- Mock at boundaries only: mock external I/O (network, disk), never internal logic
- A test may not fabricate the wiring it claims to cover: building an object with
object.__new__and assigning the private attributes the code reads cannot detect a wrong attribute name — the test simply agrees with the typo. Calibration tests did exactly that and stayed green while every real turn raisedAttributeError. Any test whose stated purpose is wiring must construct the real object and drive the production path. - Multi-client behavior needs multi-client tests: session routing,
status(), stream metadata, and client-lease changes require two sessions in two workspaces asserting that neither sees the other's identity, usage, or turn state. Single-session tests cannot observe cross-client leakage, which is why a leak shipped with a green suite. - Co-evolution experiments need longitudinal counterfactual evidence: exercise an unchanged baseline, an irrelevant delta that is correctly rejected, a task-relevant delta already satisfied by the live catalog, and an unmet requirement that traverses the real governed lifecycle. Use isolated profiles and real integration seams; deterministic reruns are not independent environmental units. Freeze the experiment configuration and corpus, preserve an immutable evidence bundle with a digest and known threats, and report the evidence level, unit count, no-op/rejection outcomes, mutations, and final steady state.
- Change-scoped validation: Run the most specific relevant tests first, then broaden only as needed: CLI/TUI changes require CLI/TUI tests; leapd changes require daemon RPC/lifecycle tests; storage or memory changes require persistence tests; gateway, IM, event-source, or approval changes require connector lifecycle, event normalization, routing, idempotency, self-message filtering, security/approval, and failure-recovery tests; plugin contract, registry, lifecycle, sandbox, marketplace, or trust changes require the plugin reload, scoped-registry, fiber/effect-scope, sandbox, marketplace-signing, trust-learning, and architecture-contract tests; skills, learning, perception, and copilot changes require their lifecycle or pipeline tests.
- Recovery strategy isolation: Each
RecoveryStrategymust be testable in isolation — verifycan_applypredicates,decideoutputs, and side-effect-state gating independently of the coordinator and other strategies. - Budget boundary tests: Verify that recovery budgets exhaust correctly (per-category, per-turn, deadline), that exhaustion produces a deterministic halt decision, and that cost accounting is exact.
- Over-engineering: if you need 3+ files for a simple feature, rethink
- Premature optimization: measure first, optimize only bottlenecks
- God objects: no class should exceed 500 lines; approaching this limit requires checking whether policy, state, rendering, protocol, storage, or adapter responsibilities should be split.
- Magic strings: use constants or enums
- Blocking the event loop: all IO must be async or
run_in_executor - Hardcoded paths, URLs, thresholds without config escape hatch
- Chinese comments in source code (English only)
- Speculative infrastructure: no hooks or extension points without a concrete consumer
- Mixing long-lived event observation into short-lived action execution;
CliBackendis for bounded commands, while streaming CLI consumers, webhooks, polling, and WebSocket subscriptions belong behindBackendEventSource-style contracts. - Activating IM agents on all inbound messages by default, skipping self-message filtering, or allowing cross-chat/proactive sends before Progressive Trust and approval policies explicitly permit them.
- Putting third-party app business code into platform core: do not add vendor scopes, lark-cli/SDK JSON parsing, auth command construction, console-specific recovery instructions, or resource-specific branching to gateway-wide modules such as action registries, capability ledgers, approval gates, or engine recovery paths.
- Shortcut-style natural-language fitting and large intent-handler taxonomies; use stable runtime gates plus capability manifests instead. Rule-based keyword/action-verb/alias matching is prohibited by default and requires explicit human second confirmation before implementation when unavoidable.
- Scattered if/break/continue recovery decisions inside the agent loop body; all error handling enters the
RecoveryCoordinatoras aFailureEnvelope - Magic retry counts or unbounded retry loops without budget constraints and deadline enforcement
- Feeding unstructured error text back to the LLM without classification, recoverability assessment, or side-effect awareness
- Multiple parallel error-handling paths for the same failure domain (LLM errors in one handler, tool errors in another, security errors in a third); use a unified classification and coordination pipeline
- Widening a provider call's
tryblock to cover bookkeeping, or classifying a Python-level defect by matching its message text against provider conditions - Answering "the caller's session" with "the most recently active session", or letting a client adopt a session id the daemon happened to report
- Hand-rolled confirmation prompts, per-tool approval allow-lists, or any second gate beside
ApprovalOrchestrator; a sensitive capability declares anActionDescriptorand evaluates it - Treating a missing, unbound, or raising approval gate as permission to proceed
- Putting a secret, token, or config value into
ApprovalRequest.detail— it is rendered to the user and persisted to the audit log - Extending
ApprovalDecisionwithout updating the daemon normalizer, TUI modal, and RPC in the same change - Importing a tool implementation from plugin core, or re-creating a
leapflow.toolsshim for a relocated plugin-subsystem module - Module-level I/O, network calls, or runtime-service imports in a plugin module; dependencies arrive through
bind_runtime - Registering a process-global interceptor, subscription, or background task without a matching cleanup effect on the plugin's
EffectScope - Reloading a plugin by injecting into
sys.pathinstead of a file-backed import spec, or overwriting a live handler to claim a tool name another plugin owns - Adding per-turn cost for plugin governance (trust, stats, health, advisor, proposals) — governance is cold-path
- Treating an environment delta, a model-authored evolution hypothesis, or a successful experiment run as authorization to mutate the framework; each must still pass evidence validation, live resolution, policy, approval, sandboxing, lifecycle, and trust gates
- Using an experiment harness to register plugins directly, forge observations/proposals, bypass approval, or write synthetic evidence into production state without an explicit experiment-profile boundary
- Answering a question about LeapFlow's own capabilities from documentation or memory instead of a live registry read
- Bare
except:clauses — always specify the exception type # TODO: implementstubs — implement or don't commit
- Files:
snake_case.py— noun for types (signal_event.py), verb for actions (compress_context.py) - Classes:
PascalCase— Protocols suffixed with purpose (SignalSource,SkillStore) - Functions/Methods:
snake_case— verb-first (filter_noise,retrieve_context) - Constants:
UPPER_SNAKE_CASE— grouped in module-level or dedicatedconstants.py - Private: single underscore prefix (
_internal_method) — never double underscore - Modules/Packages: short, singular nouns (
perception,causal,memory) - Events: past-tense domain verbs (
SignalReceived,SkillVerified,ContextCompressed) - Config keys:
dot.separated.lowercasein env/settings (copilot.idle_threshold_ms)