From dd86a7574cb02db998319489388a05b38161b2bf Mon Sep 17 00:00:00 2001 From: sohams25 <86193836+sohams25@users.noreply.github.com> Date: Wed, 22 Jul 2026 18:17:30 +0530 Subject: [PATCH] docs: align README and header with v0.3.0 - intro rewritten: what earlyon does, then the execution limitation (eager routing; all-exits ONNX is not conditional early exit; staged runtime is narrow) before any example - install: primary command is the v0.3.0 GitHub tag; PyPI shown only as a clearly labelled pending note - banner/demo SVGs: subtitle now 'Train, calibrate, and benchmark early-exit vision models'; illustrative-only labelling; numeric per-exit percentages and the deprecated computation_used field removed; accurate alt text - per-head calibration section + pipeline flow; explicit-enablement wording - benchmark section: single bounded-run table with negatives preserved, noise 2.05x labelled synthetic best-case, legacy v0.2 tables and appendix removed (quarantined in docs/benchmarks.json) - ERRATUM (#6): the run's estimated FLOPs fraction came from the flagged low-confidence uniform fallback (reuse detector false-positive on torchvision ResNet shared ReLU); fvcore attribution gives ~0.92 not 0.88 (~8% estimated saving). Noted in README, CUDA_EVIDENCE.md and appended to the evidence JSON without altering recorded fields. No core code changed. - 'Why not just deploy a smaller model?' section; execution/export modes and Jetson/TensorRT status sections; roadmap narrowed with explicit feature freeze; custom_ee claims scoped to the tested contract - preferred examples use estimated_backbone_flops_fraction; archived marketing copy quarantine-labelled --- README.md | 316 +++++++++++++++++-------------- assets/banner.svg | 8 +- assets/demo.svg | 23 +-- docs/evidence/CUDA_EVIDENCE.md | 13 ++ docs/evidence/cuda_evidence.json | 9 + docs/marketing/earlyon-reddit.md | 6 +- 6 files changed, 214 insertions(+), 161 deletions(-) diff --git a/README.md b/README.md index f4cddf8..84716f5 100644 --- a/README.md +++ b/README.md @@ -11,12 +11,20 @@ license

-> **`earlyon`** — early exits for PyTorch CV models. Small classifier heads -> partway through the network let easy images stop computing once an exit is -> confident; only hard images run every layer. Attaches to your existing -> backbone via forward hooks. No rewrite, weights load unchanged. +`earlyon` adds intermediate classifier heads to PyTorch vision models so +suitable inputs can stop before the full backbone runs. It includes +exit-head training, independent calibration for every head, threshold +selection, checkpoint migration, fair benchmarking and a reference staged +runtime. ---- +Two things to know up front. Dynamic routing currently runs in eager +PyTorch. The standard ONNX export computes all exits and is not, by itself, +a conditional early-exit runtime. For supported architectures, the +reference staged runtime can genuinely skip later stages; it is +intentionally narrow and is not a universal model-partitioning system. + +The primary target is batch-1, latency-sensitive inference; batched routing +exists but is conservative. ```python import torch @@ -26,25 +34,29 @@ from earlyon.models import resnet50_ee model = resnet50_ee(num_classes=10, pretrained=True).eval() result = model(torch.randn(1, 3, 224, 224), mode="inference") -result.exit_taken # which head answered; -1 = full network -result.estimated_backbone_flops_fraction # static estimate, not a measurement -result.confidence # softmax confidence where it answered +result.exit_taken # which head answered; -1 = full network +result.estimated_backbone_flops_fraction # estimated fraction of backbone FLOPs + # executed — not measured latency or energy +result.confidence # softmax confidence where it answered ``` -earlyon is a research-to-deployment toolkit. Its primary target is batch-1, -latency-sensitive inference; batched routing exists but is conservative. -

- earlyon demo: an easy image exits at the first head using 12% of the FLOPs, a hard image runs the whole network + earlyon demo with illustrative values: an easy image exits at the first head so later layers never run; a hard image runs the whole network

## Install +PyPI publication is pending trusted-publisher configuration. The v0.3.0 +GitHub tag is the current installable release: + ```bash -pip install earlyon +pip install git+https://github.com/sohams25/earlyon@v0.3.0 ``` -Or from source for development: +(Once PyPI publication completes, `pip install earlyon==0.3.0` will work; +until then it does not.) + +From source for development: ```bash git clone https://github.com/sohams25/earlyon.git @@ -62,15 +74,18 @@ a car and a blurry, half-occluded bird both pay for all fifty layers, even though the car is decided after a handful of them. At batch size 1 on an edge device, that flat cost is your latency and your power budget. -The fix has been in the literature for years: BranchyNet in 2016, then eight -years of follow-ups collected in a 2024 ACM survey (10.1145/3698767) showing -1.3 to 2.5× compute savings at single-sample edge inference. What was -missing is the tool. Each paper ships custom code for one architecture; -earlyon is the `pip install` version. +The idea has been in the literature for years: BranchyNet in 2016, then +eight years of follow-ups collected in a 2024 ACM survey (10.1145/3698767) +reporting 1.3 to 2.5× *estimated compute* savings at single-sample edge +inference. Each paper ships custom code for one architecture; earlyon is an +installable toolkit for the same idea — and its own benchmark below shows +that estimated savings do not automatically become wall-clock savings. ## What it does -Six ready-made factories, or wrap anything: +Built-in factories are provided for the documented architectures. A custom +wrapper is also available when the model's intermediate features, input +signature and classifier outputs can be identified. | Factory | Backbone | Exits | |---|---|---| @@ -81,10 +96,13 @@ Six ready-made factories, or wrap anything: | `cifar_resnet_ee` | CIFAR-native ResNet (He et al. 2015) | 3 (3×3 stem, no maxpool, native 32×32) | | `vit_b_16_ee` | torchvision ViT-B/16 | 2 (after encoder blocks 3 and 9) | -`custom_ee` attaches exits at named layers of any `nn.Module` and infers each -head's width from one dry-run forward. Conv (4D) and transformer token (3D) -features both work; the backbone must already return `(B, num_classes)` -logits. +`custom_ee` attaches exits at named layers of a user-provided `nn.Module`, +inferring each head's width from one validated dry-run forward. The tested +contract: conv (4D) and transformer token (3D) features; positional and +keyword example inputs; tuple/dict layer outputs via `feature_extractors`; +the backbone must return `(B, num_classes)` logits. Reused, missing or +out-of-order exit layers fail explicitly — custom support is not automatic +for every PyTorch model. ```python from earlyon.models import custom_ee @@ -97,91 +115,87 @@ Routing has two policies. `"confidence"` (default) exits when entropy drops below a threshold, which reads the whole distribution instead of just the top class. Whether an exit may fire at all is an explicit per-exit boolean (`enabled_exits`) — a disabled exit can never fire, not -even at softmax confidence exactly 1.0. Calibration and -`save_wrapper`/`load_wrapper` are policy-aware, so a calibrated entropy -model reloads as one; checkpoints carry a versioned schema (v2) and older -files migrate with a warning. +even at softmax confidence exactly 1.0, and never through a numerical +sentinel threshold. Calibration and `save_wrapper`/`load_wrapper` are +policy-aware, so a calibrated entropy model reloads as one. Training is two-stage by default: train the backbone exactly as you already do, freeze it (parameters and BatchNorm stats), then train the small heads for a few epochs. `joint_train_backbone_and_exits` does end-to-end training instead when you want the last bit of accuracy and have the budget. -Temperature scaling (Guo et al. 2017) can be fit before calibration to fix -the usual softmax over-confidence; pass `fit_temperature=True` and each head -— every exit and the final classifier — gets its own fitted temperature. + +## Per-head calibration + +Each exit and the final classifier are calibrated independently. `earlyon` +does not apply one global temperature to every head because their +confidence distributions can differ substantially — a shallow head is +usually far more over-confident than the final classifier. Pass +`fit_temperature=True` and every head gets its own fitted temperature (Guo +et al. 2017); threshold search runs after the fit, on a separate +calibration split, and the test split is never touched by either step. + +```text +train exit heads + ↓ +collect calibration logits + ↓ +fit one temperature per head + ↓ +search exit thresholds + ↓ +evaluate on a separate test split + ↓ +benchmark routed inference +``` + The full split discipline (what may touch which data) is written down in [`docs/CALIBRATION_AND_BENCHMARK_CONTRACT.md`](docs/CALIBRATION_AND_BENCHMARK_CONTRACT.md). -## Measured results (legacy: earlyon v0.2 methodology) - -CIFAR-10, trained on an RTX 4050 Laptop GPU (6 GB) from ImageNet-pretrained -weights: 3–4 epochs for the backbone, 3–4 for the exit heads, thresholds -calibrated to a 1% accuracy budget. **These runs predate v0.3**: they used a -single shared temperature and the pre-fair-runner benchmark; treat them as -indicative, not as v0.3 results. Regenerate with -`python scripts/run_benchmarks.py` (records land in -`docs/benchmarks.json` under `runs`; these live under `legacy_v0_2`). - -| Model | Test Acc | Baseline Acc | Est. FLOPs fraction | % exited early | -|-------------|---------:|-------------:|--------------------:|---------------:| -| ResNet18 | 94.42% | 96.32% | 89.88% | 35.3% | -| ResNet50 | 95.88% | 97.60% | 81.42% | 58.2% | -| MobileNetV2 | 93.31% | 95.08% | 93.90% | 8.5% | - -The FLOPs fraction is a per-image *estimate* over the real test set — a -static analysis of the backbone that excludes the exit heads' own small cost -and all routing overhead. For ResNet50 it reads as ~19% of estimated compute -gone for a 1.7% accuracy cost, with 58% of images never reaching `layer4`. -An estimated FLOPs saving is not a latency claim: wall-clock speedup is -hardware- and input-dependent; see the -[latency appendix](#appendix-wall-clock-latency-legacy) before quoting one. -When judging a deployment, also benchmark a smaller static model at matched -accuracy (pass it as a third entry to `benchmark_models`) — if a plain -ResNet18 matches your routed ResNet50, ship the ResNet18. - -The MobileNetV2 row is a loss, and it's in the table anyway. An -already-compressed backbone leaves little for early exit to skim: only 8.5% -of images leave early, so the wrapper still does ~94% of the work. Rule of -thumb from these runs: the deeper and heavier the backbone, the more there is -to save. - -Reproduce with `python scripts/run_benchmarks.py` (trains and benchmarks) or -`python scripts/re_evaluate.py` (from checkpoints). Raw numbers live in -[`docs/benchmarks.json`](docs/benchmarks.json). - -## Measured results (v0.3 fair runner — bounded evidence run) - -One seeded, deliberately small validation run (2 epochs per stage, RTX 4050 -Laptop GPU, CIFAR-10 at 224px, disjoint train/temperature/calibration/test -splits, identical samples and boundaries for all three models). Full data -and honest interpretation: [`docs/evidence/CUDA_EVIDENCE.md`](docs/evidence/CUDA_EVIDENCE.md). +## Measured results (one seeded, bounded run) + +One seeded validation run (2 epochs per stage, RTX 4050 Laptop GPU, CIFAR-10 +at 224px, seed 42, disjoint train/temperature/calibration/test splits, +identical samples and timing boundaries for all three models). Full +methodology: [`docs/evidence/CUDA_EVIDENCE.md`](docs/evidence/CUDA_EVIDENCE.md). +This is one bounded run on one GPU, not a universal speedup claim. | Model | Test acc | p50 | p95 | Throughput | Est. FLOPs fraction | |---|---:|---:|---:|---:|---:| | ResNet-18 backbone | 94.07% | 1.33 ms | 1.51 ms | 712 ips | 1.00 | -| ResNet-18 early-exit | 92.92% | 1.43 ms | 1.46 ms | 782 ips (**1.10×**) | 0.88 (estimate) | +| ResNet-18 early-exit | 92.92% | 1.43 ms | 1.46 ms | 782 ips (**1.10×**) | ~0.92 (estimate†) | | MobileNetV2 static baseline | 92.67% | 1.40 ms | 1.48 ms | 702 ips (0.99×) | 1.00 | -Read the negative parts too: the early-exit **median** latency is worse than -the backbone's (routing overhead on the 70% of images that don't exit); the -gain is in throughput and the tail (p95/p99), and the static baseline is -competitive at this training budget. That's the honest shape of early exit -on a fast GPU — reproduce with `python scripts/evidence_run.py`. - -## How it compares - -Compression makes the model cheaper for every input; early exit spends -compute per input. They stack. - -| | earlyon | per-paper research code | pruning / distillation | -|---|---|---|---| -| works on your existing backbone | yes, via hooks | one architecture each | retrain required | -| adapts compute per input | yes | yes, for that model | no, fixed cost | -| full accuracy still reachable | yes, hard inputs run everything | yes | no, capacity is gone | -| pip install, tests, CI | yes | rarely | yes, mature tools | - -If you already prune or distill, wrap the compressed model and take both -savings. +† The run originally recorded 0.88, produced by a flagged low-confidence +estimator fallback ([#6](https://github.com/sohams25/earlyon/issues/6)); +fvcore attribution gives ~0.92, i.e. roughly an 8% estimated-compute saving. +See the erratum in the evidence doc. Measured columns are unaffected. + +In this run, early exiting improved throughput by about 10%, but median +latency was slightly worse because exit-head and routing overhead offset +part of the saved backbone computation on the ~70% of images that ran the +full network; the gain shows up in throughput and the tail (p95/p99). A +static MobileNetV2 baseline remained competitive. The result shows why +early-exit systems should be measured end to end rather than judged only by +estimated FLOPs. A noise-input variant reached 2.05× — that is a synthetic +best-case bound (trained heads can fire spuriously on noise), not practical +performance. Reproduce with `python scripts/evidence_run.py`. Older pre-v0.3 +measurements are quarantined under `legacy_v0_2` in +[`docs/benchmarks.json`](docs/benchmarks.json); they are not methodologically +comparable and are not shown here. + +## Why not just deploy a smaller model? + +Often you should. A static model has a fixed cost for every input; early +exit adapts computation per input. When input difficulty varies a lot, +adaptation can win; a smaller static model may still be faster or simpler. +Early exit is useful when input difficulty varies, but it is not a +replacement for testing smaller fixed-cost models. The right baseline is +part of the benchmark, not an afterthought — which is why +`benchmark_models` takes the static model as a first-class entry and feeds +every compared model identical samples with identical timing boundaries. +Compression (pruning, distillation) and early exit are also composable: a +compressed backbone can be wrapped, and the benchmark decides whether the +combination actually helps. ## Usage @@ -212,7 +226,8 @@ calibrate_thresholds(model, val, target_accuracy_drop=0.01) # 4. deploy model.eval() result = model(torch.randn(1, 3, 224, 224), mode="inference") -print(result.exit_taken, result.confidence, result.computation_used) +# estimated fraction of backbone FLOPs executed — not measured latency or energy +print(result.exit_taken, result.confidence, result.estimated_backbone_flops_fraction) ``` The inference path runs under `torch.inference_mode()` internally, so the @@ -226,7 +241,8 @@ from earlyon.core.thresholds import calibrate_thresholds_for_budget result = calibrate_thresholds_for_budget(model, val, target_computation=0.8) result.budget_met # False (plus a warning) if 0.8 is unreachable -result.avg_computation_used # measured on the calibration set +result.avg_computation_used # estimated FLOPs fraction, averaged over the + # calibration set — an estimate, not a measurement ``` An unreachable budget warns and reports `budget_met=False` rather than @@ -303,6 +319,38 @@ config). The full contract is in The reasoning for each choice, including the ugly parts, is in [`docs/DESIGN_DECISIONS.md`](docs/DESIGN_DECISIONS.md). +## Execution and export modes + +Three ways to run a calibrated model. They are not interchangeable: + +- **Eager routed execution** (`model(x, mode="inference")`) — dynamic + routing that can stop after an exit. This is the primary reference + behavior. It carries Python control-flow and host-synchronization + overhead, and `torch.compile` refuses it with a clear error. +- **All-exits ONNX export** (`export_all_exits_to_onnx`) — a static graph + that computes **every** configured exit. Useful for export and analysis; + it does not save later-stage computation and is not conditional early + exit. +- **Reference staged execution** (`earlyon.staged`) — splits supported + (Sequential-shaped) models into stages so later stages genuinely never + run after an exit fires. Tested for eager-equivalence, including a test + that later stages do not execute. It supports only the documented + contract; it is not a universal graph partitioner, and no TensorRT + performance is claimed. See + [`docs/STAGED_DEPLOYMENT.md`](docs/STAGED_DEPLOYMENT.md). + +## Jetson and TensorRT status + +Jetson monitoring and measurement procedures are documented +([`docs/STAGED_DEPLOYMENT.md`](docs/STAGED_DEPLOYMENT.md)): tegrastats +profiling with energy integrated over the timed window when valid telemetry +exists, and unavailable telemetry reported as unavailable (`None`), never as +zero. **No Jetson performance result has been published** — none has been +measured. TensorRT deployment performance has not been measured either; the +staged runtime is a framework contract, not proof of TensorRT deployment. +The CUDA table above is a laptop-GPU run and does not predict Jetson +behavior. + ## Limitations - `forward(mode="inference")` is batch-size 1 — the latency-sensitive edge @@ -312,33 +360,32 @@ The reasoning for each choice, including the ugly parts, is in - Eager routing has real overhead: every enabled exit evaluates its head and synchronises the host (`.item()`) to decide. On a GPU that synchronisation can cost more than the skipped layers save, particularly on small - backbones — theoretical FLOP savings do not guarantee latency savings. + backbones — estimated FLOP savings do not guarantee latency savings. The fair runner reports both so you can see it. -- `torch.compile` cannot trace the routing control flow. The wrapper raises a - clear error instead of silently falling back; compile the raw backbone if - you need it. -- ONNX export (`export_all_exits_to_onnx`) writes a static graph that - computes **every** exit and leaves routing to the caller — it does not - short-circuit compute. For a deployable split that genuinely skips later - stages, see [`docs/STAGED_DEPLOYMENT.md`](docs/STAGED_DEPLOYMENT.md) - (reference implementation for Sequential backbones). -- `computation_used` / `estimated_backbone_flops_fraction` is a static - estimate: exit-head cost and routing overhead excluded; backbones that - reuse modules degrade to a warned low-confidence uniform estimate. +- `estimated_backbone_flops_fraction` is a static estimate of backbone + computation only: exit-head and routing costs are excluded, it is not + measured latency, power or energy, and backbones that reuse modules + degrade to a warned low-confidence uniform estimate. Ambiguous + architectures may not have a precise estimate at all. - Compute budgets from `calibrate_thresholds_for_budget` hold as an average over the calibration distribution. Nothing caps FLOPs per sample. -- Benchmarks are CIFAR-10 on a laptop GPU, and the published table predates - the v0.3 fair runner (see the legacy labels). ImageNet-scale numbers and a - real Jetson table don't exist yet; treat any wall-clock claim accordingly. +- The published benchmark is one bounded CIFAR-10 run on one laptop GPU. + ImageNet-scale numbers and a real Jetson table don't exist yet; treat any + wall-clock claim accordingly. ## Checkpoint migration -Checkpoints are versioned (`format_version: 2`). Files written by earlyon -≤ 0.2 load with an automatic, deterministic migration (scalar temperature -broadcast per head; legacy "disabled" threshold sentinels become explicit -`enabled_exits=False`) and a warning describing what changed — pinned -against a genuine v0.2-written fixture in the test suite. Details and code -changes: [`docs/MIGRATION.md`](docs/MIGRATION.md). +Checkpoints are versioned (`format_version: 2`) and record the exit +configuration, per-head temperatures, enabled exits, thresholds and routing +policy, plus reconstruction metadata where supported. Files written by +earlyon ≤ 0.2 load with an automatic, deterministic migration (scalar +temperature broadcast per head; legacy "disabled" threshold sentinels become +explicit `enabled_exits=False`) and a warning describing what changed — the +migration is verified against a genuine prior-version checkpoint in the +test suite. Keep a backup before migrating checkpoints you care about, and +read [`docs/MIGRATION.md`](docs/MIGRATION.md) for the code-level changes +(including where backward-compatible aliases like `computation_used` are +documented). ## Reproducibility @@ -362,11 +409,14 @@ Bug reports and benchmark results from your hardware are welcome. Start with ## Roadmap -- ImageNet-scale benchmark runs, and a measured Jetson Orin table to replace - the "profile it yourself" answer. -- Masked per-sample routing inside a batch (the v0.3 target). -- Grow the backbone factory list as people ask; `custom_ee` covers the gap - meanwhile. +Broad feature development is currently frozen. Planned work is narrow and +evidence-driven: + +- real Jetson measurements (the documented procedure, run on a device); +- TensorRT integration only when backed by measured evidence; +- user-reported compatibility fixes and PyTorch/torchvision compatibility; +- broader staged-runtime support only when justified by real demand; +- additional reproducible hardware evidence (including ImageNet-scale runs). ## Used By @@ -393,26 +443,6 @@ BranchyNet (Teerapittayanon et al. 2016) and the ACM 2024 early-exit survey (10.1145/3698767) for the ideas; torchvision (BSD) for the backbones; fvcore for FLOPs accounting. earlyon itself is MIT. -## Appendix: wall-clock latency (legacy) - -**Legacy numbers**: measured with the pre-v0.3 runner, which benchmarked the -wrapper and backbone on *different* random inputs — not methodologically -comparable to `benchmark_models` results, which feed every compared model -the identical sample sequence. Kept for transparency until re-measured. - -Throughput below uses a random-noise input, which can trigger spurious early -exits, so read it as a best-case bound rather than a claim. RTX 4050 Laptop -GPU, batch 1, 224×224, 50-iteration warmup, 300 iterations. - -| Model | Backbone p50 | Wrapper p50 (noise input) | -|-------------|-------------:|--------------------------:| -| ResNet18 | 1.32 ms | 0.51 ms | -| ResNet50 | 2.81 ms | 3.03 ms | -| MobileNetV2 | TBD ms | TBD ms | - -Reproducible via `scripts/re_evaluate.py`; raw per-run data in -[`docs/benchmarks.json`](docs/benchmarks.json). - ## License MIT — see [LICENSE](LICENSE). diff --git a/assets/banner.svg b/assets/banner.svg index d2fece7..f5cef81 100644 --- a/assets/banner.svg +++ b/assets/banner.svg @@ -1,6 +1,6 @@ - + earlyon - Masthead for earlyon. An easy input enters layer1 and leaves the network at exit 0 (12% of compute); layer2, layer3 and layer4 stay unused. + Masthead for earlyon. Illustration of early exit: an easy input enters layer1 and leaves the network at exit 0; layer2, layer3 and layer4 stay unused. @@ -18,7 +18,7 @@ earlyon - Early-exit inference for PyTorch CV models. + Train, calibrate, and benchmark early-exit vision models. @@ -47,6 +47,6 @@ exit 0 - 12% + taken diff --git a/assets/demo.svg b/assets/demo.svg index 08dffd2..0805b04 100644 --- a/assets/demo.svg +++ b/assets/demo.svg @@ -1,6 +1,6 @@ - - earlyon — inference session - Looping Python REPL. Beat 1: an easy input exits at head 0, using 12% of the network. Beat 2: a hard input takes no head and runs to the final classifier. A spine diagram beside it lights the short path, then the full path. + + earlyon — inference session (illustrative) + Looping Python REPL with illustrative values. Beat 1: an easy input exits at head 0 and later layers never run. Beat 2: a hard input takes no head and runs to the final classifier. A spine diagram beside it lights the short path, then the full path. @@ -13,9 +13,9 @@ - + - + @@ -40,17 +40,17 @@ r = model(car_closeup, mode="inference") - r.exit_taken, r.computation_used, r.confidence - (0, 0.12, 0.97) - # confident at the first head. 88% of the network never ran. + r.exit_taken, r.estimated_backbone_flops_fraction + (0, 0.12) # illustrative estimate + # confident at the first head; the later layers never ran. r = model(blurry_bird, mode="inference") - r.exit_taken, r.computation_used, r.confidence - (-1, 1.00, 0.91) + r.exit_taken, r.estimated_backbone_flops_fraction + (-1, 1.00) # no head was sure enough. hard inputs go the distance. @@ -107,10 +107,7 @@ exit 0 - 12% exit 1 - 40% exit 2 - 74% diff --git a/docs/evidence/CUDA_EVIDENCE.md b/docs/evidence/CUDA_EVIDENCE.md index 4dfa315..29f1873 100644 --- a/docs/evidence/CUDA_EVIDENCE.md +++ b/docs/evidence/CUDA_EVIDENCE.md @@ -1,5 +1,18 @@ # Bounded CUDA evidence run — 2026-07-22 +> **Erratum (2026-07-22, [#6](https://github.com/sohams25/earlyon/issues/6)).** +> The "estimated backbone FLOPs fraction" figures in this run come from the +> *low-confidence uniform fallback*, not the fvcore attribution: the reuse +> detector false-positives on torchvision ResNet's shared in-block (zero-FLOP) +> ReLU modules, so `resnet18_ee` was assigned uniform per-exit fractions +> (0.333/0.667) with `FlopsEstimate.reliable=False` and a RuntimeWarning. +> Under fvcore leaf-walk attribution the per-exit fractions are 0.547/0.774 +> and the run's average estimated fraction is **0.918, not the 0.880 recorded +> below** — i.e. the run saved roughly **8%** of estimated backbone compute, +> not 12%. Accuracy, latency, throughput and exit-distribution numbers are +> measurements and are unaffected. Raw JSON is preserved as recorded, with an +> `erratum` field appended. + Machine-readable source: [`cuda_evidence.json`](cuda_evidence.json). Reproduce: `python scripts/evidence_run.py` (seed 42; ~10 min compute on the hardware below after the CIFAR-10 download). diff --git a/docs/evidence/cuda_evidence.json b/docs/evidence/cuda_evidence.json index 1f8c69e..b1ca9ba 100644 --- a/docs/evidence/cuda_evidence.json +++ b/docs/evidence/cuda_evidence.json @@ -259,5 +259,14 @@ "gpu": "NVIDIA GeForce RTX 4050 Laptop GPU", "earlyon": "0.2.0", "os": "Linux-6.8.0-106-generic-x86_64-with-glibc2.35" + }, + "erratum_2026_07_22": { + "issue": "https://github.com/sohams25/earlyon/issues/6", + "affected_fields": [ + "test.early_exit_estimated_flops_fraction", + "benchmark_real_input.*.avg_estimated_flops_fraction", + "benchmark_noise_input.*.avg_estimated_flops_fraction" + ], + "note": "estimated FLOPs fractions in this run derive from the low-confidence uniform fallback (reuse detector false-positive on torchvision ResNet's zero-FLOP shared in-block ReLU; FlopsEstimate.reliable=False). Under fvcore leaf-walk attribution the per-exit fractions are layer2=0.5474, layer3=0.7736 and the test-set average estimated fraction is 0.918 instead of the recorded 0.880 (~8% estimated saving, not 12%). Measured accuracy/latency/throughput/exit-distribution values are unaffected. Original recorded fields left unmodified." } } \ No newline at end of file diff --git a/docs/marketing/earlyon-reddit.md b/docs/marketing/earlyon-reddit.md index a219cb6..0afb28e 100644 --- a/docs/marketing/earlyon-reddit.md +++ b/docs/marketing/earlyon-reddit.md @@ -1,3 +1,7 @@ +> NOTE (v0.3.0): this is archived v0.2-era launch copy. The numbers below +> predate the v0.3 fair benchmark runner and are superseded by +> docs/evidence/CUDA_EVIDENCE.md; do not reuse them. + # r/MachineLearning launch draft Flair: [P] (Project). Post as text, not a link post; r/ML buries bare links. @@ -34,7 +38,7 @@ model = resnet50_ee(num_classes=10, pretrained=True).eval() result = model(x, mode="inference") result.exit_taken # which head fired (-1 = full network) -result.computation_used # fraction of FLOPs actually run +result.estimated_backbone_flops_fraction # estimated fraction of backbone FLOPs (not measured latency) result.confidence # how sure it was when it left ```