Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion skills.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -151,7 +151,7 @@
- name: performance-patterns
maintainer: "napetrov"
external-repo: "https://github.com/intel/intel-performance-skills"
external-commit: "e9d0b6410fb1ad7a50fb81e0868fd23ae886882c"
external-commit: "3e33aecabb956ec48526b3f842f4f70ae96d98ff"
external-path: "skills/performance-patterns"
external-license: "MIT"
# Empty, and it earns the warning the validator prints for that. This skill covers
Expand Down
2 changes: 1 addition & 1 deletion skills/performance-patterns/.source.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"repo": "https://github.com/intel/intel-performance-skills",
"path": "skills/performance-patterns",
"commit": "e9d0b6410fb1ad7a50fb81e0868fd23ae886882c",
"commit": "3e33aecabb956ec48526b3f842f4f70ae96d98ff",
"license": "MIT",
"modified-files": [
"patterns/cold-path-annotation.md"
Expand Down
17 changes: 5 additions & 12 deletions skills/performance-patterns/patterns/cold-path-annotation.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,18 +58,11 @@ when it matches any of these patterns:

## Why the hot path suffers without it

Without the annotation the compiler has no way to know how likely each branch
is. It interleaves the cold-path instructions with the hot-path instructions in
program order. This has two costs for the caller:

1. **Instruction-cache pollution.** The cold-path instructions occupy cache
lines. Every time the hot path runs, it potentially evicts useful hot-path
instructions to make room for code that almost never executes.

2. **Branch predictor pressure.** The compiler generates generic branch
sequences. With the annotation it can emit the branch in a form the CPU's
static branch predictor recognizes as "almost never taken," saving a
mis-prediction penalty.
Without the annotation, the compiler interleaves cold-path instructions with
hot-path instructions in program order, which pollutes the instruction cache
(cold code evicts useful hot-path lines) and denies the compiler the chance to
emit a branch sequence the CPU's static predictor recognizes as "almost never
taken."

---

Expand Down
43 changes: 8 additions & 35 deletions skills/performance-patterns/patterns/cv-thundering-herd.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,41 +52,14 @@ shows normal hold/wait times). The key differentiator:

## Why this is slow

When `notify_all()` fires with N waiters:

```
Thundering Herd: notify_all with N waiters

t0: notify_all() → N threads become runnable simultaneously
t1: All N race for mutex re-acquisition (mandatory by CV semantics)
→ only 1 wins, N-1 immediately block on the mutex
t2: Each of the N-1 losers pays a full futex round-trip
(wake → schedule → attempt acquire → fail → sleep)
t3: Threads wake one-by-one as holder releases mutex
Most re-check predicate, find nothing, go back to sleep
```

The cost per `notify_all` call:
- O(N) context switches
- O(N) futex syscalls
- O(N) scheduler IPI dispatches (cross-core interrupts)
- Burst of RFO traffic on the mutex cache line as N cores attempt acquire

With a sequential `notify_one` loop, each call is a separate futex syscall +
IPI + context switch. At 160 threads: `T_wakeup ≈ N × 5µs ≈ 800µs` of pure
wakeup overhead per dispatch round.

**Why this worsens super-linearly with core count.** If T threads wake for J
jobs (J << T):

```
Wasted syscalls per round = T - J (grows with T)
Failed mutex acquisitions ≈ T - J (grows with T)
yield()/re-block calls ≈ T - J (grows with T)
```

On a 64-core system with 40 jobs, 24 threads waste — modest. On 160 cores,
120 threads waste — 75% of all wakeup effort is pure overhead.
`notify_all()` wakes all N waiters, but CV semantics force them to serialize on
mutex re-acquisition — only one wins, the rest pay a full futex round-trip
(wake → schedule → fail → sleep again) for nothing. Cost per call: O(N) context
switches, O(N) futex syscalls, O(N) scheduler IPIs. At 160 threads this is
roughly `T_wakeup ≈ N × 5µs ≈ 800µs` of pure overhead per dispatch round — and
if only J << T threads have actual work, the wasted fraction (`T - J`) grows
with core count, so a 160-core system with 40 jobs wastes ~75% of all wakeup
effort.

---

Expand Down
11 changes: 5 additions & 6 deletions skills/performance-patterns/patterns/false-sharing.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,12 +32,11 @@ causing spurious invalidations.

## Why false sharing hurts under scaling

A cache line is the smallest unit of coherence — typically 64 bytes. When thread A
writes field X and thread B writes field Y — even though X and Y are completely
unrelated — both writes invalidate each other's cached copy of the **entire** line.
Every write forces all other holders to reload the full 64 bytes from the L3 or
memory. This traffic grows linearly with thread count, which is why the function
only becomes prominent in a multi-core profile.
A cache line (typically 64 bytes) is the smallest unit of coherence, so a write
to unrelated field X by thread A and field Y by thread B still invalidates the
whole line for the other thread, forcing a full reload from L3/memory. This
traffic scales linearly with thread count, which is why it only shows up as
prominent in a multi-core profile.

---

Expand Down
26 changes: 4 additions & 22 deletions skills/performance-patterns/patterns/missing-restrict.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,28 +48,10 @@ portable.

## Why this is slow

The C standard allows any two pointers of compatible type to alias each other.
When the compiler sees:

```c
void add(float *a, float *b, float *dst, int n) {
for (int i = 0; i < n; i++)
dst[i] = a[i] + b[i];
}
```

It cannot prove that `dst` doesn't overlap with `a` or `b`. A write to
`dst[i]` could change the value that `a[i+1]` or `b[i+1]` reads on the next
iteration. To handle this safely the compiler either:

1. Generates **two loop versions** and a runtime overlap check — the vectorized
path is taken only when the check confirms no aliasing. This adds ~10–20
instructions of preamble before every call and increases i-cache footprint.
2. Stays **fully scalar** if the compiler's cost model decides the versioned
approach is not worth it.

Either outcome wastes cycles on alias bookkeeping that the programmer knows is
unnecessary.
Without `restrict`, the compiler must assume `dst` could alias `a` or `b`, so it
either emits a runtime overlap check plus two loop versions (~10–20 instructions
of preamble, extra i-cache pressure) or abandons vectorization entirely — cycles
spent on alias bookkeeping the programmer already knows is unnecessary.

---

Expand Down
19 changes: 6 additions & 13 deletions skills/performance-patterns/patterns/missing-vzeroupper.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,19 +38,12 @@ on the first SSE instruction after the AVX section.

## Why this is slow

When the CPU sees a write to a YMM (or ZMM0–15, which aliases the same state)
register, it marks those registers' upper bits as potentially "dirty". Legacy SSE
instructions treat the upper 128 bits of XMM registers as undefined and do not
zero them. When the CPU transitions from a code section that has dirty upper bits
to a section that uses legacy SSE, it must save and restore the full 256-bit (or
512-bit) register state, even though the SSE instruction only needs 128 bits.

The penalty is paid on the **first SSE instruction** after the dirty-upper-bit
state is set, and costs hundreds of cycles on some microarchitectures.

`vzeroupper` explicitly zeroes the upper 128 bits of all YMM registers (clearing
the dirty state) at negligible cost (~1 cycle). It must be executed before any
return or call path that may reach SSE code.
Writing to a YMM/ZMM0-15 register marks its upper bits "dirty"; legacy SSE
instructions leave the upper 128 bits of XMM undefined, so the CPU must save
and restore the full register state on the **first SSE instruction** after the
dirty state is set — costing hundreds of cycles on some microarchitectures.
`vzeroupper` clears the dirty state for ~1 cycle and must run before any
return/call path that may reach SSE code.

---

Expand Down
42 changes: 7 additions & 35 deletions skills/performance-patterns/patterns/mutex-to-rwlock.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,41 +40,13 @@ they do not conflict with each other.

## Why this is slow

A mutex is an **exclusive lock**: every acquisition moves the lock cache line to
Modified state via a LOCK CMPXCHG (Read-For-Ownership). Even two threads that
only read the protected data must take turns:

```
Mutex with N readers (all serialize):

Thread 1: [LOCK CMPXCHG → M state] read data [unlock → store]
Thread 2: ← waits (RFO pending) → [LOCK CMPXCHG → M state] read data [unlock]
Thread 3: ← waits ──────────────── ← waits (RFO pending) → [acquire] read [unlock]
...
Thread N: ← waits for all N-1 predecessors
```

At HCC scale (100+ cores), this serialization is catastrophic:

1. **O(N) wait time per reader** — each reader must wait for all preceding
readers to release, even though no data is being modified
2. **OSQ spin burns cycles** — the kernel mutex optimistic spin queue keeps
threads spinning on their MCS nodes while the holder is running; with many
readers each holding briefly, the aggregate spin time is enormous
3. **Cache-line bouncing** — the mutex's internal state transitions
(locked → unlocked → locked) force the lock cache line to bounce between
cores via the LLC/CHA, adding coherence latency to every handoff

The Linux kernel mutex implementation has three acquisition phases:
1. **Fast path** — single LOCK CMPXCHG; succeeds if mutex is unlocked
2. **Midpath (OSQ)** — optimistic spinning on a per-CPU MCS node while the
owner is running on another CPU; avoids the cost of sleeping
3. **Slow path** — thread is added to the wait queue and calls `schedule()`
(sleeps); woken by the holder on unlock via `wake_up_process()`

When `osq_lock` dominates perf, threads are stuck in the midpath — spinning
because the mutex holder is running (doing its read-only work) but not
releasing fast enough for the queue of waiters.
A mutex forces every acquisition — even read-only ones — to take exclusive
ownership of its cache line via `LOCK CMPXCHG`, so N concurrent readers fully
serialize. At HCC scale (100+ cores) this means O(N) wait time per reader,
kernel mutexes burning cycles in optimistic (OSQ) spin, and the lock's cache
line bouncing between cores on every handoff. When `osq_lock` dominates
`perf`, threads are stuck in this midpath spin because the holder (doing brief
read-only work) isn't releasing fast enough for the queue of waiters.

---

Expand Down
14 changes: 6 additions & 8 deletions skills/performance-patterns/patterns/parallel-accumulator.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,14 +29,12 @@ independent.

## Why this is slow

Modern CPUs can issue multiple FP operations per cycle, but only if those
operations are **independent**. A single accumulator forces strictly sequential
execution: the add at iteration `i+1` cannot begin until the add at iteration `i`
retires. The CPU's out-of-order engine stalls waiting for the dependency to
resolve, giving throughput limited by FP add **latency** (~4–5 cycles) rather than
FP add **throughput** (~0.5 cycles). This is an 8–10× gap on modern hardware.

Example serial pattern:
A single accumulator forces strictly sequential execution — each `+=` depends
on the previous result — so throughput is limited by FP add **latency**
(~4–5 cycles) instead of **throughput** (~0.5 cycles), an 8–10× gap on modern
out-of-order hardware that can otherwise issue multiple independent FP ops per
cycle.

```c
float sum = 0.0f;
for (int i = 0; i < n; i++)
Expand Down
15 changes: 6 additions & 9 deletions skills/performance-patterns/patterns/per-cpu-stats.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,15 +31,12 @@ false sharing, where threads write different fields on the same line.

## Why frequent shared stats hurt under scaling

Every atomic increment on a shared counter requires the updater to hold the cache
line **exclusively**. With N threads all updating the same counter:

- N-1 threads must wait for the exclusive transfer on every update
- The cache line bounces between LLC slices at near-memory latency (~100–300 ns)
- The more threads, the more bouncing — this is the defining signature of
scaling-limited true sharing
- Even `memory_order_relaxed` does not help: the **hardware** still enforces
exclusive ownership for any write, regardless of the software memory order
Every atomic increment on a shared counter requires exclusive ownership of its
cache line, so with N threads updating it, N-1 must wait for each exclusive
transfer and the line bounces between LLC slices at near-memory latency
(~100–300 ns) — worse as thread count grows. `memory_order_relaxed` doesn't
help: the hardware still enforces exclusive ownership for any write regardless
of software memory order.

---

Expand Down
25 changes: 5 additions & 20 deletions skills/performance-patterns/patterns/ttas.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,26 +30,11 @@ high cache-coherence traffic under contention.

## Why this is slow

A simple Test-and-Set loop:

```c
/* Test-and-Set — DO NOT USE under contention */
while (!cmpxchg(&lock, UNLOCKED, LOCKED))
_mm_pause();
```

`cmpxchg` always acquires the cache line **exclusively**, even on failure. When
thread A holds the lock and threads B and C are spinning:

- B and C repeatedly race each other for exclusive ownership of the lock's cache line
- The cache line bounces between B and C at high frequency
- This *also* steals the line away from thread A — even when A is trying to
release the lock
- If protected data shares the same cache line as the lock, that data is caught
in the same bounce (shows as false sharing in `perf c2c`)

The more waiters, the worse this scales — bus traffic grows as O(N²) under
contention.
`cmpxchg` always acquires the cache line **exclusively**, even on failure, so
every spinning waiter fights the others (and the lock holder) for the line.
Bus traffic grows as **O(N²)** under contention, and any protected data sharing
the lock's cache line gets caught in the same bounce (visible as false sharing
in `perf c2c`).

---

Expand Down
20 changes: 9 additions & 11 deletions skills/performance-patterns/triggers/from-profile.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,10 +62,10 @@ Read `patterns/missing-vzeroupper.md`.

The accumulate instruction (e.g., `vaddss`, `vaddpd`, `vmulss`, `vfmadd213ps`)
appears at the top of the `perf annotate` cycle-count column for a tight loop.
IPC from `perf stat` is well below 1.0, yet cache-miss rates are low — the CPU
is not waiting for memory, it is waiting for the previous iteration's result.
Cycles-per-iteration is at or above the FP latency of the operation (typically
4–5 cycles for `vadd`/`vfma`), even though the loop body is short.
IPC from `perf stat` is well below 1.0 with low cache-miss rates — a dependency
stall, not a memory stall. Cycles-per-iteration is at or above the FP latency
of the operation (typically 4–5 cycles for `vadd`/`vfma`), even though the loop
body is short.

Read `patterns/parallel-accumulator.md`.

Expand Down Expand Up @@ -139,13 +139,11 @@ Read `patterns/fast-crc32c.md`.
### SIMD sort

`perf report` shows `std::sort`, `_introsort_loop`, `__gnu_cxx::__ops`,
`std::__introsort_loop`, or `std::__sort` among the hottest symbols, and the
sorted data type is a numeric primitive (`float`, `double`, `int32_t`,
`uint32_t`, `int64_t`, `uint64_t`). `perf stat` may also show elevated
`branch-misses` — the comparator-driven branches of introsort are notoriously
hard for the branch predictor. The bottleneck is comparison and partitioning
overhead, not memory bandwidth; replacing with x86-simd-sort gives 3–8×
speedup by vectorizing both steps with AVX-512/AVX2.
`std::__introsort_loop`, or `std::__sort` among the hottest symbols, on a
numeric primitive type (`float`, `double`, `int32_t`, `uint32_t`, `int64_t`,
`uint64_t`). `perf stat` may also show elevated `branch-misses` — introsort's
comparator branches are hard to predict. See `patterns/simd-sort.md` for the
x86-simd-sort replacement and expected speedup.

Read `patterns/simd-sort.md`.

Expand Down
Loading
Loading